Keyword Research as a Proxy for AI Search Demand: How Far It Actually Holds

Every AI visibility product has a denominator problem, and most of us would rather not talk about it.
Before you can tell a brand it is missing from AI answers, you have to decide which answers to check. That means choosing topics, and choosing topics means starting from some model of what people ask. Nearly everyone in this category starts where we did: a keyword universe, seeded from the customer's own site and their competitors', expanded through search advertising APIs into a few thousand phrases with volumes attached.
It is a reasonable place to start, and it is a proxy for something it cannot see directly. A keyword database holds demand that has already been typed into a search box, and nothing else. Whether that still describes how people ask AI engines is an empirical question, so this year we went and measured it.
Further than the sceptics claim, it turns out. Nowhere near far enough to build on unexamined.
If you only take one line from this:
Keyword data is a proxy for AI demand, and how far it holds depends entirely on how much the engine rewrites the prompt before retrieving. It holds well for one major engine, badly for another, and may not apply at all to a third. Treat it as the territory instead and you break three things: which topics you measure, what a visibility score means, and which words end up in your content.
- Keyword coverage behaves differently on every engine. It predicts Perplexity well, ChatGPT badly, and on Claude it may not apply at all, since Claude often answers with no search behind it.
- How you rank keyword groups decides what you end up measuring. Rank groups by their loudest keyword and you systematically select product categories over service businesses, which can leave a specialist measured against a market they do not operate in.
- Never let the string you measure be the string you publish. A keyword is a measurement handle and a piece of copy. One field cannot do both jobs without one of them corrupting the other.
1. What search advertising data can and cannot contain
Ours starts from search advertising data, which is the most thoroughly measured demand signal marketing has. Seeds come from the customer's domain and their competitors', expanded through Google Ads' keyword APIs and a third-party suggestion service, then embedded, clustered into a two-level taxonomy, and labelled. We have written before about the clustering half of that pipeline, including why k-means, HDBSCAN and the Leader algorithm all fell over on dense embedding spaces and Leiden community detection did not.
That data is real, measured at enormous scale, and for search it is exactly the right instrument. It also has a defined edge, and the edge is structural rather than a question of quality: an advertising platform can only record a volume against a phrase somebody has already typed into a search box, so every row in that universe is, by construction, a thing people search for. Knowing precisely where that edge sits is what lets you build on the data with confidence instead of hope, and that is what the rest of this is about.
2. How much of AI demand a keyword database can see
Two pieces of research from this year answer that, and between them they measure the proxy from opposite ends: how much AI demand keyword data misses, and how far the demand it does capture survives contact with an engine. Neither is ours, and both are large enough to take seriously.
The first is a 17-month clickstream analysis running to February 2026, covering more than a billion lines of US panel data, which cross-referenced observed ChatGPT prompts against a database of over 27 billion keywords. For most of the study period, somewhere between 65% and 85% of prompts matched no traditional search keyword at all.
The example that study uses makes the gap obvious. A search query is "best project management software". The ChatGPT version of the same need reads more like: "I manage a 12-person remote engineering team and we're constantly missing sprint deadlines. What should I change about our weekly standups?" No keyword database contains that sentence. None ever will, because nobody types it into a search box.
The same research carries two caveats that most write-ups skip, and both cut against the headline.
The gap is closing rather than widening. Prompts using recognisably traditional search language nearly doubled between October 2025 and February 2026, going from 18.9% to 34.9%. As AI engines absorb navigational and transactional traffic, prompts get shorter and start to look more like queries.
Retrieval often isn't the path at all. As of February 2026, ChatGPT turned on web search for only 34.5% of queries, down from 46% in late 2024. Most responses lean on training data, where there is no retrieval step to optimise for and no citation to win. That is a ceiling on this entire category, ours included.
The study that changed how we build tracked 10,000 prompts across 14 days in spring 2026, running each against ChatGPT, Perplexity and Copilot, and captured the queries each engine actually sent to its own retrieval layer.

Word overlap between the user's prompt and the query the engine issues, alongside how often those queries repeat. Source: a 10,000-prompt fan-out study run over 14 days in spring 2026.
This is the finding that reframed the problem for us, because it says the proxy does not fail evenly. It fails per engine, in an order you can measure.
Perplexity behaves like a literal retriever. Its internal queries are close to a tidied-up version of what the user typed, and it fires largely the same ones run after run. For Perplexity, keyword coverage genuinely does predict what will be searched. Copilot compresses: roughly half the original words survive, qualifiers get stripped, and the prompt is squeezed into something closer to a conventional query. ChatGPT is doing something different in kind. At 13% overlap, its most divergent query shares about one word in eight with the prompt, and it almost never fires the same string twice.
What about Gemini and Claude?
That study covered three engines. Two of the biggest names in the category are missing from it, and they behave differently enough that leaving them out would be misleading.
Google's AI Mode, running on Gemini, fans out far harder than anything else measured. Separate 2026 research testing Gemini 3 across 501 prompts found an average of 10.7 sub-queries per prompt, against roughly 2.3 to 2.8 for ChatGPT in the same tests. Google's AI Mode breaks a question into intents, sub-intents and related questions, then runs them in parallel and synthesises across the whole set rather than any single one.

How many searches each engine runs behind a single prompt. The Gemini and ChatGPT figures come from one 2026 study of 501 prompts; the Perplexity and Copilot figures come from the 10,000-prompt fan-out study, so the two are not strictly like for like.
Because the Gemini figure comes from a separate 501-prompt benchmark rather than the 10,000-prompt study behind the other engines, treat these absolute counts as directional rather than measured on one identical baseline. The order-of-magnitude difference is the finding, not the decimal.
The practical consequence is not subtle. On Google's AI Mode, a single question opens roughly ten separate doors into the index, and your page only has to be the best answer behind one of them to end up in the synthesis. Breadth of coverage matters more there than precision on any one phrase. On Perplexity, where one or two near-literal queries do all the work, the opposite holds.
Claude is a different case again, and the interesting part is that it often does not search at all. Anthropic's own documentation is unusually direct about this: Claude searches for recent events, current prices or statistics, and information about organisations, people or products that may have changed. It answers directly, with no retrieval and therefore no citation, for established facts, creative work, analysis of material already in the conversation, and ordinary conversational turns. When it does search, it writes its own query rather than passing the user's words through, and it can search several times in a turn, refining as it goes.
For anyone measuring visibility, that is a real constraint rather than a footnote. A share of Claude conversations about your category offers no citation opportunity to anyone, because no retrieval happens. Absence from those answers is not a content gap you can close.
What nobody has published, as far as we can find, is a word-overlap measurement for Gemini or Claude equivalent to the three-engine study above. We are not going to estimate one. The honest position is that we know how much three of the five major engines rewrite, we know how broadly a fourth fans out, and we know the fifth frequently declines to search at all.
Pulling it together:
| Engine | What happens to your prompt | What that means for keyword coverage | What to optimise for |
|---|---|---|---|
| Perplexity | Keeps 88% of the words, repeats the same queries run after run | Strong proxy. What you rank for is close to what it searches | Exact-match keyword clustering. Conventional targeting still works |
| Copilot | Keeps about half the words, strips qualifiers, reshapes into a conventional query | Partial proxy | The unqualified core of each phrase, since qualifiers are dropped |
| ChatGPT | Keeps 13% of the words, almost never repeats a query, runs 2 to 3 searches | Weak proxy. It explores adjacent vocabulary you never targeted | Latent semantic coverage: adjacent concepts and alternate framings |
| Google AI Mode (Gemini) | Splits into around 10 parallel sub-queries and synthesises across all of them | Breadth beats precision. Coverage of a topic matters more than any single phrase | Entity-dense topic graphs. Ten shallow doors beat one deep one |
| Claude | Often answers with no search at all; writes its own query when it does search | No retrieval means no citation. Part of the category is simply unwinnable | Off-page brand mentions and PR. There is no page to retrieve |
We reached the same conclusion from a different direction while building gap analysis, where we had to key every gap by topic and provider rather than collapsing providers into one number. Two unrelated arguments landing on provider as a first-class dimension is about as much corroboration as this field offers.
So "keyword research is dead for AI search" is wrong, and we should stop repeating it. The accurate version is narrower and a lot more useful:
Keyword coverage is a strong proxy for Perplexity, a partial one for Copilot, and a weak one for ChatGPT. On Gemini it is the wrong shape of question entirely, and on Claude it may not apply at all. Any system treating it as one engine-independent denominator is averaging across five very different behaviours.
3. When search volume picks your topics
The first leak is in topic selection, and it is the more damaging of the two, because everything downstream inherits the topics.
The customer was a marketing agency specialising in social media strategy. Onboarding handed it five discovery topics: Data Analytics, Business Intelligence, Digital Strategy, Influencer, and PPC. Its actual practice, which is social media strategy, digital transformation, PR and reputation, employee advocacy, social listening and experience design, got nothing.
Three things combined to produce that outcome, and the third is where the real mechanism sits.
The brand alignment classifier doesn't discriminate. Every cluster is tagged for how well it fits the brand. On this project, 457 clusters came back tagged "on brand", 239 "mostly on brand", 88 "mostly off brand", and 5 "off brand".

A tier that holds 58% of the input cannot order what sits inside it. The dashed line marks where a gate on these tiers would fall.
So the sort does nothing. Cluster ordering is documented as alignment tier first, volume second. It is implemented as a stable sort on the tier alone, and a stable sort leaves items that tie in whatever order they were already in. These arrive already ordered by volume, because the query that fetches them sorts on the largest single keyword volume in each group before handing them over, and nothing re-sorts them afterwards. So when 457 groups tie on tier, that volume ordering underneath simply survives the sort untouched. "Alignment first" quietly becomes "volume first". The sorting logic is correct and does exactly what it was written to do; it simply has no work left, because the field it sorts on carries effectively one value.
And the volume it falls back on is not the cluster's total demand. This is the part we had wrong ourselves until we went back and read the query that produces the ranking. Clusters are ordered strictly by the search volume of their single biggest keyword, rather than by the cluster's aggregate demand.
That one detail decides the outcome.

Ranking on a cluster's largest member, not its total. Bar shapes are illustrative; the 74,000 figure is a real keyword from this project's universe.
Ranking a cluster by its largest member rewards clusters shaped like a product category, because product categories have one enormous head term with a tail hanging off it. It punishes clusters shaped like a professional service practice, which are broad and flat: lots of mid-volume phrasings, no single dominant term. A cluster holding "power bi software" at 74,000 searches a month beats a cluster of specialist terms carrying more demand in total but no comparable spike.
Look at what won. "power bi software" at 74,000 searches a month, "power business intelligence" (74,000), "analytics google analytics" (49,500), "googletagmanager" (18,100). Software product names and job titles. Precisely the shape a max-based ranking rewards.
There is a fourth contributor that makes all of this worse. Seeds drift into product categories in about two hops. The service "data analytics" expanded into the business intelligence software category, and that expansion spawned fresh seeds of its own: "Business Intelligence", "Power BI Analytics", "Business Intelligence Platforms". Nobody decided to research BI platforms. The expansion decided, one plausible step at a time.
The result was 0 out of 45 prompts present. And the entities that did show up in those answers were real business intelligence consultancies, which is to say the answers were correct. The engines did their job perfectly. We asked business intelligence questions, got business intelligence answers, and filed a social media agency's absence from the business intelligence market as an AI visibility failure.
A 0% built on off-brand topics is a tautology, not a finding.
The obvious response is to add a gate: only let the topic-choosing step see groups tagged "on brand" or "mostly on brand". On the distribution above, that admits 696 of 789 groups. It strips only the bottom 12% and leaves the failure entirely intact, because it inherits the same classifier that could not discriminate in the first place. It is an easy move to reach for and it feels like progress, which is what makes it a trap: a filter is only ever as good as the signal it filters on, so the work has to go into the signal itself.
A ranking that holds up needs three things. Score brand fit as a continuous number rather than four coarse buckets, by measuring how close a keyword group sits to the brand's own description rather than asking a model to drop it in a labelled box. Then rank on fit and volume together instead of one after the other, so that a group scoring twice as well on fit can outrank one with somewhat more search volume. Something of this shape:

Ranking on brand fit and total volume together, rather than on a tier followed by a group's loudest member.
Two things change here, and both matter. Brand fit becomes a real number from an embedding comparison rather than a bucket assigned by a model, so groups stop tying. And total volume replaces the largest single member, wrapped in a logarithm so that a 74,000-volume head term contributes more than an 18,000 one without swamping the fit score entirely. That combination is what lets a specialist's practice areas compete with a product category on equal terms. The third part is to cap how far a seed may spawn further seeds, so a service cannot quietly become a product category two hops later.
This generalises well beyond us. Every visibility number is a measurement of a chosen prompt set, and any tool choosing that set from search volume inherits the same bias. Which makes it a fair question to put to any vendor in this category, ours included: the first thing to ask isn't how many providers a tool tracks, but how it decided what to ask in the first place, and what it does when your business has no head term. A tool that cannot answer that is reporting on a prompt set nobody chose deliberately.
4. When a keyword row becomes published copy
The second leak is what happens when the same string is asked to be both a measurement and a sentence. Every step in the chain behaves correctly, and the article still comes out wrong.
The case is one of our own, generated by our pipeline against our own keyword universe, which is the only reason we can show it at this level of detail. The focus keyword was "content AI".
That is not a typo. It is a row in our keyword universe, word for word, sitting in a cluster of related rows that are all names for one competing WordPress SEO plugin's paid AI writing module. Although it is the genuine, current name of that product and spelled exactly right, as an English phrase it is backwards compared to the natural "AI content".
It went into the published draft untouched. Seven occurrences across 2,187 words, including in H2 and H3 headings.
Three independent guards existed for exactly this. All three had shipped about a week earlier. All three missed it, for the same reason.
| Guard | What it does | What it did here |
|---|---|---|
| Suggester prompt | Tells the model to emit a corrected form when the matched keyword is malformed, concatenated, misspelled or outdated | "content ai" is none of those, so the model fell back on the stronger instruction sitting in the same paragraph, which says the focus keyword must be one of the keywords in the universe, and echoed the row |
| Keyword normalizer | An LLM pass whose whole job is repairing bad keywords before use | Looked at "content ai", changed the capitalisation to "content AI", marked it valid, and passed it on |
| Keyword policy | Emits usage rules, then counts usage in the output | Its rule explicitly permits rewording to "AI content". Ignored. It then counted 7 uses against a cap of 3, logged that, and did nothing with it |
All three shared one blind spot. Every guard enumerated spelling defects: concatenated, misspelled, outdated, mis-cased. "content ai" is spelled perfectly. What is wrong with it is the meaning: the word order is backwards, and it is somebody else's product name. No layer had a concept for either.

Every guard checked spelling. The defect was meaning.
That is the general trap in building guardrails over a data source you don't control. You write rules against the failures you've already seen. A well-formed string that's wrong for reasons of meaning walks straight through a wall built out of spell checks.
Every layer did exactly what it was told, and four structural details downstream turned that one bad keyword into a bad article.
The brief is the transmission vector. The suggester is asked to express per-section keywords as natural search intent, never as raw query strings. It broke that in all five sections and wrote the literal phrase into each one. That brief is then injected verbatim into the generation prompt, so the raw string reaches the writer five times as concrete, in-context instruction. A concrete example in a prompt beats an abstract rule in the same prompt every single time.
The rubric only rewards presence.

The scorecard an article is graded against. Nothing on it costs you anything for repetition, which is how a keyword used seven times against a limit of three still comes out "Strong".
The draft had no secondary keywords at all, which scores as a pass on an empty list. A free 10 out of 10 for having none. With a single keyword, and rules demanding it in the title, the intro and a heading, there's no alternative phrasing to rotate through. The rubric doesn't just fail to punish repetition here. It pushes you toward it.
And none of it bought anything. The presence check was already lenient: it takes an exact phrase or a semantic token window. The article's real title never contained the phrase "content AI" at all, and passed regardless. Rewording it naturally would have scored the same. The pipeline carried a broken string all the way to publication in exchange for nothing at all.
There is a general lesson underneath all four. A generate-evaluate-repair loop only closes if the repair step is actually wired to act on what the evaluation found, and it is easy to build one where the score is computed, stored, and never consulted. That 76 out of 100, "Strong" was not a judgement anybody overruled. It was a number written down after the last point at which anything could still change.
5. Two strings, not one
The obvious move is to fix the keyword in place. Rewrite "content AI" to "AI content" on the way in, move on. That instinct is wrong, for a reason that transfers well beyond keywords.
The keyword is quietly doing two jobs. One is as a measurement handle: the thing tying this article back to the specific gap that triggered it, so visibility scoring can later ask whether writing it changed anything. The other is as copy, words a human reads in a heading. Rewriting in place fixes the second job by silently breaking the first, since the article would no longer be attributable to the demand it was written for.
So we carry both. One field holds the measured keyword, stored exactly as it came out of the database. A second holds the written keyword, the natural phrasing that actually goes into the article. The suggestion card shows only the written one. The scoring layer reads the measured one. A third field carries alternative phrasings, so a writer has something to rotate through instead of repeating one phrase in the title, the intro and a heading.
Concretely, one field doing three jobs becomes three fields doing one each:

The string you measure and the string you publish are not the same string.
Names here are illustrative rather than ours. The first field is what visibility gets measured against, because it is the string that came from the gap. The second is what the writer actually uses. The third is what stops the writer painting itself into a corner. Before the change, one field was doing all three jobs.
Two other decisions generalise.
Deterministic detection rather than another model call. An LLM normalizer was already in the path and it passed the keyword, so the flaw lay in the validation rule rather than in the mechanism running it. Bolting a second model call on to check the first compounds the problem rather than solving it. Deterministic detectors also have a property no LLM rewrite has: you can dry-run them across your entire history and see exactly what they would have caught and what they would have broken, before shipping anything.
Reword third-party product names to the generic capability, rather than dropping them. The demand behind that keyword row is real, since people do search for AI content optimisation tooling. Dropping the row loses the demand. Keeping it verbatim borrows a competitor's brand. Mapping it to the capability keeps one and drops the other.
6. How this squares with outside research
We built all of this from our own telemetry, which means the numbers are ours and the sample is small. What raises confidence is independent research arriving at adjacent conclusions for unrelated reasons.
| External finding | What it corroborates |
|---|---|
| 65% to 85% of ChatGPT prompts matched nothing in a 27-billion-keyword database (17-month clickstream study, over 1B lines of US panel data, through Feb 2026) | Why the keyword universe cannot be the denominator for AI demand. Ours is entirely search-advertising sourced, so it is structurally blind to the majority of prompts |
| Prompts using traditional search language nearly doubled Oct 2025 to Feb 2026, from 18.9% to 34.9%, same study | Why we are not deprecating the keyword universe. The proxy is improving rather than collapsing. What needs measuring is the error term |
| Word overlap between prompt and issued query: Perplexity 88%, Copilot 50%, ChatGPT 13% (10,000 prompts, 14 days, spring 2026) | Why proxy quality has to be assessed per provider instead of globally, and why one blended visibility number hides three different error rates |
| Query uniqueness across runs: ChatGPT 91%, Copilot 47%, Perplexity 14%, same study | Why run-to-run volatility on ChatGPT is a property of the engine rather than noise in our measurement, and why smoothing it away would hide real signal |
| ChatGPT enabled web search on 34.5% of queries as of Feb 2026, down from 46% | A ceiling on this whole category. Retrieval-shaped optimisation only addresses the minority of prompts where retrieval actually happens |
| Adding statistics, citations and quotations lifted citation visibility by 30% to 40%, while keyword density showed minimal effect (GEO: Generative Engine Optimization, peer-reviewed, KDD 2024) | Why a rubric spending 50 of 85 points on keyword presence was optimising the wrong property entirely |
The number to remember: adding statistics and source citations lifts generative visibility by 30% to 40%. Keyword density lifts it by roughly nothing.
Which reframes what a content scorecard is for. Any rubric weighted toward keyword presence is measuring the wrong property, however carefully it is calibrated, because it is counting a signal the engines largely ignore. Points spent on evidence density, on whether a passage carries a statistic, a source or a quotation, are points spent on something that demonstrably moves citation probability.
7. Where this leaves us
So, how far does it hold? Far enough to keep, and not far enough to trust blindly. It is a proxy with an error term that is now measurable rather than assumed, and that error term varies by engine, by the shape of your business, and by what you are asking the data to do. The data itself was never the problem; the damage came from forcing a single proxy to serve three incompatible jobs.
As the demand denominator, it holds up. With the per-engine caveats above, and improving over time.
As the topic selector, it needs a second signal beside it. Volume alone systematically picks product categories over service practices, and a coarse alignment tier is not strong enough to counteract that on its own.
As literal copy, it was never meant to serve. A database row is a measurement artifact. Nobody intended it to be read by a human, and the moment one turns up in a heading, something upstream has mistaken an index for a sentence.
If you are building on keyword data, four properties are worth designing for from the start, because each one is expensive to retrofit:
- Brand fit as a continuous score, never coarse buckets. Buckets that most of the input lands in cannot order anything, and a sort on them silently becomes a sort on whatever came next.
- Ranking on fit and volume jointly, never one after the other, and on a group's total rather than its loudest member.
- A hop limit on seed expansion, so a service cannot become a product category two steps later without anyone choosing that.
- Separate fields for the string you measure and the string you publish, so neither job can quietly corrupt the other.
The bigger piece, mining a prompt space directly instead of deriving one from keyword data, is a sibling of the keyword universe rather than a layer on top of it. Deriving an intent space from a keyword database inherits the aperture problem from section 2 and presents it more confidently, which is worse than not having one.
Further Reading
- GEO: Generative Engine Optimization. Aggarwal et al., Princeton, Georgia Tech and the Allen Institute for AI, KDD 2024, peer-reviewed. The GEO-bench study behind the statistics-and-citations lift figures and the finding that keyword density barely moves citation probability.
- How We Made Sense of Thousands of Keywords. Our write-up of the clustering half of this pipeline, and why graph community detection beat every threshold method we tried.
- From Analysis to Action: Turning AI Visibility Gaps Into a Content Roadmap. How we key gaps by topic and provider, and why the citation rather than the keyword cluster turned out to be the actionable unit.

