Keyword Research as a Proxy for AI Search Demand: How Far It Actually Holds

27 min read
Rakesh Menon
Keyword Research as a Proxy for AI Search Demand: How Far It Actually Holds

Every AI visibility product has a denominator problem, and most of us would rather not talk about it.

Before you can tell a brand it is missing from AI answers, you have to decide which answers to check. That means choosing topics, and choosing topics means starting from some model of what people ask. Nearly everyone in this category starts where we did: a keyword universe, seeded from the customer's own site and their competitors', expanded through search advertising APIs into a few thousand phrases with volumes attached.

It is a reasonable place to start, and it is a proxy for something it cannot see directly. A keyword database holds demand that has already been typed into a search box, and nothing else. Whether that still describes how people ask AI engines is an empirical question, so this year we went and measured it.

Further than the sceptics claim, it turns out. Nowhere near far enough to build on unexamined.

If you only take one line from this:

Keyword data is a proxy for AI demand, and how far it holds depends entirely on how much the engine rewrites the prompt before retrieving. It holds well for one major engine, badly for another, and may not apply at all to a third. Treat it as the territory instead and you break three things: which topics you measure, what a visibility score means, and which words end up in your content.

  • Keyword coverage behaves differently on every engine. It predicts Perplexity well, ChatGPT badly, and on Claude it may not apply at all, since Claude often answers with no search behind it.
  • How you rank keyword groups decides what you end up measuring. Rank groups by their loudest keyword and you systematically select product categories over service businesses, which can leave a specialist measured against a market they do not operate in.
  • Never let the string you measure be the string you publish. A keyword is a measurement handle and a piece of copy. One field cannot do both jobs without one of them corrupting the other.

1. What search advertising data can and cannot contain

Ours starts from search advertising data, which is the most thoroughly measured demand signal marketing has. Seeds come from the customer's domain and their competitors', expanded through Google Ads' keyword APIs and a third-party suggestion service, then embedded, clustered into a two-level taxonomy, and labelled. We have written before about the clustering half of that pipeline, including why k-means, HDBSCAN and the Leader algorithm all fell over on dense embedding spaces and Leiden community detection did not.

That data is real, measured at enormous scale, and for search it is exactly the right instrument. It also has a defined edge, and the edge is structural rather than a question of quality: an advertising platform can only record a volume against a phrase somebody has already typed into a search box, so every row in that universe is, by construction, a thing people search for. Knowing precisely where that edge sits is what lets you build on the data with confidence instead of hope, and that is what the rest of this is about.

2. How much of AI demand a keyword database can see

Two pieces of research from this year answer that, and between them they measure the proxy from opposite ends: how much AI demand keyword data misses, and how far the demand it does capture survives contact with an engine. Neither is ours, and both are large enough to take seriously.

The first is a 17-month clickstream analysis running to February 2026, covering more than a billion lines of US panel data, which cross-referenced observed ChatGPT prompts against a database of over 27 billion keywords. For most of the study period, somewhere between 65% and 85% of prompts matched no traditional search keyword at all.

The example that study uses makes the gap obvious. A search query is "best project management software". The ChatGPT version of the same need reads more like: "I manage a 12-person remote engineering team and we're constantly missing sprint deadlines. What should I change about our weekly standups?" No keyword database contains that sentence. None ever will, because nobody types it into a search box.

The same research carries two caveats that most write-ups skip, and both cut against the headline.

The gap is closing rather than widening. Prompts using recognisably traditional search language nearly doubled between October 2025 and February 2026, going from 18.9% to 34.9%. As AI engines absorb navigational and transactional traffic, prompts get shorter and start to look more like queries.

Retrieval often isn't the path at all. As of February 2026, ChatGPT turned on web search for only 34.5% of queries, down from 46% in late 2024. Most responses lean on training data, where there is no retrieval step to optimise for and no citation to win. That is a ceiling on this entire category, ours included.

The study that changed how we build tracked 10,000 prompts across 14 days in spring 2026, running each against ChatGPT, Perplexity and Copilot, and captured the queries each engine actually sent to its own retrieval layer.

Grouped bar chart comparing three engines on two measures. Perplexity keeps 88 percent of the prompt's words and only 14 percent of its queries are unique across runs. Copilot keeps 50 percent with 47 percent unique. ChatGPT keeps 13 percent with 91 percent unique.

Word overlap between the user's prompt and the query the engine issues, alongside how often those queries repeat. Source: a 10,000-prompt fan-out study run over 14 days in spring 2026.

This is the finding that reframed the problem for us, because it says the proxy does not fail evenly. It fails per engine, in an order you can measure.

Perplexity behaves like a literal retriever. Its internal queries are close to a tidied-up version of what the user typed, and it fires largely the same ones run after run. For Perplexity, keyword coverage genuinely does predict what will be searched. Copilot compresses: roughly half the original words survive, qualifiers get stripped, and the prompt is squeezed into something closer to a conventional query. ChatGPT is doing something different in kind. At 13% overlap, its most divergent query shares about one word in eight with the prompt, and it almost never fires the same string twice.

What about Gemini and Claude?

That study covered three engines. Two of the biggest names in the category are missing from it, and they behave differently enough that leaving them out would be misleading.

Google's AI Mode, running on Gemini, fans out far harder than anything else measured. Separate 2026 research testing Gemini 3 across 501 prompts found an average of 10.7 sub-queries per prompt, against roughly 2.3 to 2.8 for ChatGPT in the same tests. Google's AI Mode breaks a question into intents, sub-intents and related questions, then runs them in parallel and synthesises across the whole set rather than any single one.

Bar chart of searches run per prompt. Perplexity and Copilot sit at roughly 1.4 to 2, ChatGPT at 2.3 to 2.8, and Google AI Mode running Gemini at 10.7, about five times ChatGPT's count.

How many searches each engine runs behind a single prompt. The Gemini and ChatGPT figures come from one 2026 study of 501 prompts; the Perplexity and Copilot figures come from the 10,000-prompt fan-out study, so the two are not strictly like for like.

Because the Gemini figure comes from a separate 501-prompt benchmark rather than the 10,000-prompt study behind the other engines, treat these absolute counts as directional rather than measured on one identical baseline. The order-of-magnitude difference is the finding, not the decimal.

The practical consequence is not subtle. On Google's AI Mode, a single question opens roughly ten separate doors into the index, and your page only has to be the best answer behind one of them to end up in the synthesis. Breadth of coverage matters more there than precision on any one phrase. On Perplexity, where one or two near-literal queries do all the work, the opposite holds.

Claude is a different case again, and the interesting part is that it often does not search at all. Anthropic's own documentation is unusually direct about this: Claude searches for recent events, current prices or statistics, and information about organisations, people or products that may have changed. It answers directly, with no retrieval and therefore no citation, for established facts, creative work, analysis of material already in the conversation, and ordinary conversational turns. When it does search, it writes its own query rather than passing the user's words through, and it can search several times in a turn, refining as it goes.

For anyone measuring visibility, that is a real constraint rather than a footnote. A share of Claude conversations about your category offers no citation opportunity to anyone, because no retrieval happens. Absence from those answers is not a content gap you can close.

What nobody has published, as far as we can find, is a word-overlap measurement for Gemini or Claude equivalent to the three-engine study above. We are not going to estimate one. The honest position is that we know how much three of the five major engines rewrite, we know how broadly a fourth fans out, and we know the fifth frequently declines to search at all.

Pulling it together:

EngineWhat happens to your promptWhat that means for keyword coverageWhat to optimise for
PerplexityKeeps 88% of the words, repeats the same queries run after runStrong proxy. What you rank for is close to what it searchesExact-match keyword clustering. Conventional targeting still works
CopilotKeeps about half the words, strips qualifiers, reshapes into a conventional queryPartial proxyThe unqualified core of each phrase, since qualifiers are dropped
ChatGPTKeeps 13% of the words, almost never repeats a query, runs 2 to 3 searchesWeak proxy. It explores adjacent vocabulary you never targetedLatent semantic coverage: adjacent concepts and alternate framings
Google AI Mode (Gemini)Splits into around 10 parallel sub-queries and synthesises across all of themBreadth beats precision. Coverage of a topic matters more than any single phraseEntity-dense topic graphs. Ten shallow doors beat one deep one
ClaudeOften answers with no search at all; writes its own query when it does searchNo retrieval means no citation. Part of the category is simply unwinnableOff-page brand mentions and PR. There is no page to retrieve

We reached the same conclusion from a different direction while building gap analysis, where we had to key every gap by topic and provider rather than collapsing providers into one number. Two unrelated arguments landing on provider as a first-class dimension is about as much corroboration as this field offers.

So "keyword research is dead for AI search" is wrong, and we should stop repeating it. The accurate version is narrower and a lot more useful:

Keyword coverage is a strong proxy for Perplexity, a partial one for Copilot, and a weak one for ChatGPT. On Gemini it is the wrong shape of question entirely, and on Claude it may not apply at all. Any system treating it as one engine-independent denominator is averaging across five very different behaviours.

3. When search volume picks your topics

The first leak is in topic selection, and it is the more damaging of the two, because everything downstream inherits the topics.

The customer was a marketing agency specialising in social media strategy. Onboarding handed it five discovery topics: Data Analytics, Business Intelligence, Digital Strategy, Influencer, and PPC. Its actual practice, which is social media strategy, digital transformation, PR and reputation, employee advocacy, social listening and experience design, got nothing.

Three things combined to produce that outcome, and the third is where the real mechanism sits.

The brand alignment classifier doesn't discriminate. Every cluster is tagged for how well it fits the brand. On this project, 457 clusters came back tagged "on brand", 239 "mostly on brand", 88 "mostly off brand", and 5 "off brand".

Bar chart of brand alignment tiers across 789 clusters. On brand 457 at 58 percent, mostly on brand 239 at 30 percent, mostly off brand 88, off brand 5. A marked line shows the gate admitting the first two tiers, 696 clusters or 88 percent.

A tier that holds 58% of the input cannot order what sits inside it. The dashed line marks where a gate on these tiers would fall.

So the sort does nothing. Cluster ordering is documented as alignment tier first, volume second. It is implemented as a stable sort on the tier alone, and a stable sort leaves items that tie in whatever order they were already in. These arrive already ordered by volume, because the query that fetches them sorts on the largest single keyword volume in each group before handing them over, and nothing re-sorts them afterwards. So when 457 groups tie on tier, that volume ordering underneath simply survives the sort untouched. "Alignment first" quietly becomes "volume first". The sorting logic is correct and does exactly what it was written to do; it simply has no work left, because the field it sorts on carries effectively one value.

And the volume it falls back on is not the cluster's total demand. This is the part we had wrong ourselves until we went back and read the query that produces the ranking. Clusters are ordered strictly by the search volume of their single biggest keyword, rather than by the cluster's aggregate demand.

That one detail decides the outcome.

Two cluster profiles side by side. The product category cluster has one bar at 74,000 with a small tail and a total around 106,000, marked as ranked first. The service practice cluster has eight flat mid-volume bars topping out at 18,000 with a larger total around 118,000, marked as ranked below it.

Ranking on a cluster's largest member, not its total. Bar shapes are illustrative; the 74,000 figure is a real keyword from this project's universe.

Ranking a cluster by its largest member rewards clusters shaped like a product category, because product categories have one enormous head term with a tail hanging off it. It punishes clusters shaped like a professional service practice, which are broad and flat: lots of mid-volume phrasings, no single dominant term. A cluster holding "power bi software" at 74,000 searches a month beats a cluster of specialist terms carrying more demand in total but no comparable spike.

Look at what won. "power bi software" at 74,000 searches a month, "power business intelligence" (74,000), "analytics google analytics" (49,500), "googletagmanager" (18,100). Software product names and job titles. Precisely the shape a max-based ranking rewards.

There is a fourth contributor that makes all of this worse. Seeds drift into product categories in about two hops. The service "data analytics" expanded into the business intelligence software category, and that expansion spawned fresh seeds of its own: "Business Intelligence", "Power BI Analytics", "Business Intelligence Platforms". Nobody decided to research BI platforms. The expansion decided, one plausible step at a time.

The result was 0 out of 45 prompts present. And the entities that did show up in those answers were real business intelligence consultancies, which is to say the answers were correct. The engines did their job perfectly. We asked business intelligence questions, got business intelligence answers, and filed a social media agency's absence from the business intelligence market as an AI visibility failure.

A 0% built on off-brand topics is a tautology, not a finding.

The obvious response is to add a gate: only let the topic-choosing step see groups tagged "on brand" or "mostly on brand". On the distribution above, that admits 696 of 789 groups. It strips only the bottom 12% and leaves the failure entirely intact, because it inherits the same classifier that could not discriminate in the first place. It is an easy move to reach for and it feels like progress, which is what makes it a trap: a filter is only ever as good as the signal it filters on, so the work has to go into the signal itself.

A ranking that holds up needs three things. Score brand fit as a continuous number rather than four coarse buckets, by measuring how close a keyword group sits to the brand's own description rather than asking a model to drop it in a labelled box. Then rank on fit and volume together instead of one after the other, so that a group scoring twice as well on fit can outrank one with somewhat more search volume. Something of this shape:

The proposed ranking, written as score of a group equals the cosine similarity between the group centroid and the brand embedding, multiplied by the natural log of one plus the group's total volume. The first half is annotated: how close the group sits to the brand's own description, a real number, so groups stop tying. The second half is annotated: total volume, not the largest single keyword, and the log stops one 74,000-volume term swamping the fit score.

Ranking on brand fit and total volume together, rather than on a tier followed by a group's loudest member.

Two things change here, and both matter. Brand fit becomes a real number from an embedding comparison rather than a bucket assigned by a model, so groups stop tying. And total volume replaces the largest single member, wrapped in a logarithm so that a 74,000-volume head term contributes more than an 18,000 one without swamping the fit score entirely. That combination is what lets a specialist's practice areas compete with a product category on equal terms. The third part is to cap how far a seed may spawn further seeds, so a service cannot quietly become a product category two hops later.

This generalises well beyond us. Every visibility number is a measurement of a chosen prompt set, and any tool choosing that set from search volume inherits the same bias. Which makes it a fair question to put to any vendor in this category, ours included: the first thing to ask isn't how many providers a tool tracks, but how it decided what to ask in the first place, and what it does when your business has no head term. A tool that cannot answer that is reporting on a prompt set nobody chose deliberately.

4. When a keyword row becomes published copy

The second leak is what happens when the same string is asked to be both a measurement and a sentence. Every step in the chain behaves correctly, and the article still comes out wrong.

The case is one of our own, generated by our pipeline against our own keyword universe, which is the only reason we can show it at this level of detail. The focus keyword was "content AI".

That is not a typo. It is a row in our keyword universe, word for word, sitting in a cluster of related rows that are all names for one competing WordPress SEO plugin's paid AI writing module. Although it is the genuine, current name of that product and spelled exactly right, as an English phrase it is backwards compared to the natural "AI content".

It went into the published draft untouched. Seven occurrences across 2,187 words, including in H2 and H3 headings.

Three independent guards existed for exactly this. All three had shipped about a week earlier. All three missed it, for the same reason.

GuardWhat it doesWhat it did here
Suggester promptTells the model to emit a corrected form when the matched keyword is malformed, concatenated, misspelled or outdated"content ai" is none of those, so the model fell back on the stronger instruction sitting in the same paragraph, which says the focus keyword must be one of the keywords in the universe, and echoed the row
Keyword normalizerAn LLM pass whose whole job is repairing bad keywords before useLooked at "content ai", changed the capitalisation to "content AI", marked it valid, and passed it on
Keyword policyEmits usage rules, then counts usage in the outputIts rule explicitly permits rewording to "AI content". Ignored. It then counted 7 uses against a cap of 3, logged that, and did nothing with it

All three shared one blind spot. Every guard enumerated spelling defects: concatenated, misspelled, outdated, mis-cased. "content ai" is spelled perfectly. What is wrong with it is the meaning: the word order is backwards, and it is somebody else's product name. No layer had a concept for either.

Four checks, none of them looking for the actual defect. The string "content ai" travels left to right through four gates and passes every one: concatenated words, misspelled, outdated product name, and wrong casing, where it is changed to "content AI". It arrives at the published headings unchanged in meaning. Below, marked as what was actually wrong and bypassing every guard, sit two unchecked defects: the word order is inverted, since the natural form is "AI content", and it is a third party's product name, spelled correctly and current.

Every guard checked spelling. The defect was meaning.

That is the general trap in building guardrails over a data source you don't control. You write rules against the failures you've already seen. A well-formed string that's wrong for reasons of meaning walks straight through a wall built out of spell checks.

Every layer did exactly what it was told, and four structural details downstream turned that one bad keyword into a bad article.

The brief is the transmission vector. The suggester is asked to express per-section keywords as natural search intent, never as raw query strings. It broke that in all five sections and wrote the literal phrase into each one. That brief is then injected verbatim into the generation prompt, so the raw string reaches the writer five times as concrete, in-context instruction. A concrete example in a prompt beats an abstract rule in the same prompt every single time.

The rubric only rewards presence.

Segmented bar showing an article's 85 available quality points. 50 points go to keyword presence: title 15, intro 15, heading 10, secondary coverage 10. The remaining 35 cover word count, headings, FAQ and metadata. A highlighted callout states that zero points are deducted for repeating the keyword.

The scorecard an article is graded against. Nothing on it costs you anything for repetition, which is how a keyword used seven times against a limit of three still comes out "Strong".

The draft had no secondary keywords at all, which scores as a pass on an empty list. A free 10 out of 10 for having none. With a single keyword, and rules demanding it in the title, the intro and a heading, there's no alternative phrasing to rotate through. The rubric doesn't just fail to punish repetition here. It pushes you toward it.

And none of it bought anything. The presence check was already lenient: it takes an exact phrase or a semantic token window. The article's real title never contained the phrase "content AI" at all, and passed regardless. Rewording it naturally would have scored the same. The pipeline carried a broken string all the way to publication in exchange for nothing at all.

There is a general lesson underneath all four. A generate-evaluate-repair loop only closes if the repair step is actually wired to act on what the evaluation found, and it is easy to build one where the score is computed, stored, and never consulted. That 76 out of 100, "Strong" was not a judgement anybody overruled. It was a number written down after the last point at which anything could still change.

5. Two strings, not one

The obvious move is to fix the keyword in place. Rewrite "content AI" to "AI content" on the way in, move on. That instinct is wrong, for a reason that transfers well beyond keywords.

The keyword is quietly doing two jobs. One is as a measurement handle: the thing tying this article back to the specific gap that triggered it, so visibility scoring can later ask whether writing it changed anything. The other is as copy, words a human reads in a heading. Rewriting in place fixes the second job by silently breaking the first, since the article would no longer be attributable to the demand it was written for.

So we carry both. One field holds the measured keyword, stored exactly as it came out of the database. A second holds the written keyword, the natural phrasing that actually goes into the article. The suggestion card shows only the written one. The scoring layer reads the measured one. A third field carries alternative phrasings, so a writer has something to rotate through instead of repeating one phrase in the title, the intro and a heading.

Concretely, one field doing three jobs becomes three fields doing one each:

Before and after comparison of a suggestion record. Before: a single field, "keyword" set to "content AI", annotated as one field both measured against and written with. After: three fields. "measured_keyword" holds "content ai", read by visibility and coverage measurement and never by the writer. "written_keyword" holds "AI content optimization", read by the brief suggester and the writer. "alternate_phrasings" holds a list, read by the writer and scored by the quality rubric.

The string you measure and the string you publish are not the same string.

Names here are illustrative rather than ours. The first field is what visibility gets measured against, because it is the string that came from the gap. The second is what the writer actually uses. The third is what stops the writer painting itself into a corner. Before the change, one field was doing all three jobs.

Two other decisions generalise.

Deterministic detection rather than another model call. An LLM normalizer was already in the path and it passed the keyword, so the flaw lay in the validation rule rather than in the mechanism running it. Bolting a second model call on to check the first compounds the problem rather than solving it. Deterministic detectors also have a property no LLM rewrite has: you can dry-run them across your entire history and see exactly what they would have caught and what they would have broken, before shipping anything.

Reword third-party product names to the generic capability, rather than dropping them. The demand behind that keyword row is real, since people do search for AI content optimisation tooling. Dropping the row loses the demand. Keeping it verbatim borrows a competitor's brand. Mapping it to the capability keeps one and drops the other.

6. How this squares with outside research

We built all of this from our own telemetry, which means the numbers are ours and the sample is small. What raises confidence is independent research arriving at adjacent conclusions for unrelated reasons.

External findingWhat it corroborates
65% to 85% of ChatGPT prompts matched nothing in a 27-billion-keyword database (17-month clickstream study, over 1B lines of US panel data, through Feb 2026)Why the keyword universe cannot be the denominator for AI demand. Ours is entirely search-advertising sourced, so it is structurally blind to the majority of prompts
Prompts using traditional search language nearly doubled Oct 2025 to Feb 2026, from 18.9% to 34.9%, same studyWhy we are not deprecating the keyword universe. The proxy is improving rather than collapsing. What needs measuring is the error term
Word overlap between prompt and issued query: Perplexity 88%, Copilot 50%, ChatGPT 13% (10,000 prompts, 14 days, spring 2026)Why proxy quality has to be assessed per provider instead of globally, and why one blended visibility number hides three different error rates
Query uniqueness across runs: ChatGPT 91%, Copilot 47%, Perplexity 14%, same studyWhy run-to-run volatility on ChatGPT is a property of the engine rather than noise in our measurement, and why smoothing it away would hide real signal
ChatGPT enabled web search on 34.5% of queries as of Feb 2026, down from 46%A ceiling on this whole category. Retrieval-shaped optimisation only addresses the minority of prompts where retrieval actually happens
Adding statistics, citations and quotations lifted citation visibility by 30% to 40%, while keyword density showed minimal effect (GEO: Generative Engine Optimization, peer-reviewed, KDD 2024)Why a rubric spending 50 of 85 points on keyword presence was optimising the wrong property entirely

The number to remember: adding statistics and source citations lifts generative visibility by 30% to 40%. Keyword density lifts it by roughly nothing.

Which reframes what a content scorecard is for. Any rubric weighted toward keyword presence is measuring the wrong property, however carefully it is calibrated, because it is counting a signal the engines largely ignore. Points spent on evidence density, on whether a passage carries a statistic, a source or a quotation, are points spent on something that demonstrably moves citation probability.

7. Where this leaves us

So, how far does it hold? Far enough to keep, and not far enough to trust blindly. It is a proxy with an error term that is now measurable rather than assumed, and that error term varies by engine, by the shape of your business, and by what you are asking the data to do. The data itself was never the problem; the damage came from forcing a single proxy to serve three incompatible jobs.

As the demand denominator, it holds up. With the per-engine caveats above, and improving over time.

As the topic selector, it needs a second signal beside it. Volume alone systematically picks product categories over service practices, and a coarse alignment tier is not strong enough to counteract that on its own.

As literal copy, it was never meant to serve. A database row is a measurement artifact. Nobody intended it to be read by a human, and the moment one turns up in a heading, something upstream has mistaken an index for a sentence.

If you are building on keyword data, four properties are worth designing for from the start, because each one is expensive to retrofit:

  • Brand fit as a continuous score, never coarse buckets. Buckets that most of the input lands in cannot order anything, and a sort on them silently becomes a sort on whatever came next.
  • Ranking on fit and volume jointly, never one after the other, and on a group's total rather than its loudest member.
  • A hop limit on seed expansion, so a service cannot become a product category two steps later without anyone choosing that.
  • Separate fields for the string you measure and the string you publish, so neither job can quietly corrupt the other.

The bigger piece, mining a prompt space directly instead of deriving one from keyword data, is a sibling of the keyword universe rather than a layer on top of it. Deriving an intent space from a keyword database inherits the aperture problem from section 2 and presents it more confidently, which is worse than not having one.

Further Reading

Share this article

Is your brand visible in AI search?

Discover how ChatGPT and Perplexity talk about your brand. Get weekly insights and recommendations to improve your AI presence.

Related Articles