Two things about generative engine optimization are true at the same time. The measurement underneath it is unreliable enough that most of the numbers being reported to boards this year would not survive a statistics review. And the brands waiting for better measurement before they start will be a year or two behind the ones that don't.
That tension is worth working through carefully, because a large amount of money is about to move on the strength of numbers almost nobody can audit.
What Are Companies Actually Buying?
Generative engine optimization (GEO), also called answer engine optimization (AEO), is the practice of trying to influence whether and how a brand appears when a customer asks an AI assistant a question. The adjacent practice, AI visibility monitoring, tries to measure how often that happens.
The framing matters. GEO is not a channel. It is a layer of influence, and the click from a citation inside ChatGPT is the smallest and least interesting part of it. The relevant question is not how many sales an agent completed. It is how many purchases an agent shaped -- the same distinction that separates agentic commerce from agent-executed commerce, and social commerce from in-platform checkout.
That distinction is also the first clue about why this is so hard to measure.
Why Is The Measurement So Unreliable?
Six problems, and they all compound.
- The engine regenerates the answer every time. Ask the same assistant the same question twice and you often get different brands and different sources. A study published on arXiv in June measured how much the cited source set overlapped across repeated runs of identical queries: roughly 30% on Gemini, 33% to 40% on SearchGPT, and about 50% on Perplexity. Instability is worst exactly where it matters most. Purchase-intent prompts were the least consistent query type tested in an analysis of 14,000 API calls, with only about four of every 10 brands appearing across two runs of the same prompt, according to Conductor.
- Most tools measure a different product than the customer uses. Tracking platforms typically call a developer API. Customers open a consumer app carrying memory, personalization, location, model routing and its own logic about when to search the web. The gap shows up even inside a single product: only 25% of cited sources overlapped between ChatGPT's fast Instant mode and its slower Thinking mode, and the two sometimes recommended different brands for the same prompt, in an analysis run by Semrush with Kevin Indig. Who the customer says they are matters too. One index covering 126 software categories found a brand moving from winning 34% of answers to 90% depending on whether the buyer described themselves as a freelancer or an enterprise.
- The tracked prompt is not the prompt that runs.Assistants decompose a question into sub-queries and run them in parallel, a process Google calls query fan-out. Google has said its Deep Search feature can issue hundreds of sub-queries for a complex question. A dashboard tracking one phrase is grading a test the engine never took.
- Citations are not sources. Many systems generate an answer and attach supporting links afterward. Researchers call the failure mode post-rationalization: the citation is plausible but was not what produced the claim. The gap is measurable. Across 278 answers covering 95 tracked brands, ChatGPT named a brand in its prose 552 times but cited that brand's own domain only 375 times, and the two overlapped just 185 times, according to an analysis by Position Digital. Being cited and being recommended are different events. A citation report measures one of them.
- The noise floor exceeds the effect. The same arXiv research put bootstrap confidence intervals on citation share and found typical intervals spanning three to seven percentage points. Its blunt conclusion: an apparent move from 8% to 11% cannot be distinguished from random variation. Reaching a defensible estimate required 40 to 50 queries per topic on one platform and 150 or more on another. Most brand dashboards run a handful of prompts weekly and report to one decimal place.
- Visibility is not an outcome. A mention is not a sale. The referral traffic assistants send directly is a small fraction of site visits, and the real influence is almost certainly much larger than what shows up in a referral log -- which is precisely the problem. Neither number tells a chief financial officer whether the work paid for itself.
Why Doesn't Getting It Right Stay Right?
This is the part executives find hardest, and the one that should reset expectations.
Search is close to deterministic. The same query returns roughly the same ranked list to everyone, and a position earned through good work is a durable asset -- often six months or more of advantage until the algorithm updates. That durability is what made SEO investable.
Generative systems are probabilistic. The answer is composed fresh on every request. There is no position to hold and no state to defend. Give a brand a magic wand today that revealed its exact standing and the precise levers to improve it, and the picture would drift by next week without anyone touching anything -- because the model changed, the retrieval layer changed, a competitor published, or the customer's own history changed.
GEO is closer to running a continuous experiment than to holding a ranking. Teams that budget for it as a project will be disappointed. Teams that budget for it as an ongoing measurement function will not.
What About Vertical Versus Horizontal Agents?
Most GEO conversations quietly assume one kind of agent. There are two.
Horizontal agents -- ChatGPT, Gemini, Claude, Perplexity, Copilot -- are trained on the open web and sit closest to discovery. Vertical agents are retailer-owned and catalog-bound -- Amazon's Alexa for Shopping (formerly Rufus), Walmart's Sparky, and the assistants appearing inside grocery, pharmacy and specialty apps -- and they sit closest to the transaction. Alexa for Shopping is trained on Amazon's catalogue and its Prime customer behavior. ChatGPT is not. Visibility in one says very little about visibility in the other.
Almost no tool attempts both. The monitoring category is built around horizontal assistants because those can be prompted from the outside. Retailer agents cannot be queried at scale and return almost no data to brands. So the surface nearest the purchase is the one few brands try to measure, and the surface furthest from it generates all the dashboards.
What Kind Of Data Is The Measurement Built On?
Two fundamentally different methods, and the distinction rarely appears on a sales slide.
Synthetic prompt monitoring generates a prompt set, runs it against the engines on a schedule, and reports what came back. Repeatable and auditable. Its weakness is that someone had to guess what customers ask.
Real prompt monitoring uses consented consumer panels and observes conversations that actually occurred. One such panel from Measure Protocol, spanning apps, browsers, search and AI assistants, reported in 2025 that more than one in five ChatGPT conversations showed commercial intent. Several GEO vendors now build panels of this kind into their products. Panel data is the closest thing this field has to observed rather than stated preference, and observed preference is almost always worth more.
Neither is sufficient alone. Synthetic testing shows how engines respond to a controlled stimulus. Panels show what people actually ask. The useful programs run both.
One caution applies to every vendor here: very few disclose how the data is collected. Which surface, which model build, logged in or out, which region, how many repetitions per prompt, how variance is handled. A buyer who cannot get those answers in writing is not buying a measurement. They are buying a number.
What Happens As The Budgets Grow?
The fraud grows with them, and the category already contains tactics with no evidence behind them. Ninety-seven percent of published llms.txt files received zero requests in May, and among files that did get traffic the largest single category of requester was SEO audit tools checking whether the file existed, according to an Ahrefs analysis of server logs from 137,000 domains. Google has said the file is not needed. It is still being sold.
Expect worse as spending rises: fabricated benchmarks, content seeded to game retrieval, and proprietary scores only the vendor's own methodology can improve. Expect the pace of change to accelerate rather than settle.
So Why Build The Practice Now?
Because in 2026, the deliverable is not the number. It is the organization.
- Executives are already running prompts. A chief executive who typed the category into ChatGPT last week saw something and drew a conclusion. If a competitor came up first, that conclusion is already forming.
- Every vendor will send a visibility report to someone senior, and they will all disagree, because they are measuring different surfaces with different methods. Without an internal number, a company is refereeing a fight among people who profit from the fight.
- A consistent number with a published methodology beats an accurate number nobody shares. Teams need something to argue from rather than about.
- Change management is the binding constraint, as with every commerce shift. Getting merchandising, brand, content and IT into one conversation takes roughly 18 months. Start while the stakes are low.
- Most of the underlying work carries no regret. Accurate, structured, machine-readable product content improves organic rank, conversion, retail media efficiency and vertical agent performance. That case does not depend on mass AI shopping adoption. It depends on the rest of the business.
There is real evidence the influence is worth chasing. Users who saw a brand recommended by ChatGPT were 2.5 times more likely to visit that brand's site within seven days, with 56% of that traffic arriving through branded search rather than an AI referral, according to Similarweb clickstream research. Similarweb labels the finding correlation rather than causation, and the study covers United States desktop users across three verticals. It looks a great deal like measuring a billboard: lift, not clicks.
One caveat belongs on all of it. Nearly every study cited here, on both sides, was produced by a company that sells software in this category. That does not make the findings wrong, but it should shape how confidently they are read. The most rigorous work is academic, and its central conclusion is that everyone else's numbers need error bars.
What Should A Team Do This Quarter?
Write the methodology down before writing the number down, and publish both together. Report ranges rather than point estimates. Build the prompt set from observed customer language -- site search logs, service transcripts, review text, panel data -- not from what the brand team wishes people asked. Track vertical and horizontal agents separately and expect them to disagree. Keep visibility and business outcomes in separate columns, then measure incrementality the only way that has ever worked, through holdouts and matched-market tests. Ask every vendor, in writing, how the data is collected. And do the unglamorous work first: clean product data, content that answers real customer questions, and a firewall that is not quietly blocking the crawlers.
The measurement will be wrong, in ways that can be enumerated. That is the useful kind of wrong. The capability being built -- a team that can reason about probabilistic, agent-mediated customer behavior -- will matter for considerably more than chatbot mentions.
Just don’t put three decimal places on it.