In January 2026 SparkToro published the most careful public test we know of. Six hundred volunteers ran recommendation prompts in their own AI tools — 2,961 runs in total. The same list of brands appeared twice in fewer than one run in a hundred. Identical ordering appeared in fewer than one in a thousand. Narrower categories were somewhat more stable. The authors' advice was blunt: measure visibility as a percentage across dozens or hundreds of runs, and be careful with any tool that reports a rank. 1
That finding is not a flaw in one engine. It follows from how answers are made: retrieval draws a slightly different set of pages each time, and generation samples from a distribution. A screenshot of ChatGPT naming your company is real. It is also unrepeatable.
What private aviation has been given instead
Three artifacts currently stand in for an industry benchmark.
The first is EpicEdits' Private Aviation AI Visibility Index: 14 buyer prompts run across ChatGPT, Perplexity and Google AI Mode on 31 May 2026, scored as "named in n of 14", with VistaJet at 11, NetJets 8, Flexjet 7, Wheels Up 6 and PrivateFly 5. The authors describe it as a snapshot and note that answers vary by location, personalization and time. 2 It is honest about what it is; what it is, is one day.
The second is Everything-PR's "Private Aviation Citation Share Index 2026", which ranks NetJets 93, Flexjet 84, VistaJet 82 and so on out of 100. Its methodology assigns 20 of those points to "estimated AI engine retrieval signal" and states that there were no logged query runs. 3 It is a model of what an engine might say, not a record of what it said.
The third is Resocial's 25-brand digital-maturity ranking, in which AI citation carries a 20% weighting, produced by a single agency. 4
None of the three covers FBOs, MROs, aircraft management, sales or vendors. None publishes its prompt list in full. None reports how much the result would change if it were run again tomorrow. Semrush's large-scale index, by contrast, averages 126 million prompts over months and reports per-platform agreement 5 — but it is general-purpose and does not see this industry's questions.
What good measurement looks like
A benchmark that a skeptic can trust needs five properties.
- A fixed, public prompt panel, written in buyers' words, balanced across segments and buying stages, and frozen for a defined period so that month-to-month changes mean something.
- Repeated runs — at least three per prompt per API-measured engine per period, in fresh sessions, so that "named" can be expressed as a share rather than a coin flip.
- Logged answers with engine, measurement surface, region and timestamp, and human extraction of who was named, with a second reviewer where runs disagree.
- Published variance. If share of voice for a prompt cluster is 33%, the reader should see the interval and the per-prompt spread.
- Published data, to the extent provider terms permit, so that anyone — a journalist, a competitor, a client's CFO — can re-run the analysis.
These are the rules we adopted for the Business Aviation AI Visibility Index: 120 prompts across seven segments; four engines measured through official, paid APIs with web search — OpenAI, Gemini, Perplexity and Claude — at five repeats each, monthly (2,400 logged runs); every table labeled "API-measured, not the consumer products"; Google AI Overviews, Copilot and the ChatGPT app never ranked, because they cannot be measured without scraping, and instead sampled by hand each quarter (20 prompts, two runs each) and reported only as an agreement rate with the API engines; edition 1 in January 2027 only after three complete monthly runs; and the derived run log released with every edition, verbatim text where an engine's terms allow. Until then we publish the method and the panel, and no rankings. One more honesty point: API answers are not what a buyer sees in an app — different system prompts, model routing, no personalization — which is a limitation we publish and also the reason anyone with API keys can re-run our panel. We think the temptation to publish a first-run leaderboard is exactly the failure mode SparkToro documented.
What this means for a buyer of AEO services
Ask any provider three questions. How many runs is that number based on? Can I see the log? What was the result the second time? A provider who answers with a screenshot is selling weather as climate. A before/after case study that does not use the same panel and the same number of runs before and after is not evidence of anything; it is two different measurements. 1
How we know
Sources are SparkToro (January 2026), EpicEdits (May 2026), Everything-PR (2026), Resocial (2026) and Semrush (June 2026). Our own Index has logged zero runs as of publication; this article describes the method we have committed to, not results.