We ran 15,000 controlled AI captures. APIs and consumer surfaces agreed on 22% of citation sources.
500 prompts, 5 verticals, 14,354 paired captures across three AI providers. Overall citation-domain overlap between API and consumer surface: 22%. Google API path: 0%.
Dan Johnson
Co-Founder · July 29, 2026 · 7 min read

We ran a controlled study on how three of the major AI providers answer the same question through two paths: their developer API, and the consumer chat surface a customer sees when they open chatgpt.com, gemini.google.com, or perplexity.ai.
500 pre-registered prompts across 5 verticals. 14,354 successful captures. 71,381 brand mentions extracted. 92,724 citations across 11,217 unique domains.
Overall Jaccard similarity[1] between the two paths’ citation domains was 0.225. About 22% of unique domains appeared on both paths, 78% on one but not the other.
One caveat before the findings, because it changes how you should read the numbers below. We tested OpenAI, Google, and Perplexity. Anthropic (Claude) was excluded because its consumer chat surface does not expose an accessible logged-out interface we could pair with API calls under equivalent conditions. Microsoft Copilot is testable and belongs in a follow-up run; we scoped it out of this study to keep the paired design tight.
What we actually measured
For each of the 500 prompts, we ran two paired captures per provider:
- API path: a direct call to the provider’s public developer API with the grounding tool enabled where the provider supports one.
- Surface path: the same prompt submitted through the logged-out consumer chat interface at chatgpt.com, gemini.google.com, or perplexity.ai, captured through an automated browser.
Both responses were passed through the same brand-mention extractor and the same citation parser so anything that differs between them is a property of the paths, not the measurement.
Models tested
| Provider | API model | Consumer surface model |
|---|---|---|
| OpenAI | gpt-4o-2024-08-06 | chatgpt.com default (logged-out) |
| gemini-2.5-pro | gemini.google.com default (logged-out) | |
| Perplexity | sonar-pro | perplexity.ai default (logged-out) |
A note on the surface side: providers do not always name the exact model that answers a logged-out consumer session, and they can silently change it. The paired design still holds because both paths are captured on the same day for the same prompt, but the comparison is API model vs whatever the provider chose to serve consumer traffic that day, not API model vs the same model on the surface.
One more limit worth naming: gpt-4o-2024-08-06 was already old on the OpenAI API side by the time we ran this. A newer OpenAI model would probably close some of the OpenAI numbers. The structural findings, though, look like properties of the paths, not the model versions. Google’s redirect layer at 99.8%, Perplexity’s 62% verbosity gap, BBB.org’s dominance on the surface path... those do not move when you swap out the API model. Our expectation is the direction holds and the magnitudes shift. That is a hypothesis, not a tested result.
Per-provider results
The overall 0.225 number covers a wide range across the three providers. Reported per provider below.
| Provider | API unique domains | Surface unique domains | Shared | Jaccard |
|---|---|---|---|---|
| Google (Gemini) | 2 | 2,293 | 0 | 0.000 |
| OpenAI (ChatGPT) | 3,064 | 2,856 | 888 | 0.176 |
| Perplexity | 3,823 | 3,873 | 945 | 0.140 |
99.8% of Google API citations (19,035 of 19,070) resolved to a single redirect host at vertexaisearch.cloud.google.com. Google’s Vertex AI grounding docs[2] and the Gemini API grounding docs[3] describe this as expected: the grounding tool returns Google-hosted URIs that redirect to the underlying source, rather than returning the source domain directly. Structurally, the API and the consumer surface therefore share zero unique domains at the URL level.
If you want to know which websites Gemini actually cited when it answered a prompt, the API response alone will not tell you. You have to follow each redirect and resolve the destination host, which is what the consumer surface does before it renders a link.
OpenAI
OpenAI’s web search tool[4] returned real source URLs (tagged with the ?utm_source=openai referral parameter). Overlap between the API’s citation set and chatgpt.com’s citation set was 0.176, roughly 18% shared unique domains.
Perplexity
Perplexity’s chat completions API[5] returned real source URLs on both paths. Overlap was 0.140, about 14% shared unique domains.
Consumer surfaces returned fewer brand names than the APIs
On the same prompts, the consumer surface returned fewer distinct brand mentions per response than the API. Same extractor on both paths, so the delta is a property of the response, not the measurement.
| Provider | API mentions/capture | Surface mentions/capture | Gap |
|---|---|---|---|
| 5.75 | 5.02 | +14% | |
| OpenAI | 5.15 | 4.44 | +16% |
| Perplexity | 5.97 | 3.69 | +62% |
Perplexity’s API returns 62% more brand names per response, on average, than perplexity.ai returns to a logged-out user asking the same question. OpenAI and Google both show smaller but present gaps. The consumer surface is a filtered version of what the underlying model produces, and the filter differs by provider.
The pattern held across every vertical we tested
The 500 prompts covered 5 verticals. Jaccard was computed within each vertical, across all three providers.
| Vertical | Jaccard |
|---|---|
| E-commerce | 0.177 |
| Local home / auto services | 0.209 |
| B2B SaaS | 0.222 |
| HVAC | 0.244 |
| Legal | 0.250 |
Range: 0.18 to 0.25. Every vertical showed overlap under 25%. Whatever is driving the divergence is not a quirk of one industry’s aggregator dynamics. It shows up in services, retail, and B2B alike.
Which domains showed up on each path
Top 5 domains cited through the APIs, across all providers and all verticals:
vertexaisearch.cloud.google.com(Google’s grounding redirect): 19,035 citations- reddit.com: 1,902
- facebook.com: 1,105
- youtube.com: 1,010
- yelp.com: 806
Top 5 domains cited through the consumer surfaces:
- reddit.com: 1,744
- bbb.org[6] (Better Business Bureau[6]): 852
- youtube.com: 672
- justia.com: 436
- attorneys.superlawyers.com: 429
Reddit was the one high-frequency domain shared across both paths, cited 1,902 times through the API and 1,744 times through the surface. Reddit publicly signed a data licensing agreement with Google in early 2024[7] that made Reddit content available to Google’s AI training pipeline, and other providers have similar arrangements or open scraping of the site. The high Reddit weight on both paths is consistent with those relationships.
BBB.org appeared in 852 responses across 88 distinct prompts on the consumer surface path. It did not appear in the API top 15 at all. The same pattern held for expertise.com, avvo.com, legalclarity.org, techradar.com, and goodhousekeeping.com, all frequent on the surface path, absent from the API top 15. Consumer surfaces cited trust-signal domains and specialist directories that the APIs did not lean on at the same frequency.
Why this matters if you’re trying to measure AI visibility
Businesses making decisions about AI-driven search visibility need to know which measurement path a given tool uses, because the picture differs.
A business that optimizes its Yelp listing will show up in API-based measurements. A business that builds its BBB profile will show up in consumer-surface measurements. These are different investments, and neither picture is wrong on its own terms. They answer different questions.
The API path is faster to measure and cheaper to run. It is also structurally disconnected from what a customer sees when they ask ChatGPT.com or gemini.google.com for a recommendation. The consumer-surface path is slower and more expensive to capture, but reflects what a customer actually experiences.
If you are evaluating an AI visibility tool, worth asking: does this measure my customer’s experience on chatgpt.com and gemini.google.com and perplexity.ai, or does it measure what an API returns for the same prompt?
Why we built this
At GeoReputation we measure consumer surfaces, not APIs. Every score we ship to a business traces back to a specific prompt, a specific model, and the exact response a customer would see on chatgpt.com, gemini.google.com, or perplexity.ai. Customers get access to the underlying prompts, responses, citations, and recommendations. The data is the product.
If you want to see what your business looks like on the paths customers actually use, run a free scan.
References
- [1]
- [2]Grounding overview — Google Cloud Vertex AIhttps://cloud.google.com/vertex-ai/generative-ai/docs/grounding/overviewAccessed Jul 15, 2026
- [3]Grounding with Google Search — Google AI for Developershttps://ai.google.dev/gemini-api/docs/groundingAccessed Jul 15, 2026
- [4]Web search tool — OpenAI Platformhttps://platform.openai.com/docs/guides/tools-web-searchAccessed Jul 15, 2026
- [5]Chat Completions API reference — Perplexityhttps://docs.perplexity.ai/api-reference/chat-completions-postAccessed Jul 15, 2026
- [6]
- [7]Reddit signs AI content deal with Google — Reutershttps://www.reuters.com/technology/reddit-signs-ai-content-deal-with-google-sources-say-2024-02-22/Accessed Jul 15, 2026
About Dan Johnson
Dan Johnson is the co-founder of GeoReputation, where he handles the engineering. Posts on this blog are usually grounded in data pulled live from the platform.