We ran 15,000 controlled AI captures. APIs and consumer surfaces agreed on 22% of citation sources.

500 prompts, 5 verticals, 14,354 paired captures across three AI providers. Overall citation-domain overlap between API and consumer surface: 22%. Google API path: 0%.

Dan Johnson

Co-Founder · July 29, 2026 · 7 min read

We ran 15,000 controlled AI captures. APIs and consumer surfaces agreed on 22% of citation sources.

We ran a controlled study on how three of the major AI providers answer the same question through two paths: their developer API, and the consumer chat surface a customer sees when they open chatgpt.com, gemini.google.com, or perplexity.ai.

500 pre-registered prompts across 5 verticals. 14,354 successful captures. 71,381 brand mentions extracted. 92,724 citations across 11,217 unique domains.

Overall Jaccard similarity[1] between the two paths’ citation domains was 0.225. About 22% of unique domains appeared on both paths, 78% on one but not the other.

0.225
Jaccard similarity between API and consumer-surface citation domains
n = 14,354 paired captures across 3 providers, 500 prompts, 5 verticals

One caveat before the findings, because it changes how you should read the numbers below. We tested OpenAI, Google, and Perplexity. Anthropic (Claude) was excluded because its consumer chat surface does not expose an accessible logged-out interface we could pair with API calls under equivalent conditions. Microsoft Copilot is testable and belongs in a follow-up run; we scoped it out of this study to keep the paired design tight.

What we actually measured

For each of the 500 prompts, we ran two paired captures per provider:

  • API path: a direct call to the provider’s public developer API with the grounding tool enabled where the provider supports one.
  • Surface path: the same prompt submitted through the logged-out consumer chat interface at chatgpt.com, gemini.google.com, or perplexity.ai, captured through an automated browser.

Both responses were passed through the same brand-mention extractor and the same citation parser so anything that differs between them is a property of the paths, not the measurement.

Models tested

Model versions on each path, per provider
ProviderAPI modelConsumer surface model
OpenAIgpt-4o-2024-08-06chatgpt.com default (logged-out)
Googlegemini-2.5-progemini.google.com default (logged-out)
Perplexitysonar-properplexity.ai default (logged-out)

A note on the surface side: providers do not always name the exact model that answers a logged-out consumer session, and they can silently change it. The paired design still holds because both paths are captured on the same day for the same prompt, but the comparison is API model vs whatever the provider chose to serve consumer traffic that day, not API model vs the same model on the surface.

One more limit worth naming: gpt-4o-2024-08-06 was already old on the OpenAI API side by the time we ran this. A newer OpenAI model would probably close some of the OpenAI numbers. The structural findings, though, look like properties of the paths, not the model versions. Google’s redirect layer at 99.8%, Perplexity’s 62% verbosity gap, BBB.org’s dominance on the surface path... those do not move when you swap out the API model. Our expectation is the direction holds and the magnitudes shift. That is a hypothesis, not a tested result.

Per-provider results

The overall 0.225 number covers a wide range across the three providers. Reported per provider below.

Citation-domain overlap between API and consumer surface, by provider
ProviderAPI unique domainsSurface unique domainsSharedJaccard
Google (Gemini)22,29300.000
OpenAI (ChatGPT)3,0642,8568880.176
Perplexity3,8233,8739450.140

Google

99.8% of Google API citations (19,035 of 19,070) resolved to a single redirect host at vertexaisearch.cloud.google.com. Google’s Vertex AI grounding docs[2] and the Gemini API grounding docs[3] describe this as expected: the grounding tool returns Google-hosted URIs that redirect to the underlying source, rather than returning the source domain directly. Structurally, the API and the consumer surface therefore share zero unique domains at the URL level.

If you want to know which websites Gemini actually cited when it answered a prompt, the API response alone will not tell you. You have to follow each redirect and resolve the destination host, which is what the consumer surface does before it renders a link.

OpenAI

OpenAI’s web search tool[4] returned real source URLs (tagged with the ?utm_source=openai referral parameter). Overlap between the API’s citation set and chatgpt.com’s citation set was 0.176, roughly 18% shared unique domains.

Perplexity

Perplexity’s chat completions API[5] returned real source URLs on both paths. Overlap was 0.140, about 14% shared unique domains.

Consumer surfaces returned fewer brand names than the APIs

On the same prompts, the consumer surface returned fewer distinct brand mentions per response than the API. Same extractor on both paths, so the delta is a property of the response, not the measurement.

Average brand mentions per capture, API vs consumer surface
ProviderAPI mentions/captureSurface mentions/captureGap
Google5.755.02+14%
OpenAI5.154.44+16%
Perplexity5.973.69+62%

Perplexity’s API returns 62% more brand names per response, on average, than perplexity.ai returns to a logged-out user asking the same question. OpenAI and Google both show smaller but present gaps. The consumer surface is a filtered version of what the underlying model produces, and the filter differs by provider.

The pattern held across every vertical we tested

The 500 prompts covered 5 verticals. Jaccard was computed within each vertical, across all three providers.

API-vs-surface Jaccard by vertical
VerticalJaccard
E-commerce0.177
Local home / auto services0.209
B2B SaaS0.222
HVAC0.244
Legal0.250

Range: 0.18 to 0.25. Every vertical showed overlap under 25%. Whatever is driving the divergence is not a quirk of one industry’s aggregator dynamics. It shows up in services, retail, and B2B alike.

Which domains showed up on each path

Top 5 domains cited through the APIs, across all providers and all verticals:

  1. vertexaisearch.cloud.google.com (Google’s grounding redirect): 19,035 citations
  2. reddit.com: 1,902
  3. facebook.com: 1,105
  4. youtube.com: 1,010
  5. yelp.com: 806

Top 5 domains cited through the consumer surfaces:

  1. reddit.com: 1,744
  2. bbb.org[6] (Better Business Bureau[6]): 852
  3. youtube.com: 672
  4. justia.com: 436
  5. attorneys.superlawyers.com: 429

Reddit was the one high-frequency domain shared across both paths, cited 1,902 times through the API and 1,744 times through the surface. Reddit publicly signed a data licensing agreement with Google in early 2024[7] that made Reddit content available to Google’s AI training pipeline, and other providers have similar arrangements or open scraping of the site. The high Reddit weight on both paths is consistent with those relationships.

BBB.org appeared in 852 responses across 88 distinct prompts on the consumer surface path. It did not appear in the API top 15 at all. The same pattern held for expertise.com, avvo.com, legalclarity.org, techradar.com, and goodhousekeeping.com, all frequent on the surface path, absent from the API top 15. Consumer surfaces cited trust-signal domains and specialist directories that the APIs did not lean on at the same frequency.

Why this matters if you’re trying to measure AI visibility

Businesses making decisions about AI-driven search visibility need to know which measurement path a given tool uses, because the picture differs.

A business that optimizes its Yelp listing will show up in API-based measurements. A business that builds its BBB profile will show up in consumer-surface measurements. These are different investments, and neither picture is wrong on its own terms. They answer different questions.

The API path is faster to measure and cheaper to run. It is also structurally disconnected from what a customer sees when they ask ChatGPT.com or gemini.google.com for a recommendation. The consumer-surface path is slower and more expensive to capture, but reflects what a customer actually experiences.

If you are evaluating an AI visibility tool, worth asking: does this measure my customer’s experience on chatgpt.com and gemini.google.com and perplexity.ai, or does it measure what an API returns for the same prompt?

Why we built this

At GeoReputation we measure consumer surfaces, not APIs. Every score we ship to a business traces back to a specific prompt, a specific model, and the exact response a customer would see on chatgpt.com, gemini.google.com, or perplexity.ai. Customers get access to the underlying prompts, responses, citations, and recommendations. The data is the product.

If you want to see what your business looks like on the paths customers actually use, run a free scan.

References

  1. [1]
    Jaccard indexWikipedia
    https://en.wikipedia.org/wiki/Jaccard_index
    Accessed Jul 15, 2026
  2. [2]
    Grounding overviewGoogle Cloud Vertex AI
    https://cloud.google.com/vertex-ai/generative-ai/docs/grounding/overview
    Accessed Jul 15, 2026
  3. [3]
    Grounding with Google SearchGoogle AI for Developers
    https://ai.google.dev/gemini-api/docs/grounding
    Accessed Jul 15, 2026
  4. [4]
    Web search toolOpenAI Platform
    https://platform.openai.com/docs/guides/tools-web-search
    Accessed Jul 15, 2026
  5. [5]
    Chat Completions API referencePerplexity
    https://docs.perplexity.ai/api-reference/chat-completions-post
    Accessed Jul 15, 2026
  6. [6]
    Better Business BureauBBB
    https://www.bbb.org/
    Accessed Jul 15, 2026
  7. [7]
    Reddit signs AI content deal with GoogleReuters
    https://www.reuters.com/technology/reddit-signs-ai-content-deal-with-google-sources-say-2024-02-22/
    Accessed Jul 15, 2026

About Dan Johnson

Dan Johnson is the co-founder of GeoReputation, where he handles the engineering. Posts on this blog are usually grounded in data pulled live from the platform.