🔬 Primary Lab Verification - Hermes Agent Lab

AI Search Grounding: How Google, OpenAI & Anthropic Cite Entities

Published 2026-06-16 - Hermes Agent Lab, hermes-agent.reviews

🔍 Why We Ran This Analysis

AI search platforms don't all cite the same way. Dan Petrovic ran the identical query ("best ai seo agency 2026") through Google (Gemini), OpenAI (gpt-5.5), and Anthropic (Claude) on the same day and reverse-engineered their grounding funnels. The differences are stark: Google cites every page it receives (1:1 ratio), OpenAI cites only 2 of 39 pages it reads (20:1 narrow-cite split), and Anthropic does a two-pass deep read on 14 pages and cites 9 (3:2 ratio).

These platform-specific grounding behaviors create fundamentally different optimization requirements. An entity page that Google cites successfully may be invisible to OpenAI — and vice versa. This analysis breaks down each platform's grounding mechanics and provides actionable entity optimization strategies for each.

📈 Platform Grounding Comparison

AI Search Grounding: Platform-by-Platform Breakdown (Dan Petrovic, June 13, 2026)
DimensionGoogle (Gemini)OpenAI (gpt-5.5)Anthropic (Claude)
Pages Received → Cited7 → 7 (1:1)39 → 2 (20:1)14 → 9 (3:2)
Shows Uncited Pages?NoYes (37 readable)Yes (5 unselected)
Snippet FormFull page contentSliding window (100-200 words)Encrypted blob
URLsRedirect-wrappedDirect (+utm tag)Direct
Input Tokens205 (lightest)10,85383,708 (heaviest)
Speed53s14s (fastest)138s (slowest)

🌎 Google (Gemini): The Full-Page Grounder

1:1 Cite Ratio — Every Page Must Stand Alone

Google's grounding operates on complete page content with a 1:1 received-to-cited ratio. Every page Google retrieves is cited somewhere in the output. This means entity completeness and factual density directly determine whether you appear at all — there is no "read but not cited" category.

Google sends only 205 input tokens (the lightest payload), meaning it relies on pre-indexed page content rather than re-reading full pages at query time. Structured data, entity markup, and factual density in the indexed corpus are what matter.

🎯 Optimization Strategy for Google Grounding

  • Every entity attribute must be a standalone, citable factual claim. Google cites full pages — if your entity page has weak factual density, it won't be retrieved in the first place.
  • Structured data embedded in the page body (not just JSON-LD) may influence grounding. Google uses full page content; entity markup visible in the rendered DOM could matter more than hidden schema.
  • Entity IDs and Wikidata sameAs should be present and verified. Google's Knowledge Graph integration means entity disambiguation is critical for correct grounding.
  • Date freshness signals (dateModified, "Last updated" timestamps) matter because Google's 205-token payload includes freshness metadata.

🤖 OpenAI (gpt-5.5): The Narrow-Cite Splitter

20:1 Cite Ratio — Being Read and Being Cited Are Different Wins

OpenAI receives 39 pages but cites only 2. The other 37 are readable to the model but never surface as citations. This creates a two-tier visibility model: "received" (your page is in the retrieval window) and "cited" (your page is actually referenced in the output).

OpenAI uses a sliding window of 100-200 words from each page — not the full content. The model sees a narrow snippet. If your strongest entity definition, differentiator, and key stat don't appear in the first 200 words, they never reach the model.

OpenAI sends 10,853 input tokens and is the fastest platform at 14s, suggesting aggressive content windowing and minimal re-reading.

🎯 Optimization Strategy for OpenAI Grounding

  • Front-load the entity's strongest definition, key differentiator, and one hard stat into the first 200 words. This is the OpenAI retrieval window — everything beyond it may never reach the model.
  • Use a "Quick Facts" block at the top of every entity page: one-sentence definition, key differentiator, and a hard stat. This maximizes citation probability within OpenAI's 100-200 word sliding window.
  • Entity "received vs cited" audit: test your top 10 entity pages across all three platforms. A page OpenAI reads but never cites needs structural changes different from one Google won't touch.
  • URL structure matters: OpenAI appends utm tags to cited URLs. Clean, canonical URLs without query parameters may be preferred for citation.

🧠 Anthropic (Claude): The Deep Reader

3:2 Cite Ratio — Two-Pass Deep Read Rewards Layered Evidence

Anthropic receives 14 pages and cites 9 (a 3:2 ratio). It uses the heaviest input payload at 83,708 tokens and takes the longest at 138 seconds. This suggests a two-pass architecture: first pass identifies candidate pages, second pass does a deep read of the full content.

Anthropic's snippets are delivered as encrypted blobs (not plain text), and URLs are direct (no redirect wrapping). The high cite ratio (64%) suggests that pages that survive the first-pass filter are likely to be cited if they contain layered, substantive evidence.

🎯 Optimization Strategy for Anthropic Grounding

  • Build layered entity evidence: definition → features → comparisons → use cases → citations. Anthropic's two-pass deep read rewards pages that hold up under 83K-token analysis.
  • Entity depth and cross-referencing matter. A page that defines an entity, compares it to alternatives, cites benchmarks, and links to primary sources survives both passes better than a thin definition page.
  • Citations within your content are signal. Anthropic's deep read evaluates internal citation quality. Pages that cite authoritative sources are more likely to be cited themselves.
  • Long-form content is not penalized — Anthropic reads everything. Don't truncate for Anthropic the way you would for OpenAI's 200-word window.

🔄 Rank ≠ AI Citation: The Measurement Gap

What You Rank For and What the AI Retrieves Are Different Strings

Duane Forrester identifies three transformations between prompt and result: paraphrase (the user's query is rephrased), retrieval (the rephrased query fetches pages), and model judgment (the model decides what to cite). Each transformation shifts the query away from your target keyword.

Search volume disciplines rank but has no equivalent on the LLM side — no prompt-frequency index exists. The honest substitute: "frequency of citation across a prompt set run repeatedly over time" — a directional signal, not a demand number. Direction > precision.

🛠 Entity Optimization Framework

Platform-Specific Entity Optimization: A Summary Framework
StrategyGoogle (Gemini)OpenAI (gpt-5.5)Anthropic (Claude)
Content DepthFull page — every attribute a claimFirst 200 words — front-load definition + statLayered evidence — deep read rewards depth
Entity MarkupStructured data in DOM bodyQuick Facts block at topCross-referenced citations + Wikidata sameAs
Citation StrategyFactual density per paragraphCitation-ready first paragraphInternal citation quality + authoritative sources
Freshness SignaldateModified + visible timestampsRecent publish date in first 200 wordsVersioned content + changelog references
Audit MethodCheck if entity appears in Gemini at allTrack received-vs-cited gap per entityTrack cite frequency across repeated Claude runs

📊 Entity Phrasing Sensitivity Test

Your Entity May Only Appear Under One Phrasing

Because what you rank for and what the AI retrieves are different strings, your entity may appear under one query phrasing but vanish under another. Run this test for your top 5 entity keywords:

  1. Short query: "Hermes Agent"
  2. Long question: "What is Hermes Agent and how does it compare to Gobii?"
  3. Definition query: "What is Hermes Agent?"
  4. Comparison query: "Hermes Agent vs Gobii"
  5. Review query: "Hermes Agent review 2026"

Measure entity citation presence across all 5 phrasings on all 3 platforms. If your entity only appears under one phrasing, you have an entity phrasing risk — your content is not structured to survive the paraphrase → retrieval → judgment pipeline across query variations.

📋 Directional Citation Tracking

Don't Report a Single Number — Track a Trend Line

Run your entity query set repeatedly over time (daily or weekly) and track citation frequency as a trend. Honest reporting looks like:

  • "Hermes Agent appears in 4 of 10 Claude runs and climbing" — honest.
  • "Hermes Agent is cited in position 3" — not honest (implies a stable rank that doesn't exist).

For each entity page, track both Google rank and AI citation presence side by side. The gap between them is the signal. An entity page that ranks #2 but is never cited by any AI platform has a structural entity representation problem — likely missing a citable definition or hard stat in the retrieval window.

📜 Sources & Methodology

Primary Sources

  • Dan Petrovic / DEJAN: "How AI Search Grounding Actually Works: Google vs OpenAI vs Anthropic — Platform-by-Platform Breakdown" (June 13, 2026). dejan.ai/blog/grounding/
  • Duane Forrester: "Rank and AI Citation Aren't the Same Number — The Measurement Gap Deep Dive" (June 14, 2026). duaneforresterdecodes.substack.com

All grounding data (token counts, cite ratios, snippet forms) sourced from Petrovic's June 13 platform audit. Entity optimization framework synthesized from both sources plus Hermes Lab's own entity graph research.