
Audit mentions across ChatGPT, Claude, and Gemini
<aside> 📊
An AI search mention audit is not a vanity dashboard. It is a prompt-by-prompt baseline that shows where ChatGPT, Claude, and Gemini name you, cite you, or skip you — so you can fix the underlying gap, not just watch a score move.
</aside>
When a buyer asks ChatGPT for the best tool in your category, the answer usually names one to three brands. If you are not one of them, there is no SERP position to defend and no Search Console row that explains the miss. You simply were not in the answer.
This playbook is how to audit that gap across ChatGPT, Claude, and Gemini in a way you can re-run. It covers prompt sets, a logging sheet, competitor comparison, what "missing" actually looks like, and when to stop doing it by hand. For ongoing monitoring after the baseline, see AI search tracking.
The three assistants do not share a ranking system, and they do not pull from the same evidence.
| Platform | What usually shapes answers | What a weak showing often means |
|---|---|---|
| ChatGPT | Training data plus live browsing / Bing-style retrieval when search is on | Thin third-party corroboration, weak structured entity data, or pages Bing never surfaces |
| Claude | Heavier reliance on durable training-data consensus and high-authority third-party text | Few lasting editorial mentions, review profiles, or category write-ups that survive model updates |
| Gemini | Tight coupling to Google's index, Knowledge Graph, and entity clarity | Messy schema, inconsistent NAP / Organization data, or pages Google cannot cleanly interpret |
Cross-platform disagreement is normal, not a measurement bug. Analyses of multi-million citation sets show models weight owned sites vs directories differently, and brand mentions disagree across assistants a large share of the time — often cited around 62%.[1] Auditing ChatGPT alone tells you almost nothing about Claude or Gemini.
Collapse everything into one "AI visibility score" and you lose the diagnosis.
A useful directional benchmark used in 2026 audit write-ups: aim for roughly a 20% mention rate on category-defining recommendation prompts before you call the footprint healthy.[3] Treat that as a starting bar, not a law of nature — your category concentration and competitor set matter more than the round number.
Build 25–40 prompts (enough for a first baseline; expand to 50+ once you automate). Write them the way a stranger would ask, not the way your product marketing deck reads.