SEO & AI •

AI Search Visibility Audit: Does ChatGPT Recommend You?

How to audit your AI search visibility across ChatGPT, Perplexity, Gemini, and AI Overviews: a repeatable query set, scoring rubric, and log check.

A

Anas R.

— read

AI Search Visibility Audit: Does ChatGPT Recommend You?

Direct answer: you audit your AI search visibility by building a fixed set of 30 to 50 real buyer questions, running them from fresh, logged-out sessions across ChatGPT, Perplexity, Gemini, and Google AI Overviews on a repeatable monthly cadence, and scoring each response on a simple scale: absent, mentioned, recommended, or cited with a link. No platform hands you this number today. You build it the same way SEOs tracked rankings before rank trackers existed: by hand, on a schedule, in a spreadsheet.

The trigger for this audit is usually the same anecdote every marketer has heard by now: a prospect said "ChatGPT told me to talk to you," or worse, named a competitor instead. One anecdote confirms the channel exists. It tells you nothing about how often it happens, on which questions, or whether last quarter's content changes moved anything at all.

This is the measurement half of the AI-visibility problem. Our GEO guide and ChatGPT Search checklist cover what to build. This one covers how to find out, with logged evidence instead of anecdotes, whether it worked.

Quick verdict: the protocol in five lines

  • Query set: 30-50 questions across 5 archetypes (best X for Y, X alternatives, X vs Y, is X good for Z, how much does X cost), frozen for a full quarter
  • Engines: ChatGPT, Perplexity, Gemini, and Google AI Overviews, tested from fresh, logged-out sessions with no chat history or personalization
  • Scoring: 0 absent, 1 mentioned, 2 recommended, 3 cited with a link, logged per engine per query
  • Cadence: monthly, tracked on a rolling 3-month average, since single-query swings are usually model randomness, not real movement
  • Server-side check: grep your access logs for GPTBot, OAI-SearchBot, ChatGPT-User, PerplexityBot, ClaudeBot, and Google-Extended to see whether AI crawlers reach your pages at all

Why AI Visibility Isn't a Ranking You Can Look Up

A Google ranking is a fixed position on a page you can screenshot. An AI answer is not that. Ask the same question twice in two separate sessions and you can get two different sets of sources, a different order, even a different verdict, because the model samples its next words probabilistically and because session state, location, and account history all feed into what gets generated.

That unpredictability is exactly why most teams give up on measuring this and fall back on the anecdote: a support ticket, a sales call note, a stray mention on a call. It feels like data. It is not a metric, because it has no denominator and no comparison point.

What you can actually measure

Four things hold up across repeated runs: whether your brand appears at all, whether it appears as a recommended option rather than a passing mention, whether the response attaches a clickable citation back to your site, and which competitors show up alongside you on the same query. Tracked across a fixed query set over months, these four numbers move in a real, trackable direction.

What you cannot measure, and should stop trying to

There is no stable "position 3 of 10" to chase, no query-volume number to size the opportunity, and no single run that tells you why a score changed. Chasing false precision here wastes the time the audit is supposed to save. The goal is a directional trend across a frozen query set, not a leaderboard.

Build Your Query Set: 5 Question Archetypes That Matter

Before you can score anything, you need a fixed list of questions that mirror how real buyers actually phrase what they ask an AI assistant. Five archetypes cover almost all of it.

  • Best X for Y: category plus a qualifier, for example "best AI chatbot for small business" or "best AI chatbot for Shopify"
  • X alternatives: anchored to a specific incumbent, for example "Intercom alternatives" or "cheaper alternative to Chatbase"
  • X vs Y: a direct head-to-head, for example "Heeya vs Tidio" or "AI chatbot vs live chat widget"
  • Is X good for Z: a fit question tied to a use case or industry, for example "is an AI chatbot good for a law firm" or "is RAG good for customer support"
  • How much does X cost: a pricing question, for example "how much does an AI chatbot cost" or "AI chatbot pricing 2026"

These illustrate the pattern with a chatbot vendor as the example, not a published result. Swap in your own category and the archetypes hold for almost any B2B or B2C product.

Where the actual 30-50 questions should come from

Don't invent them from a whiteboard. Pull real phrasing from four sources: your Google Search Console queries already generating impressions, notes from sales calls where a prospect describes the problem in their own words, the questions your existing comparison and alternative-to pages already target, and the threads on Reddit or industry forums where people ask for recommendations in your category.

Freeze the list for a full quarter

Swapping questions between runs destroys your baseline; you'll never know if a score moved because the market changed or because you changed the test. Lock 30 to 50 questions, split roughly evenly across the five archetypes and your top 3 to 5 product use cases, and only revise the list at a quarter boundary, adding new questions as a fresh row rather than replacing old ones.

Run the Audit: A Repeatable Protocol Across Four Engines

The protocol matters more than the tool. Run it the same way every time or the numbers are not comparable month to month.

Fresh sessions, no personalization

Log out, or open a private/incognito window, before every run. In ChatGPT, start a temporary chat with reference to chat history turned off. In Perplexity, use an incognito session or the logged-out experience. In Gemini, start a brand-new conversation on a signed-out or fresh profile. For Google AI Overviews, search in an incognito window; location still influences results for anything local, so note your test location explicitly if it applies to your category.

Ask exactly once, log the full answer

Type the question exactly as written in your list, submit, and record the first response only. Do not follow up with a leading nudge like "and what about [my company]?" That is a legitimate second test of prompted visibility, but it measures something different and will contaminate your baseline if you mix it into the same score. Copy the full response text, or at minimum the passage mentioning any brand, plus a screenshot as backup, plus the date and time of the run.

The scoring rubric

Score Label What it means
0 Absent Your brand does not appear anywhere in the response
1 Mentioned Named in passing, not framed as a recommendation, no link
2 Recommended Listed as one of the suggested options, still no clickable citation
3 Cited with a link Named and backed by a clickable source citing your domain

Score every competitor named in the same response too. A run where a competitor scores 3 and you score 0 on the exact same question is the single most useful row in your entire spreadsheet.

Record It Properly: Spreadsheet Schema, Cadence, and Real Movement vs Noise

The audit is only as useful as the log behind it. A screenshot folder is not a tracking system; a spreadsheet with consistent columns is.

The columns that matter

  • Date and engine: run date, engine name, and model or version if the interface shows one
  • Query and archetype: the exact question text and which of the five archetypes it belongs to
  • Score (0-3): using the rubric above, for your brand
  • Competitors named: every other brand mentioned in the same response, with their own score
  • Citation URL: the exact page cited, when a score of 3 is logged
  • Quoted phrase: the sentence or clause that mentions you, verbatim, for later content audits
  • Notes: anything unusual, a follow-up question that changed the answer, a stale price quoted, a wrong claim

This is exactly the structure a downloadable tracking sheet would follow. If you want one built for you rather than assembled from scratch, Heeya's free AI chatbot tools hub is a reasonable place to start looking, alongside its RAG knowledge base readiness assessment, which shares the same instinct: score what you have before you try to fix it.

Monthly cadence, not daily

Running this weekly mostly measures model randomness, not real change. Monthly is frequent enough to catch genuine shifts and infrequent enough that a single content update has time to actually get crawled, indexed, and picked up before you re-test. Run it on the same day of the month, using the same frozen query list, every time.

Telling real movement from noise

One query flipping from a 1 to a 2 between two monthly runs is noise; the same model can produce a different answer to an identical prompt on back-to-back days. Track a rolling 3-month average score per archetype and per engine instead of reading single-month deltas. A signal looks like a sustained shift across several related queries in the same archetype, not one outlier row.

Read the Citations: What Pages Actually Get Pulled

Every time your audit logs a score of 3, the citation URL tells you something concrete: what kind of content earns a link, not just a mention. Patterns hold up across independent research on this, even though the mix shifts by industry and engine.

A citation-pattern analysis by Profound, covering roughly 10,000 commercial queries, found Reddit behind close to half of Perplexity's top-10 citations, well ahead of vendor sites and review platforms. That single data point explains why a category thread on Reddit or a well-answered question on a forum can outrank your own comparison page in an AI answer, something a classic SEO checklist rarely accounts for.

The domain types AI engines pull from most

  • Comparison and "X vs Y" pages with a clear verdict and a pricing table, the exact structure covered in our GEO for B2B SaaS guide
  • Structured FAQ and how-to content marked up with schema, covered in our schema.org FAQ and HowTo guide
  • Review platforms like G2 and Capterra, where third-party volume and recency both weigh in
  • Forum and community threads, Reddit chief among them, where the phrasing already matches how people ask the question

What your own citation tells you

When your domain is the one cited, note exactly which page and which passage. It is almost always a page with a direct, self-contained answer near the top, not a page that requires scrolling past an introduction to find the fact. That single observation, repeated across enough logged citations, is a better content brief than any generic best-practices list.

Fix List, Prioritized by Effort

Once a few months of audit data exist, the fix list writes itself: work on the archetypes and engines where you score lowest, starting with the changes that cost the least.

Low effort, do first:

  • Add a direct, one-sentence answer to the top of pages targeting your worst-scoring queries
  • Add or fix FAQPage schema on existing pages that already answer these questions in prose
  • Date and source every statistic on pages that get cited, since undated numbers are the easiest thing for an engine to skip in favor of a competitor's fresher page

Medium effort, next:

  • Build or rework "X vs Y" comparison pages for the specific competitors that keep outscoring you on the same queries, following the structure in our 30-point ChatGPT Search checklist
  • Write dedicated "is X good for Z" pages for the use cases where your audit shows you absent, rather than relying on a general product page to cover them

High effort, ongoing:

  • Earn genuine reviews on the platforms your citation log shows the engines actually pulling from
  • Participate honestly in the forum and community threads that keep surfacing as citations in your niche, rather than posting once and disappearing

Server-Side Signals: What Your Logs Say About AI Crawlers

The prompt-based audit tells you what AI answers say. Your server logs tell you whether AI crawlers can reach your content in the first place, a distinct and equally necessary check.

The user agents to check for

Grep your access logs for the crawlers each major AI provider currently documents: GPTBot, OAI-SearchBot, and ChatGPT-User from OpenAI, listed on its official bots page; ClaudeBot and related agents from Anthropic; PerplexityBot and Perplexity-User from Perplexity; and Google-Extended, documented on Google's crawler overview. A one-line check looks like this:

grep -iE "GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|PerplexityBot|Google-Extended" access.log | wc -l

Training crawlers vs at-query-time crawlers

These bots split into two jobs. GPTBot, Google-Extended, and the base ClaudeBot crawl for model training, feeding future versions of the model. OAI-SearchBot, ChatGPT-User, and PerplexityBot fetch pages at query time, when a live user asks a question the model needs to look up. Blocking a training crawler in robots.txt only affects future training runs; it does not remove you from citation eligibility. Blocking an at-query-time crawler does, because that is the request that fetches the page an answer is about to cite.

Why robots.txt is a request, not a lock

A Disallow line is a voluntary signal that well-behaved crawlers choose to respect, not an access control mechanism enforced by your server. That distinction became concrete news in August 2025, when Cloudflare reported that Perplexity was routing around sites that had blocked its declared crawler by switching to an undeclared, browser-mimicking user agent, a finding detailed by Search Engine Journal, after which Cloudflare removed Perplexity from its verified-bot list. If you want real enforcement rather than a polite request, that happens at your CDN or web application firewall, not in a text file. For the broader question of what a well-formed robots.txt and adjacent files should say, see our llms.txt guide, with the caveat that no file format alone guarantees a citation; the content still has to earn it.

Low or zero hits from the at-query-time crawlers is itself a finding worth logging next to your prompt-audit scores. It usually means one of three things: the crawler has never been sent to your site, it hit a block somewhere in your stack, or your content simply has not surfaced as relevant to any query yet. Cross-reference this against your AI chatbot KPI tracking if you run a chatbot widget yourself, since the same discipline of logging instead of guessing applies to both. It is also worth knowing that answers typed inside your own widget are invisible to these same crawlers, for the same reason they are invisible to Googlebot; our piece on whether chat widgets hurt SEO and Core Web Vitals covers why that content needs a static, crawlable home too.

FAQ: Auditing Your AI Search Visibility

How do I know if ChatGPT recommends my company?

Ask it directly, from a fresh, logged-out session with no chat history referenced, using the exact questions a prospect would type: "best [category] for [use case]," "[competitor] alternatives," or "[you] vs [competitor]." Run the same fixed list monthly and log whether you are absent, mentioned, recommended, or cited with a link each time. A single run is a snapshot; a repeated, logged run across months is a trend.

Is there a Google Search Console equivalent for AI search visibility?

Not a native one from the AI platforms themselves. Third-party monitoring tools such as Profound, Otterly, and Semrush's AI Visibility module automate parts of this tracking for a subscription fee. The manual protocol in this guide, a frozen query list scored on a fixed rubric, produces the same directional insight at zero cost, and is worth running even if you later add a paid tool on top.

How often should I run an AI search visibility audit?

Monthly, using the exact same frozen list of 30 to 50 questions each time. Weekly runs mostly capture model randomness rather than real change, since the same prompt can produce a different answer on consecutive days. Track a rolling 3-month average per archetype and engine rather than reading single-month swings.

Why does ChatGPT give different answers to the same question?

Large language models generate text by sampling probabilistically from possible next words rather than retrieving one fixed, cached answer, so identical prompts can produce different phrasing, different source selection, or a different order of options across separate sessions. Account history, location, and whether web search or live retrieval is enabled all add further variation. This is exactly why the audit protocol calls for fresh, logged-out sessions and a scoring average across repeated runs instead of trusting any single response.

Does blocking GPTBot hurt my chances of being cited by ChatGPT?

Blocking GPTBot in robots.txt stops OpenAI from using your content to train future models, but it does not stop OAI-SearchBot or ChatGPT-User, the separate crawlers that fetch pages at the moment a live user's question needs an answer. If you want to stay eligible for citations, allow the at-query-time crawlers even if you choose to block the training crawler.

What's the difference between being mentioned and being cited by an AI engine?

A mention names your brand somewhere in the generated text with no supporting link, the equivalent of a passing reference. A citation names your brand and attaches a clickable source pointing to a specific page on your domain, which is both a stronger visibility signal and a potential source of direct traffic. The scoring rubric in this guide separates the two explicitly, since treating them as the same metric hides which pages are actually earning links.

Can I automate AI visibility tracking instead of doing it manually?

Partially. Paid platforms like Profound, Otterly, and Semrush's AI Visibility module query multiple engines on a schedule and store the results, saving the manual labor of running each prompt by hand. They still rely on the same underlying idea, a fixed query set scored consistently over time, so understanding the manual protocol first makes any automated tool easier to configure and to sanity-check.

Does Google-Extended affect my classic Google Search rankings?

No. Google-Extended is a separate control token used only to opt your content in or out of training Google's Gemini models. Blocking it in robots.txt has no effect on standard Googlebot crawling, your position in classic Google Search results, or your eligibility to appear in Google AI Overviews, which draw on Google's regular search index rather than on Google-Extended's training feed.

None of this replaces the strategy work in our other GEO articles. It replaces the guesswork of deciding whether that strategy work is doing anything. Run the protocol for one full quarter before judging it: three monthly passes on a frozen query list is the minimum needed to separate a real trend from a single lucky, or unlucky, model response.

Want your own site to be the source AI engines cite?

A RAG-native chatbot on your own site gives AI engines and human visitors alike a single, structured, up-to-date source to pull answers from. See how Heeya's AI RAG expertise grounds every answer in your own content.

Start free: no credit card Explore free tools

Further Reading

Share this article:
Published on September 9, 2026 by Anas R.

Ready to build your AI assistant?

Join Heeya and transform your customer service with conversational AI.