A ChatGPT visibility audit is a structured test of whether ChatGPT names your company when buyers ask questions in your category. You build a prompt set of real buyer questions, run them in ChatGPT, record whether and where you are mentioned versus competitors and why, then translate the gaps into a prioritized fix list. You can do the whole thing yourself in an afternoon, and the value is the baseline: you cannot improve what you have never measured.
I’m Andrii Byzov, a fractional CMO for B2B tech. I run this audit for clients before any AI search optimization work, because it tells you exactly where you stand inside the tool buyers now use to build their shortlist. If you would rather have it done for you, that is the ChatGPT visibility audit I offer. The method below is the same one I use.
Key takeaways
- Build a prompt set of 15 to 30 real buyer questions, not your own keywords.
- Test each in a few phrasings, browsing on and off, and log the conditions.
- Score whether, where, and how you appear versus competitors.
- Sort gaps into entity, content, and authority buckets, then prioritize.
- Re-run on a schedule: results vary and change, so one snapshot is not truth.
Step 1: Build the prompt set
The audit is only as good as its prompts, and the most common mistake is testing your own marketing language. Buyers do not ask “best AI-native revenue platform.” They ask plainer, messier questions. Write down the real ones across three tiers.
Category questions are top-of-funnel: “best [category] tools for [segment],” “top alternatives to [the obvious incumbent].” Comparison questions are mid-funnel: “[you] vs [competitor],” “is [competitor] worth it for [use case].” Problem-led questions name a pain, not a product: “how do I [job to be done] without [common headache].” Aim for 15 to 30 total. Pull the exact wording from sales calls, support tickets, and your search query data, because those are the phrasings buyers actually use.
Step 2: Run the tests and log conditions
Now run each prompt. Two rules make the results usable.
First, vary the phrasing. Ask each question two or three ways, because small wording changes can flip who gets named. Second, control for conditions and write them down. ChatGPT behavior varies and changes, so a result is only meaningful next to its context. Log these for every run:
| Condition | Why it matters |
|---|---|
| Browsing on vs off | Browsing pulls live results; off relies on training data |
| Model version | Answers differ across versions and change over time |
| Account and memory | Saved memory and history can bias responses to you |
| Date and time | The same prompt can drift week to week |
A clean way to reduce bias is to run a second pass in a logged-out or temporary chat, so saved memory is not quietly inflating your own visibility. Keep both passes.
Step 3: Score what you find
For each prompt, record three things. Whether you appear at all. Where you appear: named in the top recommendations, mentioned in passing, or absent. How you are described: accurate and on-message, vague, or wrong. Then capture which competitors were named and, where ChatGPT offers it, the reasoning or sources behind the picks.
A simple spreadsheet works. One row per prompt, columns for your placement, competitor names, the description quality, and the conditions from step two. After 20-plus rows, patterns emerge fast: maybe you show up for comparison queries but vanish on category best-of questions, or you are named but described in outdated terms. That contrast is the real output of the audit.
Step 4: Diagnose the gaps
Sort every weak result into one of three buckets. This is where an audit becomes an action plan.
Entity gaps mean the model is unsure who you are. Symptoms: it confuses you with another company, describes you inconsistently across prompts, or hedges. The fix lives in your entity facts being consistent everywhere, which connects to getting cited by ChatGPT.
Content gaps mean the answer-shaped pages do not exist or are not clear. Symptoms: you lose problem-led and comparison queries because there is nothing crawlable that addresses them directly. The fix is honest comparison and roundup content the model can read and lift.
Authority gaps mean the sources the model trusts do not mention you. Symptoms: competitors who appear on best-of lists and third-party comparison pages get named and you do not, even when your product is stronger. This is usually the biggest lever and the one teams most often under-invest in.
Step 5: Prioritize and re-test
You will not fix everything at once, so rank by reach and effort. A gap that loses you a high-volume category query usually beats one on a niche phrasing. Entity fixes are often quick wins because consistency is mostly cleanup. Authority gaps take longer because they depend on third parties, so start them early even though they pay off slowly.
Then re-run the same prompt set on a schedule. Monthly is a sensible cadence for most B2B teams. Treat the audit like rank tracking for AI: a single snapshot tells you where you are, but the trend over time tells you whether your work is paying off. The mechanics of that ongoing measurement are in my guide on how to track AI visibility.
A realistic note on cost. The DIY version is mostly your time plus a paid ChatGPT plan for newer models and browsing; exact pricing depends on the tier you choose, so check current rates rather than budgeting from a number you half-remember. Paid AI-visibility tracking tools exist and can automate the re-testing, but they are not required for a first audit.
And the honest caveat that governs all of this: acting on an audit improves your odds of being named, it does not guarantee it. The model is a moving target you do not control. What the audit gives you is a clear, evidence-based picture of where you stand and what to fix next, which is far better than guessing.
FAQ
What is a ChatGPT visibility audit? It is a structured check of whether ChatGPT names your company when buyers ask questions in your category. You build a prompt set of real buyer questions, run them, and record whether you appear, where, how you are described, and which competitors get named instead. The output is a baseline plus a list of entity, content, and authority gaps to fix.
How many prompts do I need to test? Start with 15 to 30 real buyer questions across the funnel: category best-of queries, comparison queries, and problem-led queries. Test each in a few phrasings, with browsing on and off. Fewer than that and your sample is too thin; many more and the manual effort outpaces the value for a first pass.
Will the same prompt give the same answer every time? No. ChatGPT behavior varies and changes. Results shift with your account, saved memory, the model version, browsing state, and the time you run it. That is why you test multiple phrasings, log conditions, and re-run on a schedule rather than trusting a single answer as fact.
Can an audit guarantee ChatGPT will recommend me? No. An audit shows where you stand and what to fix, and acting on it improves your odds of being named. Nothing guarantees a citation, because you do not control the model. What you control is being the clearest, best-corroborated option in your category.
If you want this run for you, with a scored baseline and a prioritized fix list, that is exactly what my ChatGPT visibility audit delivers. You can also reach me on LinkedIn.
Andrii Byzov is a fractional CMO for B2B tech, focused on AI-native marketing and AI search visibility.