Quick answer: run a fixed set of ten to twenty-five prompts across ChatGPT, Gemini and Perplexity, in a clean session, and log five things for each: whether you were named, what was claimed, whether the claim is accurate, which sources were cited, and the date. Run the same set monthly. A single check tells you nothing, because answers vary between runs — the value is in the comparison over time. And before fixing anything, establish whether the error came from a source the assistant retrieved or from the model itself, because only the first kind can be corrected.
Most companies discover what AI assistants say about them by accident: a prospect mentions it, or someone types the company name out of curiosity. By then the answer has been shaping buying decisions for months.
The alternative isn't complicated, but it has to be systematic. Below is the audit method — what to ask, what to record, and how to read what comes back.
Why one check tells you nothing about your AI visibility
The first thing to understand is that assistant answers aren't stable. Ask the same question twice and you may get different companies named, different descriptions and different sources. Ask in a different session, from a different account, or with the phrasing slightly changed, and the variation widens.
That has two consequences for the audit.
The first is that a single run is an impression, not a measurement. Finding your company named once proves it can happen, not that it usually does. Finding it absent once proves nothing at all.
The second is that the unit of measurement has to be a set, not a query. If you run twenty prompts and appear in six answers, that's a number you can compare against next month's six, eleven or two. One prompt gives you nothing to compare.
Building the prompt set for an AI brand audit
The prompt set is the whole method. Build it once, carefully, and then don't change it — changing prompts between runs destroys comparability.
Ten to twenty-five prompts is the workable range. Cover four types:
Branded prompts. Your company name directly: what it does, what it costs, who runs it, where it operates. These surface factual errors fastest, and they're the ones prospects actually run before a call.
Category prompts. The service without your name: "who does X for Y in Z". This is where you find out whether you're in consideration at all, and who else is.
Comparison prompts. Your company against a named competitor, or "alternatives to [competitor]". Uncomfortable, and the most informative.
Problem prompts. The situation a buyer describes before they know what service they need — with the constraints, the industry and the location a real person would include.
Write them the way a person actually types, not as keyword strings. "We're a dental clinic in Warsaw and our website gets no enquiries, who could help" is a prompt. "dental website development" is a search query, and assistants behave differently with the two.
What to record in an AI brand audit for every prompt
Consistency matters more than volume. For each prompt, log:
- Assistant and date — Answers differ by platform and drift over time
- Named or not — The primary metric for the whole audit
- What was claimed about you — Where factual errors surface
- Accurate / outdated / wrong — Errors get triaged by severity, not by order found
- Who else was named — Your actual competitive set in AI answers, which often differs from the one you assume
- Sources cited — The starting point for any correction
- Screenshot — Answers change; without a screenshot you can't prove what was said
A spreadsheet is enough. The tooling matters far less than running the same thing every month.
How to audit ChatGPT, Gemini and Perplexity separately
Use a clean session. Log out, or use a private window. A logged-in account carries memory of your previous conversations, and an assistant that has already discussed your company with you will name it far more readily than it would for a stranger.
Run each platform separately. ChatGPT, Gemini, Perplexity and Google's AI answers draw on different sources. Presence in one implies nothing about the others, and the same error can exist in one and not another.
Note whether search was used. ChatGPT answering from its training data and ChatGPT answering after searching the web are effectively two different systems with two different failure modes. If the answer includes citations, the web was used. If it doesn't, you're looking at the model's memory.
Follow the citations. Gemini and Perplexity show their sources, which makes tracing an error straightforward — open what they cited and you'll usually find the wrong claim sitting there in plain text. ChatGPT often doesn't, in which case you search the incorrect claim itself and see which page ranks for it.
Retrieval error or training-data error: how to tell them apart
This is the distinction that determines whether a fix is possible at all, and most guidance on the topic skips it.
A retrieval error happens when the assistant searched, found a page containing wrong information, and repeated it. The wrong claim exists somewhere public: an outdated profile, an old press release, a competitor's comparison page, a directory listing nobody has updated in three years. These are correctable, because the source is correctable.
A training-data error happens when the model states something from its own memory, with no source behind it. Your company was described a certain way in the data the model learned from, and it repeats that. These don't respond to source corrections in the short term — they persist until the model is retrained or until search-grounded answers override them.
How to tell them apart: ask the same question with and without search enabled, where the platform allows it. If the answer changes when the assistant searches, the correct information is retrievable and the problem is that your sources aren't dominant enough. If the wrong claim appears in both, it's in the model.
The practical implication is uncomfortable but worth knowing early: a fraction of what assistants say about you cannot be fixed this quarter, and effort spent trying is wasted. What can be done is making the correct facts so consistently present in retrievable sources that search-grounded answers stay right even while the model's memory lags.
How to read the results of an AI brand audit
Three questions, in order.
Are you named at all? If you appear in two of twenty answers, the problem is visibility, not accuracy, and the work is the one covered in the guide on getting recommended by AI assistants.
When named, is the description accurate? Wrong specialisation, outdated services, wrong geography, a price that hasn't been current for two years. Triage these by consequence: anything that affects a buying decision — pricing, capability, whether you serve their industry or country — comes first. Cosmetic staleness can wait.
Which sources keep appearing? If the same three pages get cited across most answers, those pages are your evidence base, whether you control them or not. That's where correction work starts, and it's often not your own website.
How often to re-run an AI brand audit and what changes to expect
Monthly is the right cadence for most businesses. Weekly produces noise — the natural variation between runs will swamp any real change. Quarterly is too slow to catch an error while it still matters.
Re-run the identical set, in the same conditions, and compare the count of answers naming you and the list of errors still present. Corrections don't appear immediately: source changes have to be recrawled before they can influence an answer, and the lag differs substantially between platforms — retrieval-based systems reflect changes in weeks, while anything grounded in a model's memory may not change at all until a retrain.
Set expectations accordingly. This is a slow-moving measurement, and the value is in the trend across several months, not in any single month's result.
How I run AI brand audits
What the work covers: building the prompt set for your market and buyers, running the baseline across platforms, tracing each error to its source, separating what's correctable from what isn't, and setting up the measurement so you can see the trend rather than guess at it.
I'll say the limits plainly: nobody can guarantee what an assistant will say, and some errors won't move until a model is retrained. What can be done is establishing what's actually being said, fixing what's fixable, and knowing which is which.
Where to start auditing what AI says about your company
Reduced to one principle: build the set once, run it in a clean session, log the same fields every time, and compare month to month.
A practical step for today: write ten prompts a buyer of yours would actually type — three branded, three category, two comparison, two problem-based. Run them through ChatGPT and Perplexity in a private window and record what comes back. That hour gives you your baseline, and everything afterwards is measured against it.
If you'd like to know what assistants are currently telling your buyers, write to me and we'll run the audit together.






