Guide · 22 min
Is AI mentioning your brand or recommending it? Learn how to measure mentions, citations and recommendations, investigate gaps and improve buyer information.

You ask an AI assistant to suggest tools in your category. Your brand appears in the answer. You take a screenshot, share it with the team, and feel that your work is starting to pay off.
Then you read the answer again.
Your company is named in a long list. A competitor gets a detailed explanation of why it suits the buyer. Another gets recommended as the best place to start. Your brand is there, but the answer gives the reader little reason to choose it.
That difference deserves a closer look.
An AI mention means your brand is named. A citation identifies a source. A recommendation presents your product as a suitable choice for the buyer’s needs. These signals can overlap, but they answer different questions.
If you measure them separately, you can see where your marketing needs attention. Perhaps your research attracts citations, but your product pages leave important buying questions unanswered. Perhaps AI understands your category but describes your product using outdated information.
This guide will help you identify those gaps, measure them consistently, and decide what to work on first.
Imagine a fictional software company called Northstar CRM. Here are three illustrative answers to questions about customer management software. These are teaching examples, not captured AI responses.
| Signal | Illustrative answer | What it tells you |
|---|---|---|
| Brand mention | “Products in this category include Northstar CRM, Product B and Product C.” | The answer names the brand. It does not explain whether the product suits this buyer. |
| Source citation | “Document your sales stages before configuring a CRM,” followed by a link to Northstar’s implementation guide. | The answer references Northstar’s content as a source. It does not necessarily endorse its software. |
| Product recommendation | “For a five person sales team that needs basic pipeline management, Northstar CRM is worth shortlisting because it supports the workflow you described.” | The answer connects the product to the buyer’s requirements and suggests considering it. |
A recommendation can include a citation to the vendor, a third party, or no visible citation. A citation can support an answer that never names the source company in its main text.
These are separate dimensions, not three compulsory steps in a funnel.
There is another distinction worth keeping: being included in a recommended shortlist is different from being the first choice. Both can matter. Neither should be confused with a neutral reference or a warning against buying.
Read the sentence around your brand name. That is where the commercial meaning lives.

Suppose your company publishes a useful report about sales productivity. An AI answer references that report when explaining why sales teams struggle with admin.
That is a worthwhile result. Your work has become part of the answer’s evidence.
But the user asked about a problem, not which CRM to purchase. Counting this citation as a product recommendation would tell your team something the answer never said.
The reverse matters too. An assistant might recommend your product and cite a review site. A report that only counts links to your own domain would miss that recommendation.
Keep these questions separate:
Accuracy is essential. A recommendation based on an integration you do not offer can create the wrong expectations before someone reaches your website.
Citation checking is also worth the effort. Research published in Findings of EMNLP 2023 found gaps between claims and supporting citations in the generative search systems it evaluated. Those results describe systems tested in 2023, not current error rates. The useful lesson is to check whether a linked page actually supports the statement beside it. [1]
Before improving a score, look at the questions behind it.
A brand may appear frequently when users ask what a CRM is, yet rarely appear when a small sales team asks which CRM to buy. Combining those answers into one percentage hides the difference.
For your first audit, organise questions by the job the buyer is trying to complete.
| Buyer’s task | Example prompt | What to inspect |
|---|---|---|
| Understand a problem | “How can a small sales team stop losing track of follow ups?” | Whether the answer explains a problem your product genuinely solves. |
| Discover options | “Which CRMs should a five person sales team consider?” | Whether your product enters the suggested shortlist. |
| Check suitability | “Which CRMs support German language workflows and our existing email tools?” | Whether your capabilities and limitations are described correctly. |
| Compare products | “How does Northstar CRM compare with Product B for a small agency?” | Which differences the answer uses to guide the decision. |
| Resolve a final concern | “What should we check before migrating to Northstar CRM?” | Whether migration, pricing and support information is accurate and useful. |
Keep questions containing your brand name separate from questions that do not. If you put your company in the prompt, its appearance in the answer is weak evidence of discovery.
Use actual sales questions, support conversations and objections as starting points. Ask your sales team which details repeatedly delay a decision. Those details make better audit questions than another broad “best software” prompt.
For a DACH audience, test German and English questions separately where both languages matter. State the intended market and record the test conditions. A prompt mentioning Berlin is a scenario about Berlin; it does not automatically reproduce the experience of every buyer located there.
An answer is an observation from a particular moment and context. Treat it that way.
A September 2026 research preprint evaluated 1,536 product query responses across ChatGPT, Gemini and Google AI Overviews, including chatbot and API access where applicable. It found that recommended products often changed across repeated requests. It also found differences between consumer interfaces and their corresponding APIs. The study concerns its sampled product queries and tested configurations, so it should not be treated as a universal benchmark for every software category. [2]
For marketers, the implication is practical: repeat your checks and record how you ran them.
Personalisation adds another consideration. OpenAI documents that ChatGPT search can use location information and relevant saved memories when producing or rewriting searches. A new conversation alone does not remove every source of personal context. [3]
You do not need to reproduce every possible buyer experience. You need a clear baseline that your team can understand and repeat.
The following is a starter method for an internal audit. It is not an industry standard or a claim about how every visibility platform calculates its metrics.
Start with four discovery questions, four suitability questions and four comparison or evaluation questions.
Keep the wording natural. Include the constraints buyers actually care about: budget, team size, integrations, location, language or implementation effort.
Maintain separate groups for prompts that name your brand and those that do not. Save the exact wording so you can repeat it later.
As a manageable pilot, run each question three times across two relevant AI platforms. That produces 72 planned observations.
Three repetitions can reveal obvious inconsistency. They are not enough to establish a precise market benchmark. Expand the sample before making expensive decisions based on small differences.
Record the platform, visible model or mode, date, language, market, search setting and whether you used an API or consumer interface. Start each independent run in a fresh conversation and record memory or personalisation settings where available.
If a run fails, record the failure separately. A timeout is not a response in which your brand was absent. If an answer requests clarification or declines to recommend anything, retain that completed response and classify it accordingly.
Use a worksheet with these fields:
| Field | What to record |
|---|---|
| Prompt and intent | Exact question, buyer task, and whether the brand was named in the prompt. |
| Test conditions | Platform, mode, interface, language, market, timestamp and run number. |
| Brand mention | Yes or no, with the sentence containing the mention. |
| Recommendation | Yes, no or ambiguous, with the wording that supports the label. |
| Recommendation strength | Shortlisted, explicitly preferred, or conditional on stated requirements. |
| Source evidence | Exact URLs, distinguishing your domain from external sources. |
| Accuracy | Correct, incorrect or unverified, with the specific claim. |
| Competitors | Other brands suggested and the reasons given. |
| Full response | Saved text or a permitted export, plus a screenshot when useful. |
Agree on the classification rules before reviewing the results. For this audit, count a recommendation when the answer affirmatively suggests considering or choosing the product for the stated task.
A suitable option in a recommended list counts. A bare name in a category description does not. “Avoid this product for your requirements” is a mention, not a recommendation.
Keep ambiguous cases visible for human review. Do not quietly change the rules because a result is inconvenient.
For a fixed group of completed responses:
For this worksheet, require a recommendation to explicitly identify the brand. Recommendations are therefore a subset of mentions.
Imagine 60 completed responses to questions that did not name your brand. Your brand appears in 24 and is recommended in nine.
Your mention rate is 40%. Your recommendation rate is 15%. The difference is 25 percentage points.
You can also say that nine of the 24 mentions included a recommendation. That is 37.5% of mentions, using a different denominator. Label it clearly if you report it.
These fictional numbers describe an audit sample. They do not represent buyer reach, market share or conversion probability.
Report each platform and intent group separately before combining anything. Otherwise, a changing mix of questions can make performance appear to improve when the underlying answers have not.

The gap is a starting point for investigation. Read the answers before deciding that you need more content.
Consider a fictional product description:
“An intelligent platform that transforms sales performance.”
It sounds positive, but leaves a buyer with basic questions.
A more useful description would be:
“Northstar CRM helps small sales teams manage contacts, track deals and schedule follow ups. Its standard plan includes email integration and shared pipelines. Advanced territory management is not available.”
Only publish details that are true. The point is to make the audience, capability and limitation easy to understand.
Review your homepage, relevant product pages and marketplace profiles. If each describes a different business, resolve those inconsistencies before publishing another guide.
Suppose the answer recommends a competitor because the buyer needs a particular integration.
Check whether your product supports it. If it does, look at the public documentation. Is the integration explained? Are prerequisites and limitations visible? Is the information current?
If your product does not support it, this may be a legitimate product gap. Content cannot honestly remove that limitation.
For capabilities you do offer, a useful page should explain what works, how it works, who can access it and where the supporting documentation lives.
This gives your marketing and product teams a concrete issue to resolve.
An educational answer may have no reason to recommend software. Do not judge a useful research citation against an objective the prompt never expressed.
If the question does ask for products, inspect your commercial pages. Do they answer the buying criteria as thoroughly as your blog explains the general topic?
Useful additions might include a migration guide, an honest comparison, a pricing explanation or a case study describing a relevant customer situation.
Connect educational resources to these pages where it helps the reader. Avoid turning every research paragraph into a product pitch.
Open the cited page and find the relevant passage. Check whether it supports the statement, contains old information, or says something different.
Correct your own documentation first. Where an external page is wrong, request a factual correction with evidence. Keep track of the affected URL and claim so you can check again later.
If there is no supporting citation, record the source as unknown. You cannot reliably infer where a model learned a fact just from the answer.
Google says pages need to be indexed and eligible to appear with a snippet to qualify as supporting links in AI Overviews or AI Mode. Its existing SEO fundamentals still apply, and it does not require special AI schema or a new AI text file. Eligibility does not guarantee inclusion. [4]
OpenAI advises publishers to allow OAI-SearchBot if they want content included in ChatGPT summaries and snippets. Check hosting and CDN rules as well as robots.txt. [3][5]
Start with the relevant page. Can it be retrieved? Is important information available as text? Can users find it through internal links? Is the canonical URL correct?
Our AI Search Optimization Checklist covers the wider audit. For the broader strategy, see our guide to generative engine optimization.
You do not need to fix every possible weakness at once. Choose one buyer segment and the questions closest to its purchase decisions.
| Week | Focus | Concrete output |
|---|---|---|
| 1 | Establish the baseline. | Saved prompts, repeated answers, agreed labels and documented conditions. |
| 2 | Investigate the most important gaps. | Three specific issues linked to answer evidence and affected pages. |
| 3 | Publish useful corrections or improvements. | Updated product facts, documentation or decision support content. |
| 4 | Repeat the original checks. | A comparison using the same prompt set, plus a log of what changed. |
A good action is specific: “Publish the supported email integrations and plan restrictions on the integration page.”
“Improve AI authority” is too vague to assign or evaluate.
Record publication dates and allow for discovery and processing. Four weeks is a working cycle, not a deadline by which recommendations must improve. If results change, consider other explanations too, including platform changes, competitor activity and normal response variability.

A recommendation is evidence of what appeared in an answer. It is not evidence that somebody read it, clicked or bought.
Track measurable referrals alongside your answer audit. OpenAI says ChatGPT referral URLs include utm_source=chatgpt.com, which can help identify incoming visits in analytics. [5]
Then inspect what those visitors do: read product pages, start trials, book demos or become customers. Ask prospects how they discovered you, leaving room for them to mention AI in their own words.
Keep answer presence, observed website activity and reported discovery as separate measures. They can inform the same business discussion without pretending to provide perfect attribution.
For example, an increase in branded searches alongside better recommendation coverage is worth investigating. It does not prove that the recommendations caused the search increase.
Your monthly report should explain which buyer questions improved, what evidence changed, what work shipped and which commercial signals you can actually observe.
It means a source is referenced in that response. Inspect the claim the citation supports. A link to your research does not automatically endorse your product, establish overall trust or explain the system’s internal reasoning.
Yes. A recommendation may cite another publisher or contain no visible source link. Record the recommendation and the source separately so your reporting does not miss one while counting the other.
There is no universal target that makes sense across all prompt sets. Results depend on the questions, platforms, markets and classification rules. Compare your own trend and relevant competitors under the same conditions. Always show the sample size.
Add FAQs when they resolve real buying questions. Use appropriate structured data that matches the visible page. Google does not require a special schema type for its AI search features. Neither format guarantees a product recommendation.
There is no fixed timetable. Discovery, source updates and answer generation are outside your direct control. Track whether information becomes accessible, whether descriptions become more accurate and whether recommendations change across repeated checks.
The next time your brand appears in an AI answer, read past the name.
What does the answer say you are good for? Is that accurate? Does it give the buyer a reason to consider you? Which sources support that reason?
Those questions turn a screenshot into useful marketing work. They can reveal a missing explanation, an outdated profile, a documentation gap or a genuine product limitation.
Choose one important buyer question this week. Save several answers, inspect the evidence and fix the clearest issue you can substantiate. That is a manageable place to begin.
MentionX helps teams examine brand presence across AI answers, compare competitors and inspect the supporting evidence. Use that visibility to decide what deserves attention, then check the answers again after you make changes.
Explore MentionX or continue with How ChatGPT Recommends Brands.