Blog Generative AI

How to get your business cited by ChatGPT and AI search

Being mentioned and being cited are different jobs with different levers. What is observable, what can be measured, and the multiplier that is not in the data.

Alexander González August 24, 2026 2,782 words

A page can be the first organic result on Google for a query and never appear in ChatGPT's answer about the same thing. That is not a ranking failure. The two systems do different jobs: one orders a list of pages, the other drafts an answer and then looks for something it can attribute. The unit of optimization stops being the page and becomes the sentence, and most business sites have never written a sentence a model could safely quote.

How do you get cited by ChatGPT?

There is no published criterion and no way to request it. What is observable is that it cites sources that are already recognisable and that answer early with traceable figures. A new site rarely appears; a mid-authority one appears on niche questions; a consolidated one appears on general ones.

Quick answer

OpenAI crawlerWhat it doesBlocking it means
OAI-SearchBotIndexes pages so they can be cited in answersYou are out of the citations
ChatGPT-UserFetches a page a user specifically asked forSomeone pastes your link and it cannot be read
GPTBotCollects data to train modelsYou stay out of future training

What kind of problem this is

The model most people arrive with is that ChatGPT is a search engine with a different interface, so appearing in it must be ranking with extra steps. That belief sends the work to the wrong place.

A search engine sorts. A generative system writes an answer first and then looks for something to stand behind it. In the first you compete for a slot; in the second you compete to be the most convenient piece of evidence available. Those are different competitions, and a site can win one while never entering the other.

What changes in practice is the scope of the work. Not the page as a unit, but the individual passages that assert something checkable. A long article containing no attributable claim is invisible to this process no matter how well it is structured.

The three checks that come before any content work

These take about twenty minutes total and they gate everything else. Content work on a site that fails them produces nothing.

1. Is OAI-SearchBot allowed

Open yourdomain.com/robots.txt and read it. A rule blocking OAI-SearchBot, or a blanket rule that catches it, removes the site from the citation pool entirely.

This is worth stating plainly because the three crawlers get conflated constantly: blocking GPTBot keeps a site out of model training and has nothing to do with citations. Plenty of sites blocked all three during the 2023 wave of publisher opt-outs and never revisited the decision. If the goal is to be cited, that file is the first place to look.

2. Is the site in Bing

Bing Webmaster Tools is free, takes minutes to verify, and almost nobody in the US small-business segment has an account. A site absent from Bing's index is absent from a meaningful share of the candidate pool that generative answers draw from.

The verification also surfaces crawl problems that Google's tools report differently, which occasionally explains an indexing issue nobody could account for.

3. Is the content readable without JavaScript

If the text only exists after a script runs, assume it may not be read. This affects service pages built on page builders more often than blog posts, and it is the failure that looks like nothing at all: the page appears fine in a browser and arrives empty at anything that does not execute scripts.

What actually gets quoted

A claim with a number and a source

This is the family of techniques with the strongest measured effect. Princeton's GEO study, which is the paper that introduced the term, tested nine optimization methods across roughly 10,000 queries and found that adding sourced evidence (statistics, quotations, expert references) raised visibility in generated answers by 30 to 40 %, with its best methods reaching 41 % on the position-adjusted metric. The full breakdown of what it measured is in the Spanish guide on the Princeton GEO study.

The practical form is narrow. "We help businesses grow" cannot be cited. "Organic click-through rate for queries showing an AI Overview fell from 1.76 % to 0.61 % in Seer Interactive's 2025 study of 3,119 queries" can be, because it survives being lifted out of the page.

Answers that survive extraction

A passage gets used when it makes sense on its own. That means the answer sits immediately under the question, states the conclusion in the first sentence, and does not depend on a paragraph three screens up for its meaning.

The test is mechanical: copy any 50-word block from your page, paste it somewhere with no context, and see whether it still says something. Most marketing copy fails this instantly, because its meaning lives in the buildup rather than the sentence.

Being the origin rather than the repeater

When five sites carry the same statistic, the one cited tends to be the one that produced it or the one that documented it most precisely. A site with its own measurements, even small ones, has something no competitor can restate.

That is the durable advantage in this whole area, and it is also the slowest to build. Nothing about publishing your own data can be outsourced to a content vendor.

Where the return is highest

The Princeton study has a second finding that gets quoted far less than the 41 % and matters more when deciding where to spend.

The effect is not evenly distributed by starting position. Adding source citations lifted visibility by 115.1 % for pages sitting around fifth place, while pages already in first place lost 30.3 % in that same test. The pages that benefit most are the ones already competitive but not dominant, which is the opposite of where most teams focus.

There is a caution attached. The study measures presence inside generated answers, not visits. Citation and traffic are related and are not the same thing, and any promise that converts one into the other is inventing the exchange rate.

When NOT to invest in this

When the site is not indexed in Google either. Generative systems select from what is already indexed, so an indexing problem is the prior problem, and the diagnosis for it is in the Spanish guide on why Google does not index your page.

When nobody asks a machine about your category. A local plumber gets almost nothing here, and the return on the same hours spent on local SEO for small business is not close. The queries that produce citations are informational and comparative, not "near me".

When there is nothing original to publish. If every claim on the site is restated from elsewhere, the work is producing something worth citing first, and that is a different project.

The other systems, which do not work the same way

ChatGPT is one of several, and what works there does not transfer whole.

Google's own summaries are built from pages that could already appear in search, so eligibility is the same as ranking eligibility: answer engine optimization vs SEO.

Being findable at all comes first. A page that is not indexed cannot be cited by anything, and that check is in why is my page not indexed and how to set up Search Console.

And the groundwork is ordinary: answer early, name figures with their source, and be recognisable as an entity. The order of that work is in basic SEO checklist and what is SEO, with the research side in free keyword research.

For a business serving a Spanish-speaking market in the US, the language question changes the answer: SEO for Latino businesses in the USA and translating a website for SEO.

Two things called the same, with different levers

Most advice on this collapses two outcomes that look identical in a screenshot and are not the same job. A third collapse happens one level up, where the phrase "SEO for AI" gets used both for this work and for using AI tools to do ordinary SEO, two jobs separated in SEO for AI.

Being mentioned is your name appearing inside the answer text. It comes from the model having absorbed enough about you from across the web that your name is part of how it describes the category.

Being cited is your domain appearing as a linked source under the answer. It comes from a retrieval step that goes and fetches pages at the moment of answering.

The distinction decides where the work goes. Mentions respond to being written about elsewhere; citations respond to having a page that answers the specific question well and is reachable. A brand can be mentioned constantly and never cited, and a small site can be cited regularly without anyone knowing its name.

The number being repeated, and why it does not say what people think

One figure is circulating in almost every guide on this topic: a study of 75,000 brands published by Ahrefs, reporting a correlation of 0.664 between brand mentions and AI visibility against 0.218 for backlinks. It is usually summarised as brand mentions mattering "three times more".

Two problems with that summary, and neither is about the study.

Correlation coefficients do not divide. A correlation of 0.664 is not three times a correlation of 0.218 in any sense that means anything. The ratio is arithmetic performed on the wrong kind of number, and it is the version that travels.

And correlation is doing heavy lifting here. Brands that get mentioned a lot are usually big, and big brands are visible everywhere for reasons that have nothing to do with the mentions themselves. The study can show the association; it cannot separate the two.

None of that makes the direction wrong. Being written about elsewhere probably does help. What it does not support is a budget line calculated from a multiplier that was never in the data.

What can be measured, and what cannot

This is the part that decides whether the work can be defended to whoever is paying for it.

Clicks that leave a Google summary are counted, mixed in with the rest of search traffic and with no filter to separate them. What leaves no trace anywhere is the question that ends without a click, which in conversational products is most of them.

So the honest metric is the mention, not the visit. You check it by asking: take the ten questions a customer would ask before buying, put them to the system, and write down who gets cited. It is manual, it is imprecise, and it is the only method that corresponds to what actually happens.

The sample size problem is real and worth stating. Asking a question twice can produce two different sets of citations, so a single check tells you almost nothing. What tells you something is the same ten questions asked over several weeks, which is slow and is the honest cost of working in this space.

And there is no dashboard coming. No published criterion, no submission form, no verification. Anyone selling guaranteed citation is selling something that has no mechanism behind it.

Mistakes that repeat

Data and transparency

The study of 75,000 brands reporting correlations of 0.664 for brand mentions and 0.218 for backlinks was published by Ahrefs and is cited here as an example of a figure being summarised incorrectly, not as evidence for a recommendation: correlation coefficients cannot be divided into a multiplier, and the association is confounded by brand size. The absence of any submission form or verification path for ChatGPT citation, and the distinction between being mentioned and being cited, were checked against the first results for this query on 12 August 2026. No figure appears here for how often any site is cited, because none is published.

The 30 to 40 % visibility gain from adding sourced evidence (statistics, quotations, expert references) and the 115.1 % gain for pages around position five from citing sources come from the GEO study by Aggarwal, Murahari, Rajpurohit, Kalyan, Narasimhan and Deshpande, arXiv 2311.09735, presented at KDD 2024; the study's 41 % figure belongs to its best-performing methods on the position-adjusted metric, not to statistics alone. The click-through figures come from Seer Interactive's study published 4 November 2025, covering 3,119 informational queries across 42 organizations over a window from June 2024 to September 2025: organic click-through rate fell from 1.76 % to 0.61 % where an AI Overview was present, and from 2.72 % to 1.62 % where it was not, so only the difference between the two declines is attributable to the summary. The crawler names and their stated purposes are documented publicly by OpenAI. No figure appears here for how often ChatGPT cites any particular site: that number is not publicly measurable and any article giving one is estimating. The portfolio experience behind the operational judgment covers sites recording more than 300 million impressions a year in Search Console. Verified as of August 2026.

Primary sources, opened on 13 August 2026: the Seer Interactive study; the GEO paper (arXiv 2311.09735).

What this changes

The reflex when a new channel appears is to ask what to add. Here the more useful question runs the other way.

Open any page you would like to be cited on and look for one sentence a stranger could quote against your name. Not a claim about what you do, but an assertion about the world, with a number and somewhere it came from. On most business sites there is not one, on any page, and the entire optimization problem is downstream of that. The crawler settings take twenty minutes. Having something worth quoting is the part that takes a year.

Frequently asked questions

How do I get ChatGPT to cite my website?

Allow OAI-SearchBot in your robots.txt, get the site indexed in Bing, and write passages that state a specific claim with a figure and a source directly under the question they answer. Generative systems select evidence they can attribute, so the work happens at the level of individual sentences rather than whole pages.

Does blocking GPTBot stop ChatGPT from citing me?

No, those are separate crawlers with separate purposes. GPTBot collects data for model training, while OAI-SearchBot indexes pages so they can be cited in answers. Blocking the first keeps you out of training and leaves citations unaffected; blocking the second removes you from the citation pool while training access is unchanged.

Is getting cited by AI worth anything if nobody clicks?

Sometimes, and it depends on the query. A citation in an answer about a category you sell into functions as a recommendation to a buyer who is comparing, which has value even without a visit. A citation in a general informational answer often produces nothing measurable, which is why treating citations as a traffic channel disappoints.

What kind of content gets cited most often?

Content that makes checkable claims. The Princeton GEO study measured its strongest gains, 30 to 40 % more visibility, in techniques that add sourced evidence: statistics with attribution, quotations and cited sources. Marketing copy performs badly here for a structural reason: it asserts benefits rather than facts, and a benefit cannot be attributed to anyone.

Do I need a separate strategy for ChatGPT and Google?

Not a separate strategy, but an added layer. Both draw candidates from search indexes, so indexing and crawlability serve both. What differs is the finishing work: Google rewards a page that answers a query better than others, while a generative system rewards a passage it can lift out and attribute without ambiguity.

Most sites do not have a ranking problem

They have a what-happens-next problem. You can rank first and still sell nothing. The diagnostic looks at both and tells you which one is costing you money.

See the diagnostic