AI Search Optimization

How to Get Cited by ChatGPT: What OpenAI Documents and What It Doesn't

How to get cited by ChatGPT: let OAI-SearchBot in, know what OpenAI documents versus guesses, and track chatgpt.com referrals. Checked 5 Oct 2026.

Long Nguyen Avatar

Long Nguyen

Fullstack Developer · AI Engineer · Researcher

• • 4 min read •

How to get cited by ChatGPT: what is documented and what is not

You can only control the first step: make sure ChatGPT's search crawler is allowed to read your pages. Everything after that, meaning which pages get quoted and which get dropped, is not described in any OpenAI page I found. This article keeps those two things apart.

Facts below were checked against OpenAI's own documentation on .

Question Does OpenAI document it? What to do
Can any public site appear in ChatGPT search? Yes Don't block OAI-SearchBot
Does blocking GPTBot remove you from search? No: separate controls Set GPTBot and OAI-SearchBot independently
How fast do robots.txt changes apply? Yes, about 24 hours for search Wait a day before judging a change
Can you track visits from ChatGPT? Yes, via a UTM parameter Filter analytics on utm_source=chatgpt.com
Which pages get cited, and in what order? Not in anything I found Treat all advice as hypothesis
Do llms.txt or schema markup raise citation odds? Not in anything I found Don't expect a guaranteed lift

Step 1: let OAI-SearchBot crawl your site

OpenAI's publisher FAQ says any public website can appear in ChatGPT search, and that to be included in summaries and snippets you must not block OAI-SearchBot. If your robots.txt blocks it, fix that first; nothing else on this page matters until you do.

Two details from OpenAI's crawler documentation are easy to miss:

  • Timing. OpenAI says search systems can take about 24 hours to adjust after you change robots.txt. If you edit the file at 9am, don't re-test at 9:05.
  • Firewalls. OpenAI recommends allowing its published IP ranges as well as the robots.txt rule. A Cloudflare or WAF rule that challenges unknown bots can block OAI-SearchBot even when robots.txt says yes. The IP list is published at openai.com/searchbot.json.

The user-agent string currently ends in OAI-SearchBot/1.4, and OpenAI notes the version number may change. When it fetches robots.txt itself, it may add a robots.txt marker to the string, so match on the token OAI-SearchBot and not on the full string.

User-agent: OAI-SearchBot
Allow: /

User-agent: GPTBot
Disallow: /

That file lets ChatGPT search reach you and asks OpenAI not to use your pages for training. It's the most common setup for publishers who want visibility without contributing training data.

OAI-SearchBot vs GPTBot vs ChatGPT-User: which bot controls what

OpenAI runs three user agents that people constantly confuse. They answer different questions.

User agent Job Control it with
OAI-SearchBot Surfaces sites in ChatGPT search robots.txt: this is the search opt-in or opt-out
GPTBot Crawls content that may be used for model training robots.txt: disallow it to opt out of training
ChatGPT-User Fetches a page because a person asked ChatGPT to OpenAI says robots.txt rules may not apply

The practical consequence: a log line from ChatGPT-User proves a person asked about your page, not that you are in the search index. And OpenAI says search opt-outs belong on OAI-SearchBot, so blocking ChatGPT-User is not the search switch.

How to check your robots.txt for ChatGPT in 10 lines

Don't eyeball it. This script uses Python's standard library to report how a robots.txt treats the three OpenAI agents for any URL.

from urllib.robotparser import RobotFileParser

BOTS = ["OAI-SearchBot", "GPTBot", "ChatGPT-User"]

def check(site, path="/"):
    rp = RobotFileParser(site.rstrip("/") + "/robots.txt")
    rp.read()
    for bot in BOTS:
        ok = rp.can_fetch(bot, site.rstrip("/") + path)
        print(bot, "ALLOWED" if ok else "BLOCKED")

check("https://example.com", "/blog/my-post")

One gotcha I confirmed by running it. Take a robots.txt with a blanket User-agent: * plus Disallow: /, and one named group that allows only OAI-SearchBot, and add a named GPTBot disallow. Python's parser reports OAI-SearchBot allowed, GPTBot blocked, and ChatGPT-User blocked, because ChatGPT-User has no group of its own and falls through to the wildcard.

For search eligibility that outcome is fine, since OAI-SearchBot is the one that matters. But it shows how a wildcard rule quietly catches every agent you didn't name. Python's parser is not OpenAI's parser, so confirm real behaviour in your server logs after the 24-hour window.

Why a blocked page can still show up as a link

Blocking OAI-SearchBot does not guarantee invisibility. OpenAI's FAQ says that if it learns a disallowed page's URL from a third-party search provider or from crawling other pages, and sees signals the page is relevant, it may show just the link and page title. OpenAI states this for ChatGPT Atlas.

The documented fix is a noindex meta tag. There is a trap: the crawler can only read a meta tag on a page it is allowed to fetch. If you want a page out and blocked in robots.txt at the same time, the noindex is never seen. Allow the crawl, serve noindex, and let the tag do the work.

How does ChatGPT choose which sources to cite?

Honest answer: OpenAI hasn't published the selection logic. I searched its help center and developer documentation and found crawler rules and a referral parameter, but no ranking criteria, no scoring factors, and no list of signals.

That matters because the internet is full of precise-sounding numbers: percentage of retrieved pages that get cited, how much structure multiplies your odds, how many referring domains you need. Almost all of it comes from SEO and GEO tool vendors measuring correlations in their own samples. None of it is an OpenAI statement, and several posts repeat each other's figures. Use them as hypotheses to test, not as rules.

One recurring claim, that ChatGPT search leans on Bing's index, is also not confirmed in the OpenAI pages I read. Verifying your site in Bing Webmaster Tools takes minutes, so it is cheap insurance. Just don't plan around it as a fact.

How to write pages that are easy to cite

Since selection is undocumented, this section is my judgment, not OpenAI policy. The reasoning is simple: a system that assembles an answer from retrieved pages needs a passage it can lift without rewriting. Make that cheap.

  • Answer first. Put the direct answer in the first two sentences under each heading, then the nuance.
  • One question per H2. Phrase the heading the way a person would ask it, so each section stands alone as a quotable unit.
  • Be specific and dated. Versions, numbers, and a visible date give a passage something a generic page can't match.
  • Use tables for comparisons. A table is unambiguous to a parser and fast for a human.
  • Name the author and cite primary sources. It makes the page checkable, which is what trust comes down to for any reader, human or machine.
  • Keep content in the HTML. If the answer only appears after JavaScript runs or behind a login, a crawler may never see it.

Notice what's missing: tricks. There is no documented schema, file, or phrase that forces a citation. If you'd rather not work through crawler access, firewall rules, and page structure yourself, this is the kind of work covered by Netalith's SEO, AEO and GEO service.

How to track traffic that comes from ChatGPT

OpenAI says sites that allow OAI-SearchBot can measure referrals because ChatGPT adds utm_source=chatgpt.com to links it sends. In Google Analytics, filter sessions where the source is chatgpt.com, and compare the landing pages against the pages you expected to be cited.

Treat this as your feedback loop. You can't see citations directly, but you can see which URLs earned a click. Check weekly, not daily, given the 24-hour lag on crawler changes and how variable AI answers are from one asking to the next.

Do llms.txt or schema markup help you get cited by ChatGPT?

Neither appears in the OpenAI documentation I read as a factor for ChatGPT search. Schema markup has well-established value for traditional search features, and an llms.txt file is cheap to publish, so neither is a bad idea. Just don't promise yourself, or a client, a ChatGPT citation as the payoff. Anyone who sells that guarantee is selling something OpenAI hasn't documented.

A realistic order of operations

  1. Confirm OAI-SearchBot is allowed in robots.txt and not blocked at the firewall.
  2. Decide GPTBot separately, based on whether you accept training use.
  3. Wait about 24 hours, then check server logs for OAI-SearchBot.
  4. Rewrite your best pages so each section answers one question in its first two sentences.
  5. Watch utm_source=chatgpt.com referrals weekly and compare against the pages you rewrote.

If you want a second pair of eyes on steps 1 to 3, Netalith's $20 Audit + Roadmap checks your crawler access and delivers a prioritized fix list in 24 to 48 hours. Sources for this article: OpenAI's Publishers and Developers FAQ and OpenAI's crawler overview.

FAQ

Frequently asked questions

Can any website be cited by ChatGPT?

OpenAI says any public website can appear in ChatGPT search. To be included in summaries and snippets, you must not block OAI-SearchBot. Appearing is not guaranteed, and OpenAI doesn't publish how sources are chosen.

Does blocking GPTBot stop ChatGPT from citing my site?

No. GPTBot relates to model training and OAI-SearchBot relates to search. They are separate robots.txt controls, so you can disallow GPTBot and still allow OAI-SearchBot.

How long after changing robots.txt will ChatGPT notice?

OpenAI says its search systems can take about 24 hours to adjust after a robots.txt update, so wait at least a day before judging the change.

How can I see traffic from ChatGPT in analytics?

ChatGPT adds utm_source=chatgpt.com to referral links for sites that allow OAI-SearchBot. Filter on that source in Google Analytics or any tool that reads UTM parameters.

Does llms.txt or schema markup guarantee a ChatGPT citation?

No. Neither is documented by OpenAI as a ChatGPT search factor. Both can be useful for other reasons, but a citation can't be promised from either.

Can a page I blocked still appear in ChatGPT?

Possibly. OpenAI says it may show just a link and title for a disallowed page if it learns the URL elsewhere. Use a noindex meta tag, and keep the page crawlable so the tag can be read.

Stay updated with Netalith

Get coding resources, product updates, and special offers directly in your inbox.