AI Search Optimization

How to Get Cited by AI: A Source-Checked Guide for Six Answer Engines

How to get cited by AI in Claude, Gemini, Google AI Overviews, Perplexity, Grok and Copilot: what each documents, one robots.txt, and how to measure.

Long Nguyen Avatar

Long Nguyen

Fullstack Developer · AI Engineer · Researcher

• • 5 min read •

How to get cited by AI: the short answer for six engines

Every engine below can only cite a page it can fetch and, in most cases, index. That part is documented and controllable. What each engine does after that, meaning which of the fetched pages it quotes, is largely undocumented, and this guide says so wherever it applies.

Facts were checked against each vendor's own documentation on . Vendor pages change, so re-check before you act on a rule.

Engine What you can control Citation reporting the vendor documents
Claude Three separate bots: ClaudeBot, Claude-User, Claude-SearchBot None found
Gemini and Google AI Overviews / AI Mode Normal Search indexing and snippets, plus Google-Extended Search Console Performance report, Web search type
Perplexity PerplexityBot and Perplexity-User None found
Microsoft Copilot Bing crawling and robots.txt AI Performance in Bing Webmaster Tools
Grok No publisher controls found in xAI documentation None found

What every AI engine needs before it can cite you

The baseline is the same everywhere: a public page that returns its content in the HTML, is not blocked by robots.txt or a firewall, and is indexed by the engine's search layer. Google says this outright for AI Overviews and AI Mode. A page needs to be indexed and eligible to appear in Search with a snippet to be shown as a supporting link, and Google lists no additional technical requirements.

The most common failure is not content quality, it's a blanket bot-blocking rule. A firewall preset or a User-agent: * block can shut out several engines at once without anyone deciding to. Fix access first; judge content second.

How to get cited by Claude

Anthropic documents three bots, and the distinction matters because blocking the wrong one removes you from the wrong thing.

Bot Purpose If you block it
ClaudeBot Collects web content that may contribute to model training Signals your future content should be excluded from training datasets
Claude-SearchBot Crawls to improve the quality of search results for users Prevents indexing, which lowers your visibility and accuracy in Claude's search results
Claude-User Fetches pages when a Claude user asks for web access during a conversation Stops retrieval for that user's query, which reduces visibility for user-directed search

For citations, allow Claude-SearchBot and Claude-User. Whether to allow ClaudeBot is a separate training decision. Anthropic says its bots honor robots.txt, supports the Crawl-delay directive, and publishes its bot IP addresses at claude.com/crawling/bots.json for firewall allowlists.

Anthropic does not publish how Claude picks which retrieved pages to cite, and I found no Claude equivalent of a citation report. Your evidence will be server logs and referrers.

How to get cited by Google AI Overviews, AI Mode and Gemini

Google's answer is deliberately unglamorous: there is nothing special to do. Its documentation says no new machine-readable files, AI text files, or markup are needed to appear in AI Overviews or AI Mode. Standard Search eligibility is the whole requirement.

  • Query fan-out. Google says both features may issue multiple related searches across subtopics to build an answer and pick supporting pages. A page that answers one subtopic cleanly can be selected for that subtopic even if it doesn't rank first for the head query.
  • Preview controls cut both ways. nosnippet, data-nosnippet, max-snippet and noindex limit what Search can show. If a page carries a nosnippet rule you forgot about, it may be ineligible, so audit your templates.
  • Google-Extended. Google says this control limits use of your content for AI training and for grounding in other Google systems. Blocking it is a training opt-out, but read that wording before you do: grounding is the part that decides whether content is used in answers.
  • Reporting. Traffic from AI Overviews and AI Mode is counted in Search Console's Performance report under the Web search type. It is not split out in that documentation, so you can't isolate it from ordinary clicks there.

Gemini needs a caveat. Google's publisher page covers AI Overviews and AI Mode. I found no separate Gemini-app citation rules beyond the Google-Extended control above, so don't assume AI Overviews behaviour carries over to the Gemini app.

How to get cited by Perplexity

Perplexity documents two agents. PerplexityBot is built to surface and link sites in Perplexity search results, and Perplexity states it is not used to crawl content for AI foundation models. Allow it if you want to be findable.

Perplexity-User is different. It runs when a person asks Perplexity something that needs a page fetched, and Perplexity says it generally ignores robots.txt because the request is user-initiated. That has a practical consequence: a robots.txt rule will not reliably stop it, so access control for that agent belongs in your firewall. Perplexity publishes IP lists for both agents as JSON files and recommends combining user-agent matching with IP verification in a WAF rule.

How to get cited by Microsoft Copilot

Copilot is the one engine here where the vendor gives you a report. Bing Webmaster Tools has an AI Performance report covering Microsoft Copilot, AI-generated summaries in Bing, and select partner integrations. It shows total citations, average cited pages per day, page-level citation counts, and grounding queries, which are the phrases the AI used when retrieving your content. Bing added more citation views in a later update, so check what your account currently shows.

Bing's published guidance to publishers: deepen expertise on topics you're already cited for, use headings, tables and FAQ sections, back claims with examples and data, keep content accurate and current, and use IndexNow so changes are picked up faster. Bing also says it respects content-owner preferences expressed through robots.txt.

If you do one thing for Copilot, verify your site in Bing Webmaster Tools and read the grounding queries. They tell you what questions your pages are being used to answer, which no other engine here shows you.

How to get cited by Grok

Here the evidence runs out. xAI's developer documentation describes Web Search, X Search and Citations as tools, but I found no xAI page that gives site owners a crawler token, a robots.txt rule, or a citation report. Third-party sites list Grok user agents, and some describe it as unreliable about identifying itself, but those are not xAI statements, so I won't turn them into advice.

The honest guidance is the shared baseline: keep pages public, fast, and indexed in the major search engines, and watch referrers for grok.com or x.com traffic. Anyone selling a Grok-specific optimization checklist is guessing.

One robots.txt for the engines that document controls

This example allows the search and retrieval crawlers documented above and opts out of two pure training crawlers. Treat the training lines as your decision, not a recommendation.

# Search and citation crawlers: allow
User-agent: OAI-SearchBot
Allow: /

User-agent: Claude-SearchBot
Allow: /

User-agent: PerplexityBot
Allow: /

# Training crawlers: opt out (your call)
User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

Three cautions. First, a wildcard User-agent: * group with a blanket disallow catches every agent you didn't name, including Claude-User. Second, robots.txt does not govern user-initiated fetchers reliably, so Perplexity-User needs a firewall decision. Third, I left Google-Extended out on purpose, because Google's description ties it to grounding as well as training. Decide on it separately.

Some teams also publish an llms.txt file. Google explicitly says it isn't needed for its AI features, and none of the other vendors document it as a citation factor. If you want one anyway, Netalith's free llms.txt generator builds it in a minute.

How to write pages every engine can cite

This section is judgment, not vendor policy, but the vendors point in the same direction. Google advises nothing beyond good Search practice. Bing recommends clear headings, tables, FAQ sections, data, and freshness. A page that follows both is also easy for a retrieval system to quote.

  • Lead each section with its answer. Then give the caveat.
  • Make each H2 one question. Fan-out systems look for a page that answers a subtopic, not one that mentions it.
  • Carry specifics. Versions, dates, named sources and numbers separate your page from a paraphrase of someone else's.
  • Keep the answer in the HTML. Text that appears only after client-side rendering or behind a login can't be fetched.
  • Update visibly. A dated update line helps a reader and, per Bing, freshness matters.

How to measure whether AI engines cite you

Engine Where to look
Google AI Overviews / AI Mode Search Console Performance report, Web search type (mixed with normal clicks)
Microsoft Copilot Bing Webmaster Tools, AI Performance report
ChatGPT Analytics filter on utm_source=chatgpt.com
Claude, Perplexity, Grok No documented report: use server logs and referrer domains

Only two of these give you real citation data. For the rest, log analysis tells you whether the crawlers visit and which URLs get human click-throughs. Check weekly, since AI answers vary between askings and crawler changes take time to apply.

A realistic order of operations

  1. Remove blanket bot blocks in robots.txt and at the firewall; allowlist the published IP ranges.
  2. Allow the search and retrieval bots; decide on training bots separately.
  3. Confirm pages are indexed and carry no stray nosnippet or noindex rules.
  4. Verify Bing Webmaster Tools and Search Console, and read what they report.
  5. Rewrite your best pages so each section answers one question first.
  6. Review logs and reports weekly and compare against the pages you changed.

If you want help with the access and measurement steps across all of these engines, request a free quote from Netalith. Sources for this guide: Google's AI features and your website and Anthropic's crawler documentation. Perplexity, Bing and xAI details come from their own published documentation pages.

FAQ

Frequently asked questions

Do I need special markup to appear in Google AI Overviews?

No. Google says no new machine-readable files, AI text files, or markup are needed. A page must be indexed and eligible to appear in Search with a snippet to be shown as a supporting link.

Which Claude bot should I allow to be cited?

Allow Claude-SearchBot, which improves search result quality, and Claude-User, which fetches pages when a user asks. ClaudeBot is for model training, so allowing it is a separate decision.

Does Perplexity respect robots.txt?

PerplexityBot does, and it is documented as a search crawler, not a training crawler. Perplexity-User generally ignores robots.txt because it acts on a person's request, so control it in your firewall.

How can I see Copilot citations for my site?

Use the AI Performance report in Bing Webmaster Tools. It shows total citations, cited pages, page-level counts and grounding queries across Microsoft Copilot and Bing AI summaries.

Can I optimize for Grok?

I found no xAI documentation giving site owners a crawler token, robots.txt rule or citation report. Keep pages public and indexed, and watch referrer traffic rather than following unverified checklists.

Is there one strategy that works for every AI engine?

Only the baseline: keep pages crawlable and indexed, avoid blanket bot blocks, answer one question per section, and measure with the reports each vendor provides. Beyond that, citation selection is undocumented.

Stay updated with Netalith

Get coding resources, product updates, and special offers directly in your inbox.