Website tools

GEO & AI Crawler Audit

Audit AI crawler rules, llms.txt, agent files, structured data, FAQs, semantic HTML, and agent accessibility from source and rendered pages.

One URL. Crawler policy and page evidence.

Root crawler files are checked at the page origin. JavaScript evidence loads in a separate browser pass.

previewRENDERED
Audit preview
Live audit Crawler scan

How it works

Three simple steps. No account or installation required.

  1. Step 1

    Enter your URL

    Enter one public page URL.

  2. Step 2

    We run the audit

    Review source checks while the rendered-page pass completes.

  3. Step 3

    Review and export

    Inspect exact crawler rules, evidence, and prioritized actions or export the report.

More about this tool

Overview

GEO & AI Crawler Audit inspects one public page and its origin-level crawler controls. It separates search and citation crawlers from user-triggered fetchers and training controls, shows the exact robots.txt rules that apply, and compares server HTML with a rendered browser view for JavaScript-heavy sites. It also checks experimental AI files, structured data, visible FAQ patterns, entity and attribution signals, semantic HTML, and agent-relevant accessibility. Reports are factual snapshots without a GEO score, ranking prediction, stored history, or AI-generated analysis.

How to read the audit

The overview summarizes crawler access, discovery files, page signals, and rendered evidence. Expand any agent to see the exact matching group, directives, rule lines, and effective result for the submitted path.

Search & citation

Agents used to discover or cite public pages in search-backed AI experiences.

User-triggered fetch

Requests made for a user action. Some providers state that robots.txt may not apply.

Training & data

Publisher controls for model development or datasets. Allowing them is not treated as a ranking factor.

Standards and evidence

Robots matching follows RFC 9309 semantics. Crawler names, purposes, and policy caveats are maintained from vendor documentation. Experimental files are labeled separately from established crawler controls.

What the audit cannot prove

A published robots rule is a request to the named operator, not an access-control system. This tool does not send traffic from vendor IP ranges, verify server logs, bypass bot protection, certify accessibility, or predict whether a page will be cited or ranked.

Source HTML versus rendered DOM

Server HTML shows what a non-rendering crawler receives immediately. The rendered pass executes public page JavaScript in an isolated browser and rechecks FAQs, schema, semantic structure, and control labels. Large differences are useful evidence for JavaScript-heavy SaaS and CMS sites.

Frequently asked questions

What does the GEO and AI Crawler Audit inspect?

It reads robots.txt rules for major AI search, user-fetch, and training agents; checks llms.txt, llms-full.txt, agents.md, and AGENTS.md; and inspects the submitted page for indexability, FAQs, structured data, semantic HTML, attribution, and agent-relevant accessibility.

Does allowing GPTBot improve ChatGPT search visibility?

Not by itself. GPTBot is a training crawler, while OAI-SearchBot is the crawler associated with ChatGPT search visibility. The report keeps search, user-fetch, and training controls separate so one cannot be mistaken for another.

Does the audit run JavaScript?

Yes, when the isolated browser worker is available. Source HTML results appear first, then the report compares them with the rendered DOM. The source report remains usable if the browser pass is unavailable.

Is llms.txt required for AI search visibility?

No. llms.txt is an emerging convention, not a universal standard or ranking requirement. The audit reports its presence and basic Markdown structure as optional evidence.

Why are agents.md and AGENTS.md shown separately?

They are different conventions. Lowercase agents.md is used by some agentic storefront systems, while uppercase AGENTS.md is commonly repository guidance for coding agents. Neither is treated as a universal GEO requirement.

Is this a complete accessibility or WCAG audit?

No. It checks observable signals that also help browser agents, such as labels, accessible names, alternative text, landmarks, native controls, and document language. Manual accessibility testing is still required.

Does the audit crawl the whole website?

No. It inspects the submitted page, root crawler files, and up to three declared or conventional sitemap endpoints. It does not crawl sitemap entries or internal pages.

Are reports or submitted URLs stored?

No report history is stored. The URL and fetched content are processed for the current request, and exports are created locally in your browser.

Technical GEO evidence · Free, no account required