Layer one: can AI crawlers reach the site
Five checks. robots.txt has no group blocking OAI-SearchBot, ClaudeBot, PerplexityBot, or Bingbot on the pages you want cited. The CDN or WAF is not blocking AI user agents; Cloudflare in particular can block them by default. Fetching a key page with each AI user agent returns 200 and full HTML, not a challenge page. robots.txt itself returns 200 and parses. The pages are not behind authentication, geoblocking, or interstitials. SerionFlow's free AI crawler access checker and robots.txt checker cover the first four.
Layer two: is the content in the static HTML
Five checks. Fetch the page with JavaScript disabled or with curl and count words; the number should be close to what a browser shows. The H1 and every H2 are present in the raw HTML. FAQ questions and answers are in the raw HTML, not injected on click. Tables render as table elements, not as images or client-rendered grids. The canonical, title, and meta description in the raw HTML match what the browser shows. A single-page app that renders client-side will fail all five, and no amount of content work fixes that until rendering does.
| Layer | Checks | Pass criterion | Free tool |
|---|---|---|---|
| 1. Crawler access | 5 | Every AI agent gets 200 and full HTML on key pages | AI crawler access checker; robots.txt checker |
| 2. Static HTML | 5 | Raw HTML word count within ~20% of rendered; headings, FAQs, tables present | curl or browser with JS off |
| 3. Answer structure | 6 | Direct answer in first paragraph; H1 matches prompt; table; FAQ; honest negative | Manual review |
| 4. Signals | 5 | Valid schema matching visible content; accurate dates; self-canonical; sitemap listing; llms.txt present | Schema validator; sitemap checker; canonical checker |
| 5. Measurement | 4 | Fixed prompt set; monthly re-run; citation share per engine; Search Console conversational queries tracked | Citation gap finder; Search Console |
Layer three: can an engine lift an answer from the page
Six checks per priority page. The H1 is the prompt or the exact query. The first paragraph answers it directly in forty to sixty words with no preamble. There is a TL;DR or summary list whose items stand alone. At least one table carries real numbers. There is an honest section on limits or when the product is the wrong choice. There is an FAQ with verbatim questions and short answers. Pages that pass layers one and two but fail here are readable and still rarely cited, because engines prefer sources that state the answer plainly.
- H1 equals the prompt.
- Answer in the first paragraph.
- Standalone TL;DR bullets.
- Table, honest negative, FAQ.
Layer four: are the signals accurate
Five checks. JSON-LD validates and every marked-up element, especially FAQPage questions, is visible on the page. datePublished and dateModified reflect reality and the page has been updated within the last year for evergreen topics. The canonical is self-referencing. The page is in a sitemap that Search Console and Bing Webmaster Tools have accepted. An llms.txt exists at the root and lists the page if it is a key page. None of these earn citations on their own; wrong ones reduce trust.
Layer five: are you measuring
Four checks. A fixed prompt set of twenty to fifty buyer questions exists. It has been run in ChatGPT with search, Perplexity, and Google AI Overviews with every cited URL logged. It is re-run monthly on the same prompts. Search Console impressions and positions for the corresponding conversational queries are tracked alongside. Without these, an audit has no baseline and no way to show that fixes worked. SerionFlow's free citation gap finder and GEO audit tool provide a starting structure; paid plans add sampled AI visibility monitoring and AI-related query signals.
How to run the audit and what to do with the result
Audit ten priority pages rather than the whole site; the pattern generalises. Record pass or fail per check, fix layer one and two failures site-wide first because they are usually one configuration change, then rewrite the priority pages against layer three, correct signals in layer four, and establish measurement in layer five. Re-audit quarterly. When SerionFlow published its own audit in September 2026, every page failed layer two because content rendered client-side, which explained zero AI citations more completely than any content finding; fixing the substrate came before writing a single new page.
Step by step
- 01
Pick ten priority pages
Choose the pages that target your most valuable prompts, including at least one comparison and one how-to.
- 02
Test crawler access
Fetch each page as OAI-SearchBot, ClaudeBot, PerplexityBot, and Bingbot and confirm 200 with full HTML; check robots.txt and CDN settings.
- 03
Check the raw HTML
Fetch with JavaScript off and confirm word count, headings, FAQs, tables, and metadata are present in the source.
- 04
Score answer structure
For each page, check H1 equals prompt, first-paragraph answer, TL;DR, table, honest negative, and FAQ.
- 05
Validate signals
Run schema, canonical, and sitemap checks and confirm dates are accurate and llms.txt exists.
- 06
Establish measurement
Freeze a prompt set, run it across engines, log citations, and schedule a monthly re-run alongside Search Console tracking.
Clear answers
Frequently asked questions
What is the most common reason a site gets no AI citations?
+
Layer one or two failures: AI crawlers blocked at the CDN or in robots.txt, or page content that only exists after JavaScript runs. Both are invisible in a browser and both make every content improvement irrelevant until fixed. Check them first.
How do I check what an AI crawler sees on my page?
+
Fetch the page with curl or a browser with JavaScript disabled and read the HTML. If the word count is a fraction of what the browser shows, or headings and FAQs are missing, AI crawlers see the smaller version. A free AI crawler access checker confirms the fetch succeeds for each agent.
Does schema markup help with AI citations?
+
Accurate schema helps engines understand a page and does no harm; inaccurate schema, such as FAQPage markup for questions not visible on the page, reduces trust. It is a layer-four signal: worth getting right after crawler access, static HTML, and answer structure are fixed.
How often should I run a GEO audit?
+
Quarterly for the full checklist on priority pages, and immediately after any change to hosting, CDN, robots.txt, or rendering. Measurement in layer five runs monthly regardless, because the prompt re-run is what shows whether fixes worked.
Is there a free GEO audit tool?
+
SerionFlow offers a free GEO audit tool, AI crawler access checker, robots.txt checker, schema validator, sitemap checker, and citation gap finder that together cover most of the checklist without an account. The answer-structure layer is a manual review of the page text.
Continue exploring
Related SerionFlow resources
More in Guides
- GEO services for SaaS: what to buy, from whom, and what to expect
- How to check a sitemap for errors
- How to fix keyword cannibalization
- How to get cited by ChatGPT, Perplexity, and Google AI Overviews
- How to test and validate a robots.txt file
- Meta description length and best practices
- Programmatic SEO on Next.js and Vercel
- Programmatic SEO on Shopify: what works and what does not
Make the next move obvious
Let SerionFlow turn your market evidence into momentum.
Confirm the market, rank the opportunities that matter, create brand matched pages on your domain, and keep moving with controlled weekly intelligence.