What we check, and why
Twenty-two checks, two scores. Here is every one of them, what it looks for, and why it changes whether search engines and AI assistants can find you, read you and quote you.
SEO, AEO and GEO
Three acronyms for three different jobs. Most sites are built for the first and have never been looked at for the other two.
SEO
Search engine optimisation
Being found in a list of links. Someone searches, you appear, they click. This is the job most websites were built for.
AEO
Answer engine optimisation
Being the answer rather than one of ten links. The question gets answered on the spot, and only one or two sources get named.
GEO
Generative engine optimisation
Being quoted accurately when an assistant writes prose about your subject. You are not a link here - you are a fact inside somebody else's sentence.
The uncomfortable part: you can rank perfectly well and still be invisible to all three of ChatGPT, Claude and Perplexity, and nothing in your analytics will tell you.
How a page becomes an answer
Four stages. A site can fail at any one of them, and everything downstream stops.
Stage two is why crawler access carries more points than anything else we measure. If the gate is shut, stages three and four never happen, however good the page is.
AI Readiness — 11 checks, 105 points
Whether AI assistants can reach your site, read it, and trust it enough to quote. Weighted by consequence, not by how easy something is to fix.
Where the points sit
The two in yellow are the ones that make everything else moot if they fail.
-
Can AI crawlers get in?
20What we look for. Your robots.txt, checked against the AI crawlers by name - GPTBot, ClaudeBot, PerplexityBot, Google-Extended and nine others.
Why it matters. This is the gate, which is why it carries the most points. Block these and nothing else on this page matters: ChatGPT, Claude and Perplexity simply never see your site. Plenty of sites block them by accident, copying a robots.txt from somewhere else.
-
Are your words actually in the page?
15What we look for. How much of your visible text is in the HTML we receive, versus only appearing after JavaScript runs.
Why it matters. Most AI crawlers don't run scripts. If your copy is assembled in the browser, they see an empty shell. A beautiful site can be completely invisible this way, and nothing in your analytics will tell you.
-
Machine-readable business details
15What we look for. Structured data for your organisation and site, whether the types suit what you actually are, and whether any of it is broken.
Why it matters. This is how you state facts - who you are, what you sell, your address - in a form a machine cannot misread. Broken markup scores worse than none, because a tool that tries to read it gets an error instead.
-
Do you get to the point?
10What we look for. Where the substance sits on the page: near the top, or after a long run-up.
Why it matters. Assistants quote the part that answers the question. Pages that open with scene-setting give them nothing to lift, so they quote a competitor who answered in the first sentence.
-
Are you visibly a real business?
10What we look for. Contact details in any country's format, named people, and links to profiles that can be checked independently - a company register, LinkedIn, a maps listing, Trustpilot or Yelp.
Why it matters. Assistants are cautious about recommending businesses they can't verify. Address, email and an outside profile move you from 'a website' to 'a company that exists'.
-
Do you answer real questions?
8What we look for. Content shaped as questions and answers, in headings and structured data.
Why it matters. The single most quotable format there is. A question-shaped heading matches a question-shaped prompt, which is how most people talk to an assistant.
-
Is it laid out to be lifted?
6What we look for. Real lists and tables where the content suits them, rather than the same information buried in prose.
Why it matters. Machines extract from structure far more reliably than from paragraphs. A price list as a table gets quoted accurately; the same prices in a sentence get mangled or skipped.
-
Can anyone tell it's current?
6What we look for. Machine-readable dates on your pages, and whether the copyright line has been left behind.
Why it matters. Given two equally good sources, an assistant prefers the one it can tell is maintained. An undated page looks like it might be from 2016.
-
An llms.txt summary
5What we look for. A file at /llms.txt handing AI tools a plain summary of the site, and whether it's signposted.
Why it matters. Newer and not yet universal, so it's worth only a few points - but it lets you write the description of your business rather than leaving an assistant to infer one.
-
Do your pages point at each other?
5What we look for. How many links each page makes to other pages on the same site.
Why it matters. Links are how a machine works out which of your pages matter and how your topics relate. A page nothing links to reads as an afterthought.
-
Is your language declared and consistent?
5What we look for. Whether the page declares a language, and whether the spelling matches what it declared.
Why it matters. It costs nothing and removes a guess. We check consistency, not nationality - American English on a site that declares American English is exactly right.
Site Health — 11 checks, 87 points
The technical foundations. Less glamorous than the AI half, and the reason a good site quietly underperforms for years.
-
Are you accidentally invisible?
15What we look for. noindex tags and robots.txt rules that hide pages from search engines.
Why it matters. The most expensive mistake on this list, and the most common after a site launch - a staging setting shipped to production hides the whole site, and everything else becomes moot.
-
Is the whole site encrypted?
10What we look for. Certificate validity, and whether anything still loads over plain HTTP.
Why it matters. Browsers label unencrypted pages 'Not Secure', which turns visitors away before they read a word. It is also a long-standing ranking signal.
-
Can everyone read it?
10What we look for. Written descriptions on images, heading order, and a declared page language.
Why it matters. The same things that help a screen reader help a crawler: both are reading your page without seeing it. Accessibility and machine readability are mostly the same work.
-
Is there a usable sitemap?
8What we look for. Whether a sitemap exists, and whether the URLs in it are real and current.
Why it matters. A tidy list of your pages makes discovery faster and more complete. A sitemap full of dead URLs is worse than none - it teaches crawlers to distrust it.
-
Does the server respond promptly?
8What we look for. Response timing, page weight and redirect overhead, measured at the moment of the scan.
Why it matters. Honest proxy signals, not a lab measurement - we say so on the report. Slow responses cost you rankings and cost you visitors who don't wait.
-
Does each page introduce itself?
8What we look for. Titles and descriptions: present, unique to the page, and a sensible length.
Why it matters. Your title is the headline in every search result and often the label an assistant gives you. Duplicated titles make pages compete with each other instead of ranking.
-
Does it work on a phone?
8What we look for. Viewport declaration, tap target spacing and layout signals in the markup.
Why it matters. Google indexes the mobile version of your site, not the desktop one. If the phone version is the poor relation, that's the version being judged.
-
How deep is everything buried?
6What we look for. How many clicks from the homepage each scanned page sits at.
Why it matters. Pages buried five levels down get crawled rarely and rank poorly. Depth is a decent proxy for what a site treats as important.
-
Do pages load directly?
5What we look for. How many forwarding hops sit between the URL asked for and the page returned.
Why it matters. Every hop costs time and leaks a little ranking signal. Chains build up quietly over years of site changes and nobody notices.
-
Are canonical tags pointing sensibly?
5What we look for. Canonical tags, and in particular any pointing at a different domain.
Why it matters. A canonical aimed at someone else's site hands them the credit for your page. It is nearly always a copy-paste accident, and an expensive one.
-
Is there substance on the page?
4What we look for. Word counts and near-duplicate pages across the site.
Why it matters. Thin and duplicated pages give neither a visitor nor an assistant anything to take away, and dilute the pages that do.
Which assistants this affects
When we check whether crawlers are allowed in, these are the names we look for in your robots.txt. Blocking any of them removes you from the tool it feeds.
-
GPTBotChatGPT, OpenAI -
ClaudeBotClaude, Anthropic -
PerplexityBotPerplexity -
Google-ExtendedGoogle AI Overviews and Gemini -
Applebot-ExtendedApple Intelligence -
AmazonbotAlexa -
meta-externalagentMeta AI -
CCBotCommon Crawl, feeds many models -
cohere-aiCohere -
BytespiderByteDance -
Claude-WebClaude, browsing -
anthropic-aiAnthropic, legacy name
Microsoft Copilot reads from Bing's ordinary index, so it follows your normal search settings rather than a separate crawler - which is covered by the Site Health indexability check.
How the scoring works
Each check returns green, amber or red and earns all, some or almost none of its points. Your percentage is what you earned out of what was available. Findings are ordered worst-first and heaviest-first, so the top of your report is always the thing most worth doing.
Every finding gives you the plain-English version first and the technical detail underneath, always visible, never behind a click. If you want to see it working, run a scan - or read the questions people ask before they do.