Search for a generative engine optimization checklist and you will get the same twenty items recycled across a dozen blogs: write quality content, use structured data, build authority, be helpful. None of that is checkable. You cannot tick “build authority” on a Tuesday afternoon and know whether you did it, and a list you cannot tick is not a checklist. It is a mood board.
The second problem is provenance. Most of these lists are a 2019 technical SEO audit with the word “AI” pasted over the headings. Plenty does carry over, and we have mapped the overlap in our piece on SEO versus GEO. But a few things genuinely differ, and the recycled lists skip those or get them backwards.
What follows is the checklist we run on client sites: generative engine optimization explained as a sequence of checks rather than a set of opinions. Every item is specific enough to be either done or not done, and the stages are ordered deliberately. A model cannot cite a page it cannot read, and it will not cite a page that says nothing new.
The core position
GEO is a retrieval problem before it is a content problem. Most brands lose citations at the crawler and structure layer, long before anything they wrote gets judged on quality.
Stage 1: Can AI Crawlers Actually Read Your Page?
This is the stage most checklists treat as a footnote, and it is the one that silently disqualifies whole sites. Across a month of network traffic, Vercel and MERJ logged 569 million GPTBot requests, 370 million from Anthropic’s crawler and 24.4 million from PerplexityBot, and found no evidence that any of them rendered JavaScript. ChatGPT’s crawler downloaded JavaScript files in 11.5% of fetches and Claude’s in 23.84%, then never executed them. Gemini and AppleBot render; the rest read raw markup, so a page can rank first in Google and still be a blank shell to most generative engines.
Verify the main content is in the initial HTML response. Run curl -s against the URL and search the output for a sentence from the middle of your body copy. If it is absent, no non-rendering crawler sees it. Google’s own JavaScript SEO documentation still recommends server-side or pre-rendering, because not all bots can run JavaScript.
Confirm your CDN or WAF is not blocking them at the edge. A robots.txt allowance is irrelevant if bot management returns 403 before the request reaches origin.
Grep 30 days of server logs for every AI user agent. Record fetch count, status codes and mean response time per bot. A bot with zero hits is a finding, not missing data.
Find the 200 responses that contain nothing. Single-page applications routinely return a success code with an empty body. Crawl your own sitemap and flag any URL whose rendered text is under 200 words.
Keep an XML sitemap with honest lastmod values. Dates that change on every deploy are worse than none.
Decide on llms.txt deliberately, then move on. Google’s John Mueller has called it “purely speculative for now”, noting that no AI system uses it despite years of availability. It is a cheap hedge, never the project.
Stage 2: Structure That a Model Can Lift
Generative engines do not quote pages, they quote passages. The unit of optimisation is a self-contained block of text that survives being torn out of its context.
Answer the heading’s question in the first 40 to 60 words beneath it. No preamble. The opening sentence should be liftable verbatim.
Phrase H2s as the question a buyer types or speaks. “How Much Does a GEO Audit Cost?” beats “Pricing Considerations”.
Define the page’s subject inside the first 100 words, as “X is Y”, with no hedging clauses in front of it.
Enforce one idea per section. Sections that answer two questions get absorbed and paraphrased rather than cited.
Convert every comparison into a real HTML table. Tables are the densest extractable format on a page.
Show a visible published date, a visible updated date and a named author with stated credentials, in the body text as well as the schema.
Delete every internal reference. Search the draft for “as mentioned above” and “in the previous section”: each one breaks a passage that would otherwise stand alone.
Stage 3: Entity Consistency and Third-Party Authority
Models assess the aggregate of everywhere you are described, not your website alone. Semrush’s 13-week study of 230,000 prompts and more than 100 million citations found Reddit in roughly 60% of ChatGPT responses and Wikipedia in roughly 55% over the first half of the period, with LinkedIn leading in Google AI Mode.
Write one canonical 40-word brand description and use it verbatim everywhere: About page, LinkedIn, Crunchbase, G2, directories, press boilerplate. Variation dilutes the association.
Audit your name for ambiguity. If a model cannot disambiguate you, it defaults to the better-corroborated entity.
Publish an About page stating founding year, headquarters, team size, named leadership and exactly what you sell. Those five facts get extracted when a model is asked who you are.
Add Organization schema on the homepage with a complete sameAs array pointing at every profile you control, and Article schema with author, datePublished and dateModified on every editorial page.
Use FAQPage markup only where the questions and answers are visible. Google’s structured data policies prohibit marking up content readers cannot see, and incomplete markup is ineligible anyway.
Claim and complete profiles on the platforms cited in your category. Run five buyer prompts, list every domain the answers cite, and work that list instead of a generic directory list.
Build toward reference-class sources. Wikipedia-grade coverage needs independent, substantial press, so start that work well before you need it.
Stage 4: Original Information Gain
Generative engines synthesise what already exists, which makes commodity content worthless to them and original data disproportionately valuable. If a fact exists in one place only, an answer containing it has to attribute it.
Publish one number per quarter that nobody else has. Aggregate your own account data, survey your customers, or benchmark something measurable in your category.
Name your methodology and give it a stable label. A named framework becomes an entity models attach to you.
Write key statistics as complete, self-contained sentences. “Median onboarding time across 214 accounts was 11 days” travels. “Onboarding was faster” does not.
State sample size, date range and collection method next to every figure. Unsourced numbers get dropped in favour of sourced ones.
Test first-hand and show the workings. Screenshots, raw outputs and failed attempts are evidence a competitor cannot paraphrase.
Version your data pages and refresh them on a fixed schedule, so the newest figure in your category is always yours.
Stage 5: Measurement You Can Actually Run
Most GEO visibility produces no click. The Pew Research Center tracked 900 US adults across 68,879 Google searches in March 2025: they clicked a traditional result in 8% of visits where an AI summary appeared, against 15% where it did not, and clicked a link inside the summary in just 1% of visits. Measure GEO by referral sessions alone and it will look broken while it is working.
Define a fixed basket of 30 to 50 buyer-realistic prompts and run it monthly across ChatGPT, Perplexity, Claude, Gemini and Google AI Mode. Fixed wording, fixed schedule, logged results.
Record four fields per prompt: mentioned, cited with a link, described accurately, and which competitors appeared instead of you.
Create an analytics segment for AI referrers covering chatgpt.com, perplexity.ai, claude.ai, gemini.google.com and copilot.microsoft.com, reported separately from organic.
Add a free-text “How did you hear about us?” field to your primary form. Self-reported attribution is the only place assistant-driven demand shows up honestly.
Track branded search volume as a proxy. Rising branded queries against flat rankings usually means an answer engine is recommending you without a click.
The checklist above is a quarter of work. Compressed into one day, the order is not negotiable.
Curl your ten most commercially important URLs and confirm the body copy is in the raw HTML. Everything downstream is wasted if it is not, which is why our engineering-led approach to generative engine optimisation starts with the response and not the copy.
Read your robots.txt and your CDN bot rules, and unblock the retrieval crawlers you want. Fifteen minutes, and the most common reason a well-optimised site is absent from AI answers.
Run twenty buyer prompts by hand and write down what you see. You cannot improve a citation share you have never measured, and the gaps will reorder your roadmap.
Everything else compounds slowly. These three are binary, and two of them are engineering tickets rather than content briefs. That split is why we run SEO and GEO as a single programme inside our Search engine: the fixes that make you citable are mostly the fixes that make you rank, and they sit with your developers.
Maya leads the search practice at Gyrodile: technical SEO, content strategy and AI search visibility. She writes about what changes when the answer engine replaces the results page.