Let's talk

Topic

Search & GEO

How answer engines read a site, and what makes content citable.

6 articles

About this topic

Search is being replaced at the top of the funnel by systems that read pages and answer the question directly, and the two audiences want different things. A search engine can rank a page that says very little. An answer engine has to find a passage specific enough to quote and attribute, which stands up out of context — a much harder bar, and a much better one.

This site is the working example, including every part of it that went wrong. It was rebuilt from WordPress onto static HTML and cut over in a day, and the fortnight that followed produced more usable evidence about how engines actually behave than any amount of reading would have.

The clearest lesson was that crawling and indexing are not the same event, and neither is the same as being citable. In the first day after cutover the access log recorded roughly five hundred requests from ClaudeBot, four hundred from GPTBot and five hundred from bingbot — but only forty-six from Googlebot, and Google had indexed five of twelve sampled pages. A week of heavy AI crawling and a domain Google was still making its mind up about, at the same time, on the same site.

The second lesson was that a migration is judged on its redirects and nothing else. Bing kept returning a pre-cutover page with old marketing copy for days. The cause was not the content: a hundred and forty legacy URLs had been mapped with a trailing slash, engines held some of them without one, and nginx only adds a slash when a real directory exists — which a retired WordPress path does not. So the crawler asked for the old address, received a 404, learned nothing, and kept the stale entry. The access log shows the moment it was fixed: eighteen 404s on one URL, then ten 301s.

The third was that the most dangerous failures are silent and self-inflicted. Every build produced during the rebuild carried noindex on every page, because the build defaulted to the staging hostname and the noindex flag was derived from it. Deploying one of those would have removed the entire site from search, and nothing would have complained. A `robots.txt.dev-bak` file containing `Disallow: /` was sitting in the public directory for the same reason.

The fourth is the one that matters most for GEO, and it was visible in a single afternoon. Asked what the company does, Perplexity answered from the new pages and, on a competitive vendor question, listed the product with accurate operational detail. ChatGPT, with web search on, described the business the site had stopped being, and said plainly that it did not have enough public evidence to identify the platform. Same site, same day, same crawl access. The difference was whether there was anything specific enough on the page to lift.

What follows from all of that is unglamorous and mostly editorial. Say the specific thing rather than the marketable one. State limitations, because qualifying statements are what get quoted. Answer a question somebody would actually type. Serve HTML a crawler can read without executing JavaScript. And check the plumbing first, because none of the writing matters if the redirect returns 404 or the page ships with noindex.

An interior wall painted halfway, new pale paint beside old blue, with stepladders in front.

ChatGPT described the company we used to be

After a site move, one AI assistant described our product accurately and another described a business we had stopped being. What a page must contain to be quoted.

All articles