Every wrong address on the site showed the homepage
Two of our sites answered every mistyped address with the homepage and reported success. To a search engine that is thousands of identical pages, none marked real.
You type your own website’s address with a letter missing, or follow a link from an old email to a page you moved last year, and instead of “page not found” you get the homepage. It feels like a kindness. Nobody is left staring at an error. So nobody in the office ever mentions it.
What a visitor sees and what a search engine sees are different things. When a search engine fetches a page it also reads the answer the site gives about whether the page exists. This site answered “yes, here it is” to every address anybody could invent, and served the same homepage under each one. From the outside that is not one site with a friendly error page. It is a site with an unlimited number of identical pages and no way of telling which ones are real.
The two files a search engine reads first, the one that says which parts of the site to visit and the one that lists the pages, were affected in the same way. Asked for either, the site returned the homepage’s contents. So from the day it went up the site had been telling every crawler nothing useful, in a format the crawler could not read.
We found this on the site for our quotation product, while deploying a new version of it. We had found it two weeks earlier on the landing page for a children’s learning app we build. Same fault, same single line, two sites. There is no number for what it cost, because nothing measured it before the fix, and we are not going to invent one. What we know is what the site had been saying to the outside world, and that it was wrong.
What was actually going on
Some web applications are built as a single page that the browser assembles, and for those the web server is told: if the address does not match a file, serve the main page and let the application sort out what to show. That is correct for a signed-in application, where every address is a screen inside the app. It is wrong for a marketing site made of separate pages, because there is nothing to sort out. An address that matches no page is simply a page that does not exist.
Both sites had been set up with the single-page rule, one line of server configuration, applied to a site that was not a single-page application. Every unknown address fell through to the homepage with the status code that means success. The learning app’s landing page had the extra symptom that a link shared into a chat rendered as a bare grey box, because the page carried no preview image.
What we changed
On the quotation product’s site, the server now serves the multi-page build and answers an unknown address with a real “not found” status and a real not-found page. Files that carry a fingerprint in their name are cached hard, because they can never change under that name; pages are not cached at all, because they can. Exactly one path is still passed through to the application, the one that receives enquiries, and the product film from July’s outreach was kept at its old address so those links still work. We verified fourteen addresses live, including the not-found case and the enquiry path end to end, with zero failed requests, and submitted nineteen real addresses to the IndexNow service, which accepted them.
On the learning app’s landing page, unknown addresses now return not-found against a proper page, while the signed-in portal keeps its own fallback because it genuinely is a single-page application. It gained a real robots file that excludes the portal and the API, a sitemap of its three public pages, a preview image and card for shared links, and structured data describing the application, the organisation and the site. We caught our own mistake in the sitemap’s namespace before it shipped.
What it did not fix
Nothing measured the before and after, so we cannot say how many duplicate addresses a search engine had already recorded or how long it takes to forget them. Acceptance from IndexNow means the addresses were received, not that anything was indexed. And the signed-in portal still returns its shell for any address under its prefix, which is correct, so what keeps a crawler out of it is the robots file rather than a not-found answer, and a robots file is a request, not a lock.
The mechanism, in plain words
The line was nginx’s try_files $uri $uri/ /index.html. It says: look for a file at this path,
then a directory, then give up and serve the index. On a single-page application that final step is
the application. On a static site it is a soft 404: an HTTP 200 with the homepage’s HTML for every
path that does not exist, including /robots.txt and /sitemap.xml, which is why both were
returning HTML instead of the plain text and XML a crawler expects. A crawler can mint distinct
URLs without limit and every one of them answers 200 with identical content.
The rule is that an unknown path must return 404, and a single-page fallback belongs only under the prefix where the signed-in application lives. Everything else follows from that: robots and sitemap become their own files with their own content types, fingerprinted assets can be cached for a year, and the not-found page is a page rather than an accident.
Where this ends up
Sazinga Quote is the product whose site this was. A site that answers every question with “yes” has told a search engine nothing, and the fix was one line and a day of checking that the line was right.