Let's talk
engineering

Seventeen passing layout tests that could never fail

Seventeen layout tests passed against a stylesheet that made the assertion impossible to fail. Fifty-one screen tests passed on a portal that was read-only.

· · updated

A straight steel bar lying across a wooden plank on a workshop bench.

Your supplier reports that every test passed. The report is accurate. The tests did pass. But some of them could not have failed, and a release decision made on that report rests on nothing.

A responsive-layout suite asserted that the page’s scroll width never exceeded the viewport width. Seventeen screens, seventeen passes, across every breakpoint. The stylesheet set horizontal overflow to hidden on the root elements, a common backstop that stops stray content producing a horizontal scrollbar. That declaration makes the measured scroll width impossible to exceed the viewport width, so the assertion could never fail on any page in that application.

The cost of this kind of mistake is not the test. It is the false assurance: nobody looks at layout because the layout tests are green.

What was actually going on

In every case below the check was present, ran routinely, and was incapable of producing the negative result it existed to detect.

Fifty-one screen tests passed against a read-only portal. Every screen rendered and every assertion about content held. None of them exercised a write. The note in the log is exact: rendering is not working.

A deploy verification that could not fail. The deploy script requested the application immediately after restarting the process manager, inside the restart window, and reported the empty response. It did that on every healthy deploy. It was noted roughly fifteen times over seven months before anyone fixed it, because a step that always says the same thing becomes background noise. A second deploy script had the same defect and compared against a fallback page instead of the application, reporting a six-kilobyte placeholder as though it were the three-hundred-kilobyte bundle.

A build reported as successful when the success belonged to something else. The build was piped through another command, and the exit code that was checked belonged to the last stage of the pipeline rather than the compiler.

A page that could not be probed. The application is a single-page app, so the server returns a successful response for any path. Requesting a URL proved nothing about whether the page had been deployed. It went unnoticed until an external reviewer rejected the store listing for a missing page.

The same pattern appeared outside testing: a security plugin writing access rules to a file the web server never reads, a certificate renewal cron entry that disables itself, a noise-reduction filter whose strength setting did nothing, and a silence threshold carried over from another recording environment that sat below the noise floor. The last two were caught by checking that changing the input changed the output.

What we changed

The layout test was rewritten to measure actual geometry: find any element whose right edge falls beyond the viewport with no scrolling ancestor to contain it, and name it. It was then proved non-blind by injecting an oversized element into a narrow page and confirming it fired and named the offender.

The follow-up shows how easily this goes wrong twice. A test asking whether a table was stacked on mobile checked that the header height was at most one pixel. A correctly stacked header measures two. The test reported a correct screen as broken.

Deploy verification now searches the served bundle for text belonging to the specific page. Other habits followed:

  • Every regression test is watched failing before it is trusted. A security fix in one system was checked by running its new test against the unfixed code, which produced fourteen failures, three of them the specific leak being fixed.
  • The artefact is read, not the configuration. Two mobile permissions shipped because a runbook’s description of the manifest was trusted instead of the built manifest being read.
  • A cleanup verified by searching for six known phrases missed a seventh, worded differently, which shipped. The check now matches claim-shaped patterns and was confirmed to fire on the exact string that got through.
  • A completion report is treated as a claim. In one case the real number was three suites short of what was reported.

What it did not fix

A test that has passed since the day it was written is either covering something stable or covering nothing, and there is no automatic way to tell which. The habit only works on the checks somebody thinks to question. The phrase-list miss and the mobile permissions both reached a release before the habit existed.

What to ask your own team or supplier

  • When did each of our key tests last fail, and for the reason it exists?
  • Do the tests exercise writing and saving, or only that screens draw?
  • Does a deploy check prove that the new version is running, or only that something answered?
  • Was the new regression test run against the unfixed code and seen to fail?
  • When a report says all tests pass, who re-ran the gates against the code we are about to accept?

Where this ends up

The question under all of this is when a check last told you something you did not already believe. It is one of the gates on AI-first delivery, where verification is a separate act from production and the evidence for a piece of work is the running system rather than the report saying it passed.

Working on something like this?

We build this kind of software, and we staff the teams that do.

Get in touch