Let's talk
security

The production guard was a line one flag deletes

A login endpoint with no credential was fenced off by an assert. One optimisation flag strips every assert from the build, and the fence with it. Caught in review.

· · updated

A laptop on an operations desk; beside it the children's learning app's module editor with a module's text and its source material.

Somewhere in your codebase is a development shortcut that logs a test user in without a password, so that tests can run as a real account. The question a CTO has to ask is what stops that shortcut existing in production.

In this server, it was a check that refused to run outside development, and it was written as an assertion. The runtime the server is written in has an optimised mode that removes every assertion from the compiled output. Under that flag the guard is not a check that passes. It is not there. A credential-free login would have been open on a production server, protected by a line that a deployment flag deletes.

It was found in review, before it shipped, so no customer was exposed. No test could ever have noticed, because tests do not run under that flag.

What was actually going on

An assertion reads exactly like a guard. It states a condition that must hold and stops the program when it does not. But an assertion is a claim the author believes is already true, there to catch the author being wrong during development. A guard is a rule enforced against the world, including against a caller acting in bad faith, and the world does not stop at your build configuration.

If a check protects against anything other than your own mistakes, it must be a statement the runtime is not permitted to remove. The same applies to debug-only logging that is also the audit trail, a validation that only runs in a development middleware stack, and a rate limit in a proxy configuration that a fallback route bypasses. In each case the protection has a shorter lifetime than the thing it protects.

What we changed

The fix is one word: an explicit conditional raise instead of an assertion. It was proved by running the suite under the optimisation flag and watching the guard still fire.

The test applied to every safety check since is to name the configuration under which it does not execute. If none can be named, good. If one can, that configuration is a supported way to run the software without the check.

Every regression test written afterwards was watched failing before it was trusted: write the test, remove the fix, see it go red, put the fix back. It costs about ninety seconds. It paid off twice. A test for a data migration turned out to pass against an implementation that committed each row separately, which is the property it claimed to rule out. And a screen crashed at runtime with a missing function while the whole suite and the type checker were clean, because a helper had been added to a module and re-exported through an index file moments later. The test that covers it reads both files as text and compares them. It is ugly, and it was confirmed to fail on the exact real defect before being kept.

What it did not fix

Two other checks from the same period could not fire for structural reasons. One test module inserted stub modules before importing the code under test, so if the real model module failed to import, the test still passed. A test that supplies its own subject is testing the stand-in.

And fifty-one browser tests passed over a portal in a children’s learning app that was entirely read-only, with no write endpoint wired and no button doing anything. Every test was correct: a screen renders with the right data. Rendering is not working. A separate suite of three tests now encodes what working means: create a task, add a reward, approve a submission, and confirm the child’s balance rose by exactly the expected amount after a reload. Those three caught things fifty-one could not.

None of this is a bug in the ordinary sense. Each check exists, is correct, and cannot produce the outcome it is relied on for. There is no automatic way to find the next one. The question is slow the first time and quick afterwards.

What to ask your own team or supplier

  • Which safety checks in production code are assertions, and what does our production build do with assertions?
  • For each guard, under what configuration does it not execute?
  • Is there a development login or test endpoint, and what, other than one line, keeps it out of production?
  • Were the regression tests seen to fail before the fix was applied?
  • Do any tests replace their own dependencies with stand-ins, and could a broken import go unnoticed?

Where this ends up

Asking of every guard under what circumstances it fails, and whether anyone has seen it do so, is the gate described in AI-first delivery. It is reading rather than scanning, and it is where a check that cannot fail gets found.

Working on something like this?

We build this kind of software, and we staff the teams that do.

Get in touch