Let's talk
engineering

A failing test run that said nothing about the code

Three test suites failed on code nobody had changed. Leaked test users had filled a shared tenant to its seat limit, and a red run meant nothing about the release.

· · updated

A hallway coat rack with every hook full of coats and shoes lined up beneath it.

Your team tells you the release is red. Three test suites failed, and someone has to decide whether the release week slips. Then it turns out that nothing in the release was wrong. The suites failed because the shared test environment had filled up.

That is what happened here. Three suites in unrelated modules failed on code that had not changed, and the error was a seat limit: the test tenant had reached its maximum number of members. Three separate investigations started in the wrong module before anyone looked at the environment. Each one cost engineering time, and each failure was reported first as a bug in whatever module noticed it.

The risk for you is wider than one wasted morning. A team that learns its red runs are sometimes meaningless starts re-running them until they go green. At that point the suite has stopped telling you anything about the release.

What was actually going on

Two other suites were the cause. Both needed a second user to test something. One checked that a notification reaches somebody other than the person who acted. The other checked that a permission is refused to a role that lacks it. Both invited a member and never removed it, so every run left one more behind.

Nineteen had accumulated, which was enough to reach a seat limit of twenty. That limit is a real product rule, correctly enforced, and unrelated to anything the failing suites were testing.

The same tenant produced three more incidents in the same period, all of the same shape. A preference controlling whether inventory is tracked had been switched off by other work in flight, and every stock-receipt suite failed on an unrelated rule. A product soft-deleted by another process left a price-list entry pointing at nothing, so a browser test that picked the newest entry went red mid-session. And when the tenant’s trial period expired, most of the suite went red at once.

Not one of those was a code regression. A shared test tenant is global state that every suite reads and several write, so it fails like global state. The symptom appears far from the cause, the result depends on the order suites run in, and leaks stay below the threshold of notice until they arrive all at once.

What we changed

Purging the leaked members took a minute. The fix that mattered was adding seat release to both suites, so that each leaves the tenant with the same number of members it found.

The rule we now hold every suite to is idempotence: run it twice in a row with nobody resetting anything, and get the same result. It is easy to check and it catches leaks nobody listed.

One suite deliberately expires the tenant’s trial to check that the lifecycle gate refuses writes. It restores the original value in a construct that runs whether or not the test passed, because cleanup that only runs on success is cleanup that does not run when it is needed.

The lifecycle suite has since moved to a disposable tenant of its own. It tests transitions (suspended, closed, reactivated) that can leave an environment unusable, and a test of a destructive state change must never run against state anybody else depends on.

The partition made the argument by itself. When the shared tenant’s trial expired, the ten suites that stayed green were exactly the ten that create their own tenant.

What it did not fix

A batch run produced two failures that passed when run on their own. Another process was writing to the same tenant while the batch ran, and nothing in either suite was wrong.

No suite discipline removes that. When several processes write to one environment, a run measures the concurrency as much as the code. An isolated database per run is the right answer and it had not been built. Until it exists, a full run made while other work is in flight is useful for spotting obvious breakage and is not evidence. An authoritative run happens when nothing else is touching the environment, and that fact is recorded next to the number.

What to ask your own team or supplier

  • When a run goes red, how often is the cause the environment rather than the code, and who checks?
  • Do the suites share one test environment, and does each one leave it as it found it?
  • Can a suite be run twice in a row with the same result, and has anyone done that?
  • When a run is reported, does the report say whether other work was touching the environment at the time?
  • Which tests exercise destructive state changes, and what do they run against?

Where this ends up

Running a suite twice and requiring it to leave the tenant exactly as it found it is one of the gates described in AI-first delivery, where verification is kept as a separate act from production because a check that cannot fail proves nothing about either.

Working on something like this?

We build this kind of software, and we staff the teams that do.

Get in touch