Let's talk
operations

A customer's trial ends and your software locks them out

An expired trial silently put a test environment read-only for four days. The same thing will happen to the first customer whose card fails.

· · updated

A striped barrier pole lowered across the entrance to an empty car park at dusk.

A customer’s trial runs out, or their card is declined, and your software simply stops letting them save anything. There is no warning first. Nobody on your staff has a button to switch them back on. The first you hear of it is when their team cannot save an order and someone rings the support line.

We found this in our own test environment before any paying customer met it. A routine check of the whole project showed twenty-five of thirty-five test suites failing, with no release and no change to the code for weeks. The cause was not a fault. The shared test tenant’s trial had ended four days earlier, and the trial rules, working exactly as designed, were refusing every write.

What was actually going on

It was proved, not assumed. Every failing suite failed at its first attempt to save something, with the trial-expired message. The ten that passed were exactly the ten that create a fresh tenant of their own instead of using the shared one. That split was the signature, and it found the cause in minutes.

For four days the environment had been read-only. Testers were blocked, the browser tests could not run because they cannot sign in and create anything, and nobody reported it. There had been no banner, no email and no log line anyone was watching. The refusals simply began on a Tuesday.

The rule itself was correct. It expires trials and refuses writes, and it works out expiry when a request arrives, not on a schedule, which is the right design because a scheduled job that fails leaves customers on a free trial for ever. The gap was everything around it. Nothing warned a tenant three days before, or on the day. Reopening a tenant was a technical request with no screen, so only someone able to send authenticated requests by hand could unblock a customer.

What we changed

The finding moved the administration console from a finishing touch, to be built after the features, to a requirement that has to exist before the first customer does.

The rule itself had been wrong twice, and both fixes are in. It compared against the universal date instead of the tenant’s own business date, so a tenant several hours ahead lost the last hours of its trial. And it refused everything, reads included, so an expired customer could not even see the records they had entered, nor the screen telling them how to subscribe.

It now separates three states. An expired trial or a suspended tenant can read everything and change nothing, and the screen explains why. A closed tenant is blocked entirely, which had not been checked at all before. The rule is applied to any request that changes data, rather than listed route by route, so a new feature added next year is covered automatically. The call that reopens a closed tenant is exempt from the gate it lifts, otherwise closed would be a state with no way out.

Before the rule went live on the shared environment, we queried how many of its 126 tenants already had an expired trial. The answer was zero, so switching it on changed nothing that day. Had it been forty, forty tenants would have been locked out together.

What it did not fix

This work did not build the warnings before expiry or the screen for staff to reactivate a tenant. Those are what the finding calls for, and as written here they are still open. Reactivation remains a technical request with no interface, so a customer whose card fails would meet the same lock with nobody able to lift it quickly.

The pattern, for anyone who charges by subscription

Ask three things of your own software. What happens, on screen, the day a trial or payment lapses? Who on your staff can undo it without a developer? And before any new rule that refuses things goes live, how many customers will it catch today? One query answers that, and it is the difference between a release and an incident.

A lock with nobody able to lift it is not half a feature. It is an outage with the date already set.

Where this ends up

It is one of the reasons a first version settles customer accounts and lifecycle properly and writes down what was deferred. Admin tools are often the right thing to defer, but a rule that can only be lifted by hand is the point at which deferring it stops being free.

Working on something like this?

We build this kind of software, and we staff the teams that do.

Get in touch