Let's talk
security

Signed out, and still able to get in for eight hours

Signing out of this compliance system changed nothing on the server. A copied login stayed good for eight hours, and access changes left no trail for an auditor.

·

Hands cutting a copy of a brass key in a locksmith workshop with spare keys beside them.

A member of staff leaves, or moves to a role that should see less. You remove their access, or you ask them to sign out of the shared machine in the site office. Are they actually out? And when the auditor asks who changed whose access last quarter, and who reset whose password, can you show them?

In the system this article is about, a compliance product built to demonstrate exactly that kind of control to an auditor, the honest answers at the time were no and no. Signing out did nothing beyond the person’s own screen. A login that had been copied — from a shared machine, from a browser add-on, from a support session where somebody pasted a request — stayed good until it ran out on its own, which was eight hours. Nothing anywhere in the product could shorten that by a second. And the two actions an auditor is most likely to ask about, changing somebody’s access and resetting their password, left no trace at all.

The cost is not a breach; none happened. It is that a product whose purpose is proving control to a third party could not, at that point, prove it about itself, and that was found only after it had been demonstrated to several people.

What was actually going on

The system had a place to record sessions: a table designed to hold each issued login, the device it went to, when it expires, when it was last seen, and when it was withdrawn. There was even a note above it explaining that it existed for sign-out, revocation and audit. That is the correct design, written down by somebody who knew what they were doing.

Nothing used it. Signing in issued a login and recorded nothing. Every request checked that the login was genuine and not yet expired, and looked nothing up. There was no sign-out on the server at all; signing out, in the app, simply forgot the login locally.

That is what makes it worth writing about. Sessions are commonly deferred. Here the storage existed, fully specified, revocation column and all, and its presence persuaded everyone. A reviewer reading the data design concludes revocation is handled. A reviewer reading the login check assumes the lookup happens elsewhere, because there is a table for it. Neither is being careless. Well-designed storage is real evidence of intent and no evidence at all of behaviour.

Searching for tables nothing writes to found three: the sessions table, the summary table behind the dashboard’s headline figure, and the history table meant to record changes to compliance rules. All three were described as things a job or a flow would maintain. None was.

The change that needed a re-login, and the change that needed a transaction

A second problem sat beside the first. The system deliberately re-checked a person’s permissions on every request, paying an extra database query each time, so that taking a permission away would take effect immediately. But the app cached the permission list it received at sign-in and never refreshed it. The server enforced the new rules at once; the screen kept offering the old buttons, which now failed. The message returned when roles were changed said, in so many words, that the user must sign in again for the change to appear. A documented compromise, shipped as a sentence.

Changing a person’s roles was also done as two steps — remove all their roles, then add the new set — with nothing binding them together. A failure between the two leaves the person with no roles at all. Locked out rather than over-privileged, which is the better way to fail, and still an outage for somebody.

What has to change

A login that can be withdrawn, so that signing out on the server ends the session and removing a person ends theirs, whatever device still holds the copy. Access changes and password resets recorded in the audit log, because those are the two questions an auditor asks first. A role change made as one operation, so it cannot half-complete. And a way for the app to learn that its permissions have changed, so the extra cost of checking on every request actually buys the behaviour it was paid for.

What it did not fix

None of this was found by anyone being attacked. It was found by reading the sign-in path end to end, which took about an hour, and which was not done until the product had been demonstrated to several people. That is the wrong order, and I know it.

The mechanism

A token you cannot revoke is a credential with a fixed lifetime. Decide that deliberately and make the lifetime match: eight hours is a reasonable session length and an unreasonable time to be unable to withdraw access, and without a session store you only get to pick one.

Underneath the permission check sat a sharper problem. The function that resolves a user’s permissions took the organisation as an argument, marked it unused, and never referenced it. Its query joined role assignments to grants with no filter on which tenant a role belonged to, so a user assigned a role from a different organisation would have been granted its permissions. The write path guarded against that in one endpoint; nothing guarded the read path, and nothing in the database prevented it, because the assignment table lacked the composite key that makes cross-tenant references impossible elsewhere. The same query also did not exclude soft-deleted roles, while the endpoint that lists a user’s roles did. Two functions in the same file, reading the same data, with different opinions about what deleted means, and the one that grants access is the one that matters.

The rules that came out of it: search the schema for tables nothing writes to, because storage is intent and queries are behaviour, and where they disagree the schema will be believed and the code will be running. Keep one definition of deleted, in one place. Wrap replace-all writes in a transaction, because delete-then-insert is one operation in the user’s mind and two in the database, and the gap between them belongs to nobody. And if you pay for per-request authorisation, spend the rest of it: either put permissions in the token and state the staleness, or resolve them per request and give the client a way to be told they changed.

Where this ends up

Sessions belong with the data model and tenancy among the handful of decisions that are expensive to reverse once a product has users, which is why they are settled properly in a first version rather than after the demonstrations — the approach is set out under startup MVP development.

Working on something like this?

We build this kind of software, and we staff the teams that do.

Get in touch