Let's talk
engineering

An approve button that never approved anything for months

An expense system had an approval step no expense could reach. Before rebuilding it, we took out everything that was not doing work. Three tests for what is real.

· · updated

A cleared-out workshop with a nearly bare bench and discarded parts stacked by the wall.

Your expense system has an approval workflow. There is a filter for items waiting for approval, a status column, and approve and reject buttons on each expense. It was built, reviewed and demonstrated to you. Months later you discover no expense has ever waited for approval, because the system marked every expense approved the moment it was created. The buttons had never once done anything.

Nobody noticed, because a filter that returns nothing looks exactly like a filter with nothing to show. The cost is what the word “approval” meant to the business: you thought spending was being checked, and it was not, and the screen gave you no sign.

That was the clearest example of what the system had accumulated, and it set the approach for the rebuild: before adding anything, take out everything that is not doing work.

What was actually going on

We began with a read-only inventory of the system: fifty-two data models, forty-three route files and around thirty-three screens, with the specific redundancies listed instead of described. The approval step was the clearest case: the code that creates an expense set every one to approved on the way in, so the pending state the screen existed to handle could never occur. Tracing from the screen back to that code is what showed it.

What we changed

Removal came before rebuilding, in phases, code first and reversible, with the step that drops database tables held behind its own approval. The front end lost 342 files: route files, component folders, services and language files. The back end lost 139, and 39 tables were dropped once the code using them was gone. Two dashboards became one, and nine separate modules became a single workflow. Several whole areas went: a one-time-code login whose endpoints had been replaced, a recharge feature nobody used, two translations that had never been finished, and a layout component that every page overrode anyway.

We used three tests for whether a feature is real. The first: can the state it manages ever occur? Trace from the screen to the code that writes the field, and check that some path produces the value the screen exists to handle. A surprising number of workflow features fail this.

The second: is anything reading it? The recharge feature had a checkbox, a column, a filter and a set of cards on the detail page, and was never used, because the business handles that arrangement outside the system. The screens came out and the columns stayed in the database, which is the right trade: a screen costs attention to carry, and an unused column costs almost nothing.

The third: is it true? A dead login path is harmless. A screen that promises approvals to a business that has none is a false statement about the software, and false statements about software are expensive because people plan around them. The same test later removed marketing claims that were not true of the product, and a line in an app telling users their photographs would upload automatically when they reconnected, when nothing in the app listened for a reconnection.

Before the first file went, there was a safety net: a full database dump checked by listing what was inside it, a tag on the current version, and a branch from that tag, both saved off the machine. The rebuild worked against a copy of the database, with a check in the destructive scripts that refuses to run against the wrong name. The way back was one sentence: point the server at the previous database and restart, or redeploy the pre-trim branch.

What it did not fix

We got one thing wrong, which is sequencing. The front end was trimmed while a separate session trimmed the back end, and I deliberately did not touch the back end, to avoid a collision. That left a window where the two halves disagreed about which endpoints existed. It built and it ran, but nothing verified the pairing. In hindsight the two halves should have been one plan with an ordering, not two plans with a boundary.

The work was also done by several assistants working from an exact list of files, each diff checked by me and the build re-run. That worked because the list named specific files: twenty-three route files, thirty component folders and thirty services gives a change you can check, where “remove the payouts module” gives guesses at the boundary.

The pattern, for anyone who wants a rebuild and does not know what is real

Before anyone quotes you for a rebuild, ask for the inventory, and ask three things of every feature. Can the situation it handles ever happen? Does anything use it? Is it true: does it promise something the business does not do? Whatever fails goes out before anything new goes in.

And ask for the way back first. A rollback plan you have to invent under pressure is not a rollback plan. The cost of carrying a feature is not the disk it occupies. It is that everybody who reads the system afterwards, including you, believes it works.

Where this ends up

Reading a system until you know which parts are unreachable, and taking those out before rebuilding around the rest, is the first half of the work we do as enterprise software modernisation. It is also what lets the replacement happen one capability at a time rather than all at once.

Working on something like this?

We build this kind of software, and we staff the teams that do.

Get in touch