Reps' orders marked rejected when nobody rejected them
A field app marked valid orders as rejected by the office after a brief server restart. The fix meant telling a failed connection from a refusal.
A sales rep takes orders in the field. They come back to a list of orders marked as rejected by the office, and nobody in the office rejected anything. So the rep does the sensible thing and types the orders in again, and now the business has duplicates, or has to work out which copy is the real one.
That is what our field app did. During a deployment, the server returned errors for about a minute while it restarted, and a rep’s orders written in that window were marked rejected. We have no figure for the cost. The cost is duplicate orders entered by people who believed the first had been lost, and duplicates made that way are dearer to sort out than the original problem.
What was actually going on
The app saves each order on the phone and sends it when it can. The sending routine ran whenever the phone’s connection changed, and it treated any answer that was not “no connection” and not “conflict” as final. So a temporary server error and a real refusal of the order looked the same. Both produced a dead entry, with wording telling the rep the office had refused the change. One is true and one is not, and the app could not tell them apart.
There was also no retry schedule. Not a badly tuned one, none. An order that failed once for any server-side reason stayed failed until a person went looking for it.
Then came a second flaw, in our own first fix, found before it reached anyone. We gave each queued order five attempts, after which it was parked as failed, and we counted every attempt, including the ones where the phone never reached a server. A phone out of coverage for a day would have used up five attempts against nothing and parked a perfectly valid order as rejected. That is the exact silent loss the offline design exists to prevent, caused by the mechanism written to prevent it.
What we changed
Failures are now sorted into two kinds. A temporary failure means the server gave no verdict: it was busy, restarting, slow, or the connection dropped. The order is probably fine, so the app tries again later. A final failure means the server read the order and refused it, because it was badly formed, not permitted or not found. Trying again is pointless.
The key rule: only a server that answered uses up an attempt. A phone with no signal reschedules without spending anything, so it can be off the network for a week and find its queue intact. A server that keeps returning errors will still run out the attempts and be shown to a person, which is correct, because that is a real problem somebody needs to see. Where a busy server says when to come back, the app waits for that time instead of guessing.
Retry delays grow, and are randomised across the whole interval. Field staff cluster, in a branch office or a depot, and reconnect together. With a fixed schedule, every phone that failed at the same moment retries at the same moment, and keeps doing so, so the app itself creates a rush on a server that was already struggling.
An order waiting to be retried now has its own visible state, “retrying”, with its own wording and its own count on the status bar. Before, it had to show as pending, which says nothing is wrong, or as failed, which says the office refused it. A rep who can see it retrying does not report it or type it in again. The schedule is also now actually driven: the app sets a timer for the next attempt, where before it only woke when the connection changed, which is no help for “try again in ninety seconds”.
Adding the next-attempt time meant changing a database already on reps’ phones, which we cannot reverse. The change first checks what the phone already has, so an app upgrading from any earlier version ends up correct, and running it twice does no harm.
What it did not fix
A server that returns errors for long enough will still park a valid order for a person to handle. That is deliberate. The check behind the fix asserts both halves: phones with no signal lose no attempts, and answered failures do. Testing only the first half would pass on a queue that never gives up on anything.
The approach was not found by testing. It came from asking what the mechanism does to its user on the worst realistic day.
The pattern, for anyone whose reps use an app in poor signal
Ask how the app tells “the office said no” from “the message never got through”. Then ask what a rep sees in each case, and whether a rep can tell an order that is about to succeed from one that failed.
If they cannot, they will type it in again. The same question is worth asking of any retry rule, lockout, rate limit or expiry in your systems: what does it do to the person it was written for, on their worst day?
Where this ends up
The write queue in Sazinga Field works this way, so a rep who spent a day out of coverage comes back to orders that are still waiting to go, not orders the office appears to have refused.