Let's talk
concurrency

Reps took orders without signal, and some went missing

A field sales app treated every conflict reply as order already saved. Some saved orders showed as failed and others could be deleted from the phone.

· · updated

A phone held over a distributor's trade counter; beside it the Sazinga Field app home screen with its sync bar.

Your rep is standing in a customer’s shop with no signal. He takes the order on his phone. The phone holds it in a queue and sends it when the connection returns. That is the whole point of a field sales app: the rep never loses work.

Ours managed to do the opposite in two ways. An order could reach the office and be saved while the rep’s phone showed it as failed. And an order that the office had genuinely refused could be removed from the phone’s queue as though it had been saved. In both cases the rep finds out days later, when a customer asks where the order went.

What was actually going on

There were three separate faults, and they compounded.

The first was on the server. When the phone resent an order the office had already saved, which happens whenever the reply gets lost because the phone dropped out of coverage again, the server did not answer “I already have that one”. A database error was only handled in one import routine, and everywhere else it reached the phone as a generic failure. The phone saw something other than the expected reply and marked the order failed. The order was in the database. The rep was looking at a screen saying it had not gone through.

The second appeared once the server was fixed to say “conflict” properly. The phone treated every conflict as “already saved” and deleted the queued copy. But a conflict can mean other things: two customers entered with the same reference by two reps, which is a genuine refusal, or a change such as marking a dispatch delivered on a record that had already moved on. Either way the rep’s work was deleted with no record.

The third was the retry allowance. The phone had five attempts, and attempts made with no connection at all were using them up. A phone out of coverage for a working day would spend its whole allowance on failures that said nothing about the order, then park a perfectly valid order as failed.

What we changed

The server now tells its replies apart. “This id already exists” means a repeat, and it is the only reply the phone treats as proof of saving. “This value clashes with another row” is a different reply, as is a reference to something missing. Any database fault we had not understood still surfaces loudly, because catching everything and returning something plausible hides real defects.

For any other conflict on a new order, the phone now asks the server whether the order exists by its own id. If it does, the order was saved. If not, it was refused and the rep is told. A conflict on a change, which has no such id, is treated as not applied. The rule is that nothing is marked sent unless it can be proved, and anything unprovable is shown to the user rather than guessed either way. The question is needed because for customers and products the server checks the business code first, so a repeat and a genuine clash look identical.

Only an attempt the server actually answered now counts against the allowance. Busy or restarting servers and timeouts are retried after a wait, spread out so a whole branch coming back online does not all retry at once. A request that is itself wrong stops at once and is shown.

What it did not fix

Some actions cannot safely wait in a queue, so they do not. A cancellation replayed against an order that has since moved on would be reported as done, so cancelling stays online-only. An upload that depends on a short-lived link cannot be replayed from a queue of order details, so it is not queued, and the screen says so before the rep confirms.

The pattern, for anyone buying a field sales app

Ask your supplier what happens if the connection drops after an order was saved but before the phone heard about it. Ask what the phone does with an answer it cannot classify. If the answer is “it assumes”, the queue will eventually lie to somebody, in silence, and you will hear about it from a customer.

The simplest test: put a phone in airplane mode for a working day, take a dozen orders, reconnect, and count them in the office. Then check that every order the phone showed as sent is in the office, and that every one it showed as failed is not there after all. The two lists should never disagree, and if they do, the queue is guessing.

Where this ends up

That is the failure Sazinga Field is built to avoid, with the order a rep takes in a shop and the order the office sees kept as one record that the phone can prove has arrived.

This came out of building Sazinga Field

Orders, stock, dispatch and the people on the road, in one place. The problem above is one we met while building it, and what we did about it is in the product.

If you run something like this, there is one thing you can do without a call: send one day's order sheet.