Let's talk
concurrency

Field staff on poor signal: records lost or repeated

A mobile app on two bars of signal needs incremental sync. A timestamp cursor drops records silently or repeats them forever. The design that avoids both.

· · updated

A field worker in a warehouse holding a phone showing the proof-of-delivery screen in Sazinga Field, records syncing over poor signal.

Your field staff are in a warehouse with two bars of signal. The app lets them queue work offline and replay it, but it cannot ask “what changed since I last looked”. Every refresh reloads every list, which is tolerable at a desk and close to useless on that connection.

The cheap way to add incremental sync has a worse failure than slowness. A record can be missing on a phone, permanently, with no error anywhere. Or the same records can come down again on every refresh and the sync never finishes. Nobody would see either until a rep worked from a list that was wrong.

What was actually going on

The missing feature was a hard block, not a spare afternoon. The API had no way to ask for changes since a point in time. A strict query-parameter guard had been introduced so that an unrecognised parameter returned an error instead of being silently ignored, so attempts to pass a since-timestamp or cursor failed on every resource. Even the crude fallback was shut: sorting by last-updated was outside the sort allow-list on the two resources that mattered most.

That was checked against the running system and written up as required API work. Faking incremental sync by refetching everything and calling it sync would have shipped a feature that misstated its own cost.

The naive cursor is the last update time seen, asking for everything later than that. Update timestamps are not unique. In PostgreSQL the usual default is the transaction start time, so every row written by one transaction carries the same timestamp to the microsecond. A bulk import writes four hundred rows with one timestamp. A cascade writes a parent and its children with one.

If a page boundary lands inside such a group, a strictly-later cursor skips the rest of the group permanently, and nothing reports an error. A later-or-equal cursor repeats the group at the start of every following page. If the group is bigger than the page, the sync never advances, and it looks like a network fault.

What we changed

Sync became its own route family, not an extra parameter on the existing list endpoints. Offset pagination and cursor pagination are different contracts, and merging them makes a request with a sort and a cursor valid and silently wrong. Sync must return deleted rows and ordinary lists must never. The response shapes differ. The new routes share each resource’s scope helpers and permission key with its list, so the two cannot disagree about who may read what.

The cursor orders by a pair and compares it as a pair: the update timestamp and the row identifier, with a true row comparison in the query. The pair is unique, so the order is total, and a cursor holding both resumes exactly where it stopped however many rows share a timestamp. The supporting index is on tenant, update time and identifier, deliberately not a partial index that excludes deleted rows, because sync is the one path that must see them.

Deletions travel as tombstones: the identifier, the update time, a deleted flag, no payload. The cursor is opaque and carries a format version and the resource it belongs to, so a cursor from one resource replayed against another is refused instead of producing a plausible wrong result. A discovery route lists which resources the caller may sync.

The tests check both errors, because they are opposites. A paged walk asserts every row arrives exactly once. A row modified mid-walk is tested both ahead of and behind the cursor. Tombstones, a malformed cursor and a cursor from the wrong resource are covered. Data scope is proved in both directions, since an implementation that returns nothing passes the exclusion half perfectly.

What it did not fix

Transaction start time is not commit time. A long transaction can commit a row whose update time is earlier than a position a client has already passed, and that client never sees the row. For this system’s write patterns that is theoretical and has not been observed. The proper fix is a commit-order cursor, and it is written down as a known hazard with the shape of its solution. It is not built.

Catalogue resources, meaning products, variants and price lists, apply no branch narrowing on sync, because their ordinary lists apply none either. Inventing one for sync alone would have given a rep a synced catalogue that disagreed with the one the app shows online.

What to ask your own team or supplier

  • How does the mobile app learn what changed since it last synced, and what does it do when thousands of rows share an update time?
  • How does a deleted record reach a phone?
  • Has a test walked every page and proved each record arrives exactly once, not at least once or at most once?
  • Which known gaps in the sync are written down, and who owns them?

Where this ends up

Incremental sync in Sazinga Field rests on this composite cursor, so a phone on two bars pulls what changed instead of refetching every list.

This came out of building Sazinga Field

Orders, stock, dispatch and the people on the road, in one place. The problem above is one we met while building it, and what we did about it is in the product.

If you run something like this, there is one thing you can do without a call: send one day's order sheet.