Sync That Survives Real Networks.
The Real Constraint
Sync gets designed in a conference room with full wifi. It gets used in a basement.
The app that started this argument was chronic care — patients recording vitals at home, clinicians reviewing them twice a week. Half the readings came from a cellar bathroom, because that is where the scale lives, on a five-year-old handset with a radio that drops to 3G before it drops to nothing. The recording moment is the product: show a spinner after the reading they just took and the patient stops taking readings.
“Just cache it” is not an architecture. A cache answers how do I avoid refetching; sync answers how do two replicas converge after independent writes — and a cache assumes the server is always the truth, which is backwards when the device holds the only copy of six days of readings.
The constraints on the wall:
- Reads never touch the network. Every screen binds to a local query; no
await fetch()on the render path. - Writes never block. A save is a local transaction plus an outbox row: sub-100 ms, radio off.
- The radio is optional. Full CRUD works in airplane mode.
- Offline is measured in weeks, not minutes. A queue that survives a weekend is not one that survives a tunnel.
Nine days offline — plane, rural cabin, a dead phone — is a first-class state with its own tests. Assume connectivity and you have a race condition with a progress bar.
Why Last-Write-Wins Loses Data
Here is the case that got this rewritten: systolicTarget, edited concurrently on two devices.
The patient’s phone writes 130 at 03:12 after a rough night. The clinic tablet writes 140 at 09:40 during a telehealth visit — a deliberate clinician decision. Both go offline. Last-write-wins keeps whichever timestamp is larger, and the phone’s clock runs four minutes fast, because phone clocks always do. The patient’s 130 overwrites the clinician’s 140.
No error. No toast. No audit row anyone will read. At the next visit the doctor sees 130 and assumes the target was reviewed and lowered. That is not a merge — it is silent clinical data loss with a green checkmark, and it is the failure mode that ends contracts.
Per-record LWW is worse still: one offline write plus a fresher snapshot clobbers the whole row, killing three unrelated edits because one field disagreed.
If two people edited the same field and the system quietly picked one, you don’t have a sync engine. You have a data-loss engine with excellent uptime.
LWW is fine where losing the older value costs nothing: avatar URL, theme, notification toggle, lastOpenedAt. Rule of thumb: if a human would have to notice that a value disappeared, it is not an LWW field.
The CRDT Primer You Actually Need
CRDTs are not a framework you adopt; they are four or five tiny data structures, and you pick one per field. That is the whole trick.
- LWW-register with a replica tiebreak. Value plus
(seq, replicaId)— a monotonic counter and a stable device ID instead of wall-clock time. It still picks one of two concurrent writes, but deterministically and identically on every replica. - PN-counter. Per-replica increments and decrements; merge takes the max of each side — doses logged, refills consumed, offline edit counts. An integer you just add to is a bug waiting for a second device.
- Grow-only list (RGA-style). Inserts hold their position via fractional indices or a parent reference; deletes are tombstones — the reason an ordered reading history stays coherent when two people append at once.
- Move-register / LWW-element-set. Membership plus ordering: care-team assignment, status enums, tags. “Which clinic owns this patient” is a move, not a delete-then-insert — the insert can land first.
That is roughly it: a vitals record is a dozen LWW-registers, one grow-only list, one move-register for status. Three types, no lock-in, easy to property-test by permuting replicas and asserting identical convergence.
Reference data the server owns — billing, tariffs, feature flags. Single-writer documents. Anything a human resolves anyway. A full text-CRDT engine bolted onto a form of twelve numbers is a war crime against a two-gigabyte phone: metadata and tombstones grow. Use the smallest structure that fits the field.
The Local Store
The local database is not a cache of the server. It is the source of truth for reads, permanently; the server is the rehydration and fan-out layer. Every screen queries local storage; the network layer only moves bytes in and out of the queues — the shape behind every mobile engagement where the field is the primary surface.
We use WatermelonDB where the UI needs lazy observable queries: indexed, on-demand subscriptions instead of hydrating a whole dataset into memory. Teams already living in raw SQL get the same from encrypted SQLite (SQLCipher, AES-256). (patientId, recordedAt) is not optional; it is the query your app runs forty times a session.
Keys live in the Keychain on iOS and the Keystore/StrongBox on Android, scoped ThisDeviceOnly / non-exportable and wrapped by the secure enclave. Never in SharedPreferences, never in UserDefaults, never derived from a PIN. Device-bound keys mean a copied database file is ciphertext: the attacker needs the handset, not the backup.
Which raises the question: what happens when the device is wiped? The key dies with the device and the old file becomes unreadable. That is correct behavior, not a bug. Recovery is re-auth plus a re-hydrate: snapshot, then replay the delta since the last epoch. It also dictates a server requirement people forget — keep enough history to rebuild a client from nothing, because “the phone has the data” and “the phone is gone” are both normal Tuesdays.
The Sync Loop
The loop is boring on purpose. Boring is the feature.
A mutation writes the row locally and an operation into the outbox in the same transaction — no window where the data exists but the intent does not. The flusher leases a batch (a few hundred ops, coalesced per document so a hundred keystrokes become one patch), sends it under an idempotency key the client minted, and advances a resumable cursor only after the server acknowledges. Retries use exponential backoff with full jitter — random(0, min(cap, base · 2ⁿ)), one-second base, five-minute cap — because synchronized retries are a thundering herd.
On clock skew: never trust device time. The device clock is a hint for display, nothing more. Ordering comes from a monotonic per-document sequence, causal metadata carries what-saw-what, and the server stamps a hybrid logical clock on ingest. Client timestamps surface in the UI labelled “device time” — so a clinician can see that 03:12 came off a fast phone.
// every op is idempotent — replaying a batch twice is a no-op
type Op = {
id: string; // client-minted UUID: the idempotency key
doc: string; // document id
seq: number; // monotonic per-document counter, never wall clock
replica: string; // stable device id, breaks concurrent-write ties
wroteAt: number; // diagnostics only — never a merge input
patch: Record<string, unknown>;
};
export async function enqueue(db: DB, doc: string, patch: Patch) {
const op = { id: uuid(), doc, replica, patch, wroteAt: Date.now() };
await db.transaction(async (t) => {
await t.apply(doc, patch); // local commit first
await t.insert('outbox', { ...op, seq: await nextSeq(t, doc) });
});
scheduleFlush(400); // debounce, coalesce per doc
}
export async function flush(client: SyncClient, outbox: Outbox) {
const batch = await outbox.lease(200); // resumable cursor: since=<cursor>
try {
const ack = await retryWithJitter(
() => client.push(batch), // POST /sync — server dedupes on op.id
{ baseMs: 1_000, capMs: 300_000, fullJitter: true },
);
await outbox.ack(ack.cursor); // advance ONLY on ack
} catch (err) {
if (isClientError(err)) await outbox.park(batch, err); // 4xx: park it, do not loop
throw err; // network: back off, keep the batch
}
}
// pure, commutative, idempotent — safe to run twice, in any order
export function merge(local: Rec, remote: Rec): Rec {
return mergeFields(local, remote, (a, b) =>
a.seq === b.seq
? (a.replica > b.replica ? a : b) // deterministic tiebreak
: (a.seq > b.seq ? a : b),
);
}Two details outweigh their line count. Ack-then-advance: move the cursor before the ack and a dropped response is silent data loss on the next boot; move it after and a crash just replays an idempotent batch. And parking 4xx — a validation failure is not a network condition, and retrying it forever is how an outbox grows to eighty thousand operations and takes the app down.
When Conflicts Must Reach a Human
Convergence is a math problem. Deciding which value is right is a judgment problem, and for a class of fields the correct answer is: stop, and show a person both values.
For PHI we do not silently auto-merge dose, allergy, and target fields. Concurrent writes go to a clinical review queue — a first-class table, not a log line — carrying both sides, both devices, both device-clock timestamps, and what each edit had seen. The clinician gets a card: two values, a diff, one tap to resolve. The queue is the audit artifact.
The rule that makes it survivable: quarantine the field, not the document. The rest of the record syncs; only the disputed fields show an amber marker with a count. Blocking the whole record behind a modal — on a phone, mid-consult — guarantees the conflict gets dismissed unread, which is silent auto-merge with extra steps. Conflicts should be visible, cheap to resolve, impossible to lose.
If the resolution of a conflict can change what a patient takes, a human resolves it. Automate the plumbing — detection, grouping, ordering, notification — and keep the decision. That boundary is the difference between a sync engine a clinical board approves and one it will not.
Encryption and PHI
Threat model first, features second — otherwise you are encrypting things because it looks responsible.
- Lost or stolen device. Full-disk encryption plus SQLCipher with an enclave-wrapped key: a handset without a passcode gets ciphertext, an extracted backup gets ciphertext with no key material.
- Hostile network. TLS 1.3 everywhere, certificate pinning on the sync endpoint — with a pinned backup and a remote rotation, because pinning without a rotation plan is an outage generator waiting for your next renewal.
- Curious app on the same device. Scoped storage, no world-readable exports, PHI blurred in the app-switcher snapshot, no PHI in push payloads — a notification preview is a disclosure.
- Insider and audit. Append-only access logs: who read which record, which field changed, from which device, which resolution a human chose. Regulators ask for the trail, not the algorithm.
- Server compromise. The honest answer: if the server renders the data, the server can read it. Client-side field encryption is possible and expensive — scope it deliberately instead of trusting the perimeter.
All of it is boring, correct baseline — and it is what fails to a shortcut in week three, which is why we audit it before a feature ships. See it applied across our shipped work.
What We Measure
Sync that is not instrumented is folklore. These five go on the dashboard from day one; the last column is what we sign.
| Metric | What we watch | SLO we write into the contract |
|---|---|---|
| Sync success rate | 99.9% accepted on first attempt; retries counted separately | ≥ 99.9% monthly, measured server-side |
| p95 convergence | ≤ 8 s after connectivity returns, all replicas agree | ≤ 60 s at p99, from ack timestamps |
| Offline queue depth | p95 ≤ 40 ops; no growth after 72 h offline; alert at 250 | 0 unrecoverable ops — ever |
| Conflict rate | ≤ 0.3% of touched fields; above 1% is a schema bug, not a user problem | 100% surfaced, 0% silently dropped |
| Battery cost | ≤ 1.2% per 24 h attributable to sync, worst device in the fleet | ≤ 1.5% per 24 h |
The one in the contract is convergence: 99.9% of sync batches converge within 60 seconds of connectivity returning, evaluated monthly from server-side ack timestamps, with a 0.1% error budget that pages us, not the client. Note what it does not say: nothing about the client’s clock, nothing about a lab wifi network.
Six Lessons We Keep Relearning
- Client time is a decoration. Order with counters and causal metadata; let a wall clock pick a winner and you get the 130-vs-140 incident again.
- Idempotency keys come before everything. Before batching, before compression, before the clever delta encoding — without them retries are duplicates and every rare replay bug becomes corruption.
- Ack-then-advance, no exceptions. Move the cursor before the ack and you lose data on the exact crash users hit in a train tunnel. One line decides whether the rest of the design matters.
- Tombstones are cheaper than resurrected rows. Hard-deleting offline is how a record returns three days later with yesterday’s fields. Delete is an intent, not an erasure, until every replica has seen it.
- Every conflict you auto-resolve silently is a bug report you will never receive. Instrument conflict rates even for fields you do flag — the number tells you which schema assumption broke.
- Pull the cable in CI. Tests that run with full wifi on a clean clock test the wrong system. Random partitions, ±10 minutes of clock skew, a killed process mid-batch — on every pull request. That suite has caught more defects than any review.
None of this is exotic: smallest fitting CRDT, encrypted local store as read truth, idempotent ops behind an acked cursor, humans in the loop where it matters — applied with enough discipline to survive contact with a basement.
Reading About It Is the Cheap Part.
If your app ships into basements, basements have opinions about your merge strategy. Bring the problem to a principal engineer — thirty minutes is usually enough to tell whether it is a blueprint problem or a build problem.