[ Mobile · Architecture ]

Sync That Survives Real Networks.

replica A · phone · home cuff · 03:12    systolic 148 / 92
replica B · clinic tablet · 09:40        systolic 118 / 76
merge strategy: last-write-wins        result: 118, no error thrown
merge strategy: field registers + review   result: both retained, flagged
offline queue · 6 days · 1,204 ops · 0 lost
p95 convergence after reconnect: 4.2s   conflict rate: 0.31%

The Real Constraint

Sync gets designed in a conference room with full wifi. It gets used in a basement.

The app that started this argument was chronic care — patients recording vitals at home, clinicians reviewing them twice a week. Half the readings came from a cellar bathroom, because that is where the scale lives, on a five-year-old handset with a radio that drops to 3G before it drops to nothing. The recording moment is the product: show a spinner after the reading they just took and the patient stops taking readings.

“Just cache it” is not an architecture. A cache answers how do I avoid refetching; sync answers how do two replicas converge after independent writes — and a cache assumes the server is always the truth, which is backwards when the device holds the only copy of six days of readings.

The constraints on the wall:

  • Reads never touch the network. Every screen binds to a local query; no await fetch() on the render path.
  • Writes never block. A save is a local transaction plus an outbox row: sub-100 ms, radio off.
  • The radio is optional. Full CRUD works in airplane mode.
  • Offline is measured in weeks, not minutes. A queue that survives a weekend is not one that survives a tunnel.

Nine days offline — plane, rural cabin, a dead phone — is a first-class state with its own tests. Assume connectivity and you have a race condition with a progress bar.

Why Last-Write-Wins Loses Data

Here is the case that got this rewritten: systolicTarget, edited concurrently on two devices.

The patient’s phone writes 130 at 03:12 after a rough night. The clinic tablet writes 140 at 09:40 during a telehealth visit — a deliberate clinician decision. Both go offline. Last-write-wins keeps whichever timestamp is larger, and the phone’s clock runs four minutes fast, because phone clocks always do. The patient’s 130 overwrites the clinician’s 140.

No error. No toast. No audit row anyone will read. At the next visit the doctor sees 130 and assumes the target was reviewed and lowered. That is not a merge — it is silent clinical data loss with a green checkmark, and it is the failure mode that ends contracts.

Per-record LWW is worse still: one offline write plus a fresher snapshot clobbers the whole row, killing three unrelated edits because one field disagreed.

If two people edited the same field and the system quietly picked one, you don’t have a sync engine. You have a data-loss engine with excellent uptime.

LWW is fine where losing the older value costs nothing: avatar URL, theme, notification toggle, lastOpenedAt. Rule of thumb: if a human would have to notice that a value disappeared, it is not an LWW field.

The CRDT Primer You Actually Need

CRDTs are not a framework you adopt; they are four or five tiny data structures, and you pick one per field. That is the whole trick.

  • LWW-register with a replica tiebreak. Value plus (seq, replicaId) — a monotonic counter and a stable device ID instead of wall-clock time. It still picks one of two concurrent writes, but deterministically and identically on every replica.
  • PN-counter. Per-replica increments and decrements; merge takes the max of each side — doses logged, refills consumed, offline edit counts. An integer you just add to is a bug waiting for a second device.
  • Grow-only list (RGA-style). Inserts hold their position via fractional indices or a parent reference; deletes are tombstones — the reason an ordered reading history stays coherent when two people append at once.
  • Move-register / LWW-element-set. Membership plus ordering: care-team assignment, status enums, tags. “Which clinic owns this patient” is a move, not a delete-then-insert — the insert can land first.

That is roughly it: a vitals record is a dozen LWW-registers, one grow-only list, one move-register for status. Three types, no lock-in, easy to property-test by permuting replicas and asserting identical convergence.

Where CRDTs are overkill

Reference data the server owns — billing, tariffs, feature flags. Single-writer documents. Anything a human resolves anyway. A full text-CRDT engine bolted onto a form of twelve numbers is a war crime against a two-gigabyte phone: metadata and tombstones grow. Use the smallest structure that fits the field.

The Local Store

The local database is not a cache of the server. It is the source of truth for reads, permanently; the server is the rehydration and fan-out layer. Every screen queries local storage; the network layer only moves bytes in and out of the queues — the shape behind every mobile engagement where the field is the primary surface.

We use WatermelonDB where the UI needs lazy observable queries: indexed, on-demand subscriptions instead of hydrating a whole dataset into memory. Teams already living in raw SQL get the same from encrypted SQLite (SQLCipher, AES-256). (patientId, recordedAt) is not optional; it is the query your app runs forty times a session.

Keys live in the Keychain on iOS and the Keystore/StrongBox on Android, scoped ThisDeviceOnly / non-exportable and wrapped by the secure enclave. Never in SharedPreferences, never in UserDefaults, never derived from a PIN. Device-bound keys mean a copied database file is ciphertext: the attacker needs the handset, not the backup.

Which raises the question: what happens when the device is wiped? The key dies with the device and the old file becomes unreadable. That is correct behavior, not a bug. Recovery is re-auth plus a re-hydrate: snapshot, then replay the delta since the last epoch. It also dictates a server requirement people forget — keep enough history to rebuild a client from nothing, because “the phone has the data” and “the phone is gone” are both normal Tuesdays.

The Sync Loop

The loop is boring on purpose. Boring is the feature.

A mutation writes the row locally and an operation into the outbox in the same transaction — no window where the data exists but the intent does not. The flusher leases a batch (a few hundred ops, coalesced per document so a hundred keystrokes become one patch), sends it under an idempotency key the client minted, and advances a resumable cursor only after the server acknowledges. Retries use exponential backoff with full jitter — random(0, min(cap, base · 2ⁿ)), one-second base, five-minute cap — because synchronized retries are a thundering herd.

On clock skew: never trust device time. The device clock is a hint for display, nothing more. Ordering comes from a monotonic per-document sequence, causal metadata carries what-saw-what, and the server stamps a hybrid logical clock on ingest. Client timestamps surface in the UI labelled “device time” — so a clinician can see that 03:12 came off a fast phone.

// every op is idempotent — replaying a batch twice is a no-op
type Op = {
  id: string;          // client-minted UUID: the idempotency key
  doc: string;         // document id
  seq: number;         // monotonic per-document counter, never wall clock
  replica: string;     // stable device id, breaks concurrent-write ties
  wroteAt: number;     // diagnostics only — never a merge input
  patch: Record<string, unknown>;
};

export async function enqueue(db: DB, doc: string, patch: Patch) {
  const op = { id: uuid(), doc, replica, patch, wroteAt: Date.now() };
  await db.transaction(async (t) => {
    await t.apply(doc, patch);                       // local commit first
    await t.insert('outbox', { ...op, seq: await nextSeq(t, doc) });
  });
  scheduleFlush(400);                                // debounce, coalesce per doc
}

export async function flush(client: SyncClient, outbox: Outbox) {
  const batch = await outbox.lease(200);             // resumable cursor: since=<cursor>
  try {
    const ack = await retryWithJitter(
      () => client.push(batch),                     // POST /sync — server dedupes on op.id
      { baseMs: 1_000, capMs: 300_000, fullJitter: true },
    );
    await outbox.ack(ack.cursor);                    // advance ONLY on ack
  } catch (err) {
    if (isClientError(err)) await outbox.park(batch, err);   // 4xx: park it, do not loop
    throw err;                                        // network: back off, keep the batch
  }
}

// pure, commutative, idempotent — safe to run twice, in any order
export function merge(local: Rec, remote: Rec): Rec {
  return mergeFields(local, remote, (a, b) =>
    a.seq === b.seq
      ? (a.replica > b.replica ? a : b)              // deterministic tiebreak
      : (a.seq > b.seq ? a : b),
  );
}

Two details outweigh their line count. Ack-then-advance: move the cursor before the ack and a dropped response is silent data loss on the next boot; move it after and a crash just replays an idempotent batch. And parking 4xx — a validation failure is not a network condition, and retrying it forever is how an outbox grows to eighty thousand operations and takes the app down.

When Conflicts Must Reach a Human

Convergence is a math problem. Deciding which value is right is a judgment problem, and for a class of fields the correct answer is: stop, and show a person both values.

For PHI we do not silently auto-merge dose, allergy, and target fields. Concurrent writes go to a clinical review queue — a first-class table, not a log line — carrying both sides, both devices, both device-clock timestamps, and what each edit had seen. The clinician gets a card: two values, a diff, one tap to resolve. The queue is the audit artifact.

The rule that makes it survivable: quarantine the field, not the document. The rest of the record syncs; only the disputed fields show an amber marker with a count. Blocking the whole record behind a modal — on a phone, mid-consult — guarantees the conflict gets dismissed unread, which is silent auto-merge with extra steps. Conflicts should be visible, cheap to resolve, impossible to lose.

Never let a merge decide a dose

If the resolution of a conflict can change what a patient takes, a human resolves it. Automate the plumbing — detection, grouping, ordering, notification — and keep the decision. That boundary is the difference between a sync engine a clinical board approves and one it will not.

Encryption and PHI

Threat model first, features second — otherwise you are encrypting things because it looks responsible.

  • Lost or stolen device. Full-disk encryption plus SQLCipher with an enclave-wrapped key: a handset without a passcode gets ciphertext, an extracted backup gets ciphertext with no key material.
  • Hostile network. TLS 1.3 everywhere, certificate pinning on the sync endpoint — with a pinned backup and a remote rotation, because pinning without a rotation plan is an outage generator waiting for your next renewal.
  • Curious app on the same device. Scoped storage, no world-readable exports, PHI blurred in the app-switcher snapshot, no PHI in push payloads — a notification preview is a disclosure.
  • Insider and audit. Append-only access logs: who read which record, which field changed, from which device, which resolution a human chose. Regulators ask for the trail, not the algorithm.
  • Server compromise. The honest answer: if the server renders the data, the server can read it. Client-side field encryption is possible and expensive — scope it deliberately instead of trusting the perimeter.

All of it is boring, correct baseline — and it is what fails to a shortcut in week three, which is why we audit it before a feature ships. See it applied across our shipped work.

What We Measure

Sync that is not instrumented is folklore. These five go on the dashboard from day one; the last column is what we sign.

MetricWhat we watchSLO we write into the contract
Sync success rate99.9% accepted on first attempt; retries counted separately≥ 99.9% monthly, measured server-side
p95 convergence≤ 8 s after connectivity returns, all replicas agree≤ 60 s at p99, from ack timestamps
Offline queue depthp95 ≤ 40 ops; no growth after 72 h offline; alert at 2500 unrecoverable ops — ever
Conflict rate≤ 0.3% of touched fields; above 1% is a schema bug, not a user problem100% surfaced, 0% silently dropped
Battery cost≤ 1.2% per 24 h attributable to sync, worst device in the fleet≤ 1.5% per 24 h

The one in the contract is convergence: 99.9% of sync batches converge within 60 seconds of connectivity returning, evaluated monthly from server-side ack timestamps, with a 0.1% error budget that pages us, not the client. Note what it does not say: nothing about the client’s clock, nothing about a lab wifi network.

Six Lessons We Keep Relearning

  1. Client time is a decoration. Order with counters and causal metadata; let a wall clock pick a winner and you get the 130-vs-140 incident again.
  2. Idempotency keys come before everything. Before batching, before compression, before the clever delta encoding — without them retries are duplicates and every rare replay bug becomes corruption.
  3. Ack-then-advance, no exceptions. Move the cursor before the ack and you lose data on the exact crash users hit in a train tunnel. One line decides whether the rest of the design matters.
  4. Tombstones are cheaper than resurrected rows. Hard-deleting offline is how a record returns three days later with yesterday’s fields. Delete is an intent, not an erasure, until every replica has seen it.
  5. Every conflict you auto-resolve silently is a bug report you will never receive. Instrument conflict rates even for fields you do flag — the number tells you which schema assumption broke.
  6. Pull the cable in CI. Tests that run with full wifi on a clean clock test the wrong system. Random partitions, ±10 minutes of clock skew, a killed process mid-batch — on every pull request. That suite has caught more defects than any review.

None of this is exotic: smallest fitting CRDT, encrypted local store as read truth, idempotent ops behind an acked cursor, humans in the loop where it matters — applied with enough discipline to survive contact with a basement.

Put It to Work

Reading About It Is the Cheap Part.

If your app ships into basements, basements have opinions about your merge strategy. Bring the problem to a principal engineer — thirty minutes is usually enough to tell whether it is a blueprint problem or a build problem.