Skip to content

ADR-019: Single-document point saves with an asynchronous statistics fold

Context

The recovered-design reconciliation chose to update materialized statistics in the same MongoDB transaction as each Study point save (its "Primary update path" row). It left the contention threshold as an open evidence gate (item 5 of "Decisions still requiring Phase 0 evidence").

That evidence came in with PR #3855, the idle-host write benchmark of 30 September 2026:

  • Command cost. A materialized screening save issued 25 commands inside a multi-document snapshot transaction. A source-only save issued 2. Every cell failed the under-10% p95 gate, by 46% to 1,283%.
  • Contention. Every save writes shared per-project statistics documents: the control row, the project-scope current row, the publication guard and the outbox slots. At ten reviewers on different Studies, 399 of 1,000 materialized submissions exhausted their retries; source-only had none.

Decision

Replace the same-transaction point update for every screening- and annotation-dependent family with the design in async-point-fold-design.md. It sits behind the default-off flag materializedProjectStatisticsFold and a durable per-project FoldMode.

  1. Single-document pending entries.
  2. The save. A point save appends a compact pending entry (the classified transition plus any invalidation intents) to the Study in its existing single-document write. It writes no statistics document and uses no statistics transaction.
  3. The fold. A leased per-project worker in project-management folds entries into the rows, in batches of at most 32, in a transaction only it runs.
  4. Reads. Readers serve stored rows plus pending entries from one pinned snapshot, so values are exact before the fold.
  5. Rebuilds publish authoritative(B) - pending(B).
  6. Worker outages never fail a save: readers fall back past age and size thresholds.
  7. Storage-version tripwire. A fold-mode project's control carries StorageVersion = global + 1000 + FoldProtocolVersion. Every pre-fold binary already treats that as an incompatible project, so it fails closed. Rows and history never carry a fold marker. A cluster-gitops allowlist guard, bracketed by the durable narrow gate, covers the paths that do not check storage versions during a rollback.
  8. Protocol compatibility (N-1). Entries, the control stamp, heartbeats and fleet membership carry a fold protocol version.
  9. From the production baseline P0, a binary at N reads, folds and rebuilds N and N-1 entries, writes only the stamped protocol, and quarantines older entries.
  10. The stamp advances by an owned compare-and-set once every live member declares N.
  11. Only additive changes use the N-1 window. A derivation-changing change is a breaking bump that needs a reset of the affected families.

Owner decisions recorded (Chris, 30 September and 1 October 2026):

Decision
a MVP scope. All screening- and annotation-dependent families migrate in one MVP. ReviewerAnnotation and QuestionAnswers stay Stale-on-save (invalidation intents) and are served authoritatively
b Write gate. p95 overhead under 10% or at most +2 ms, and zero statistics-caused save failures or exhaustion at 1, 2, 5 and 10 reviewers, in same-Study and different-Study cells. Revised 2026-10-04 (amendment): latency gates the different-Study and 1-reviewer same-Study cells only; the same-Study 2, 5 and 10-reviewer cells are stress figures
c Version bump. The fold bumps the Study's Audit.Version when it removes entries
d Reviewer-tracking mode. It is read just before the write. A future mode-switch owner must add a writer grace period. Revised 2026-10-04 (amendment): it is read once per attempt, among the attempt's concurrent reads, with no pre-write re-read; the stage-two mode transition refuses to run inside the writer window (invariant 13)
e Per-stage reviewer tracking is deferred to #3876
f Rollback. The runbook order is followed and the allowlist GitOps guard, which existing binaries honour, is applied
g Staging pilot. Staging-only enable after slice 2, behind a positive IsStaging() allow setting. Production waits for the full MVP
h N-1 support from production P0. Between MVP slices the staging pilot resets instead

Consequences

  • The hot path. Point saves keep today's source-level shape: one Study write plus the existing non-statistics documents. Statistics add only two read-only control reads and a larger document. Gate (b) is evaluated by the benchmark's new fold arm.
  • New durable state:
  • Study.PendingStatistics and StatisticsFoldSequence;
  • control FoldMode and protocol fields;
  • per-protocol fold-worker heartbeats on the global control;
  • pmProjectStatisticsFleetMember;
  • pmProjectStatisticsFoldQuarantine (30-day TTL);
  • a partial index on pmStudy, built as a separate production operation.
  • History. History stays free of pending sets. Copies need an empty pending set, and authoritative builds commit an observation barrier.
  • Rollback requires the documented runbook. A rebuild of every family is required after any guard window.
  • Classifier fixes after production activation require a breaking protocol bump and a reset of the affected families.
  • Server version. MongoDB 4.4 or later is required.
  • Retrying a lost Study upsert. Once a fold can bump a Study's version between a writer's read and its save, a cached writer's version-guarded upsert can lose with E11000. The two retry shapes are not interchangeable:
  • StudyUpsertConflictRetry is for a whole-aggregate save that recomputes its change from a fresh uncached read. It reloads, re-applies and saves again, gated on the fold flag. A reload that differs only by the fold's own bump is retried without charge, up to a small cap. Use it for new cached writers (presence, idle-session and liveness consumers).
  • ReservationChangeSave.TrySaveAsync is for a reservation change that must commit with its statistics half in one transaction. It surfaces the lost race as ReservationChangeConflictException for its caller's own retry. Keep reservation and capacity paths on it; StudyUpsertConflictRetry also accepts that exception, so it can wrap such a call.

Amendment 2026-10-04: idle-host gate (b) result

Result. The idle-host write benchmark on Bramble (2026-10-03: 20 CPUs, load at most 0.27 per CPU, two timing runs and one counted run; results) failed gate (b) as then defined: one or two of sixteen cells passed. Statistics caused zero conflicts and zero failures in every cell. The failures were latency and the literal exhaustion clauses:

  • the uncontended fold save's only extra sequential round trip was the durable-mode re-read before the write (about 1.3 ms of a 2 ms allowance), and with ReviewEligibilityPolicy on, the mode read inside the eligibility transaction;
  • under same-Study contention the retry dominated (55-70% of the overhead): a charged fold retry was 9 commands in 5 sequential rounds against source-only's 3 in 3;
  • "fold exhausted equals source-only exhausted" failed in both directions in every contended cell, for reasons that were not statistics-caused, and varied between runs of one arm by as much as between arms.

Decisions (design owner, Chris, 2026-10-04), implemented together:

  1. Drop the pre-write mode re-read (design decision (d), revised). A fold-path attempt reads FoldMode and the durable reviewer mode once, among its concurrent reads; neither SaveScreeningOnFoldPathAsync nor a fold-path eligibility transaction reads the mode again. Why it is safe: the attempt's read still fails closed on disagreement; every entry records the mode it assumed and the fold, overlay and rebuild never apply an entry under another mode (they discard and stale); the Disabling writer grace (5 minutes) outlasts the pre-write deadline (30 seconds), so a save that lands after a disable is drained, and an older attempt reloads and falls back; and the source write is no more exposed to a mid-request mode change than today's statistics-off save, which never re-checks the mode at all. Invariant 13 still binds any future owner of a mode transition, and is now enforced in code: stage two of a mode transition refuses (Busy) until FoldPreWriteDeadline + FoldWriterGrace after stage one, and a tripwire test fails if stage one gains a production caller (today it has none, so no mode transition can happen in production).
  2. Make the retry cheaper. A charged retry reloads concurrently with its control-plane reads (two sequential rounds with the write), reads the bulk-update lock from that reload instead of a separate probe, and reads unrecorded-drop ranges only after an unknown result. Because the receipt and quarantine reads no longer follow the reload, a retry whose reload shows the Study's fold sequence moved reads the control plane again after the reload; without that, a fold landing between the two could hide a committed id (a test shows the double append it would allow). Every digest, duplicate, unknown-result and fencing rule is unchanged.
  3. Re-baseline the gate. Statistics-caused exhaustion = 0 replaces literal exhaustion equality; latency gates the different-Study cells and the 1-reviewer same-Study cells; the same-Study 2, 5 and 10-reviewer cells are stress figures with no PASS/FAIL but still must show zero statistics-caused conflicts. Why: those cells measure retry cost under lockstep contention on one Study, and the literal exhaustion comparison measured race noise; decision (b) itself only ever forbade statistics-caused failures.

Consequences. The uncontended screening fold save is 4 commands in 2 sequential rounds (was 5 in 3); a charged retry is 6 commands in 2 rounds with one namespace (was 9 in 5). Gate (b) is evaluated again by an acceptance rerun on Bramble after this change merges; production enable stays refused in code until it passes and a production rollout is separately approved.

References