Skip to content

Temporary planning review record. The report below is reproduced verbatim as returned by the independent read-only reviewer (Plan agent, Opus model, launched 3 October 2026 about 14:35 BST). Only this front matter and note were added. Its finding IDs (AC-01 to AC-35) are review IDs, distinct from the plan's acceptance-criterion IDs (AC-<release>-nn). Line numbers refer to the package as it stood when the reviewer read it. Resolutions are in the round-2 resolution matrix.

Review D: acceptance criteria and testability (adversarial, read-only)

Reviewer scope: this review covers acceptance-criteria.md and the release sections of integrated-plan.md. I traced them against the owner ledger, the §1.11 decisions, research A1–A28, PRISMA fixtures 1–8, the access-policy, lifecycle and setup fixtures, the QM v2 tracker, FEAT-011, FEAT-012, the eligibility policy and FEAT-023. I also checked the repository's actual test tooling on main. Nothing was modified. Finding IDs AC-01 to AC-35 are review IDs. Proposed criteria follow the plan's scheme (AC--, UI-, PE-).

Path abbreviations (absolute): - PKG = /home/chris/workspace/syrf/pr/pr3617.research-screening-as-specialised-annotation-gxgahs/docs/planning/integrated-review-plan-2026-10 - PLN = /home/chris/workspace/syrf/pr/pr3617.research-screening-as-specialised-annotation-gxgahs/docs/planning - MAIN = /home/chris/workspace/syrf/main - ACD = PKG/acceptance-criteria.md - IP = PKG/integrated-plan.md - LEDGER = PLN/review-form-owner-decisions-2026-10-02.md

1. Verdict

The acceptance package is a real step forward: 214 numbered criteria with verification methods, a common checklist, a Material 3 standard and pilot criteria. It is not yet a reliable basis for shipping, for three reasons.

Traceability. The "authoritative list" names only 13 of the roughly 90 ledger and §1.11 identifiers. By content: - At least ten confirmed decisions have no criterion at all: PV1, VS1, RE1, NT1, TC1, QY5, RE4's single cross-stage task, Q-10, Q-13 and amendment H. - The contracts' conformance lists are not binding at ship gates. - Several criteria assert recommended answers to open questions as if they were decided, and six say "as Q-xx decides", which can never fail.

Evidence model. Every criterion is checked at a release ship gate on an "exact head". But a release is many trunk-based PRs: each one auto-deploys to staging, and production promotion follows when an administrator merges it. Continuous properties (flags-off, legacy unchanged) are therefore checked months after dark code reaches production, and staging acceptance runs on a moving image.

Executability: - Key fixtures are one-line sentences: PRISMA 1–8 and their undefined "variants", the access-policy, lifecycle and setup lists, ASySD parity, and the 10,000-session publication. - The tooling the methods assume is missing on main: axe, visual capture, non-Chromium runs, a mixed-version rollback harness, and a benchmark harness for this programme. - The e2e stack has three personas, one of them an application administrator. - The seed plan relies on a staging mechanism that only transfers ownership, and on a legacy seeder that R0 will refuse.

Several criteria can also pass while the behaviour is wrong: - R0's compatibility check can pass with a read-only test even though Study is written by whole-document replace. - "Every qualifying candidate takes part" has no negative cases. - The concurrency criteria never force the race. - UI-3 must be verified on an M3 component path that doesn't exist yet.

The five most important changes: 1. Make the criteria traceable and binding. Add a Source/status column and a generated decision → criterion → test matrix. Add the missing criteria (§3.1). Make each contract's conformance suite a release criterion. Rewrite placeholder criteria as concrete provisional assertions that block their release until the question is answered. 2. Split the definition of done in two. Merge criteria apply to every PR: the flags-off and legacy spec set, no excluded specs in touched areas, and tests tagged with criterion IDs. Activation criteria are checked on a recorded release-candidate image SHA and kept in a per-release acceptance record. Tier the human gates by release risk. 3. Deliver the acceptance tooling (lane L17) in W0: - the persona set - axe checks inside journey specs, and screenshot capture - Firefox and WebKit smoke projects - deterministic concurrency barriers - a mixed-version harness - benchmark datasets on FEAT-024's existing gated harness, sized to its 25,000-study dataset and AF2's 1,000-question gate 4. Turn fixtures into versioned data with per-release evidence assertions, and replace the seed rule with an additive canonical "seed-if-absent" job plus three test-data tiers. 5. Make pilots measurable: entry criteria from §5.10, an invariant monitor, telemetry, severity definitions and minimum exposure. Decide how UI1 (Material 3) is verified before FEAT-023's cutover.

2. Findings

ID Severity Location Finding Evidence Recommended change
AC-01 Blocker ACD §1–§4 (:14-18, :27); PKG/contracts.md conformance lists Confirmed decisions have no binding criterion. The list names 13 decision IDs (SF6, VS2, EW1, FV4, RA5, AG2, AG3, DP2, DP5, DP6, DP7, PR1, QD1). By content, PV1, VS1, RE1, NT1, TC1, QY5, RE4 (one task across stages), Q-10, Q-13 and amendment H have none, and about 25 more are partial (§5.2). Conformance items in contracts.md that aren't repeated in ACD are not ship-gate binding, because the gate is "every criterion below". Examples: C1's A7 and A11; C2's legacy duplicate conflicts and candidate vs reconciled heads; C3's "reused answer keeps its authoring stage" and lost exposure reports; C5's order-only change and withdrawn session; C6's access-policy fixtures and A12–A16. ACD:14-18; rg of ACD for each ID; LEDGER:39 (PV1), :45 (VS1), :58 and :313-319 (RE1), :365-372 (NT1), :826-845 (TC1), :220 (QY5), :655-669 (RE4); PKG/decision-register.md:189 (Q-10), :193 (Q-13); PKG/prisma-amendments.md:139-153; PKG/contracts.md:127-129, :151-153, :167-170, :273-276, :313-314 Add a Source column to every row (ledger ID / RECOVERED / PROPOSAL / Q-xx pending / A-xx). Add the §3.1 criteria. Add one row per release, "AC--CONF: the conformance suites of contracts Cx pass in full on the release candidate", and give each conformance item an ID (C1-T01…). Generate decision → criterion coverage from a traceability file and fail docs CI when a ledger ID has no criterion (§3.3).
AC-02 Major ACD:27 and the rows listed Open questions are treated as decided, and placeholders can't fail. Asserted as fact: AC-R4a-05 (Q-36), AC-R4a-09 "target-1 creates no task" (Q-29), AC-R3a-07 "new combined steps default to Stop" (Q-24), AC-R3a-01 "collective satisfaction for an unvoted reviewer" (A-06/Q-15), AC-R3c-02 (Q-02 flow), AC-R4a-08 "most restrictive blinding" (Q-28), AC-R2d-08 (Q-20). Placeholders that can't fail: AC-R1c-03, AC-R3b-05, AC-R3c-07, AC-R4a-12, AC-R5c-01, AC-R5c-03. Q-32, E1 and E11 have no row at all. ACD:27 ("Everything else follows from confirmed decisions or from the plan"); ACD:210, :216, :222, :245, :270, :273, :274; ACD:135, :235, :250, :277, :323, :325; PKG/open-questions-and-assumptions.md:64-67, :70, :79-84 Add a status per row: confirmed / pending-Q-xx / assumption-A-xx. Write each pending row as the concrete assertion of the recommended answer. Example for AC-R4a-12: "a legacy reconciled answer exports as 'legacy reconciled (authority unknown)', can't be queried, and R4a readiness treats the study as needing canonical reconciliation". Getting the answer and re-confirming the row at the freeze gate becomes part of the release's definition of ready. A ship gate can't count a pending row as passed.
AC-03 Major ACD:16-18, :53, :61; IP:251-257, :962-963 The evidence model doesn't fit trunk-based delivery. Releases are many small PRs behind default-off flags. Each merge goes to staging automatically and to production when an administrator merges promotion. AC-ALL-01/02 (flags-off, legacy unchanged) are checked only at the ship gate, and "exact head" (AC-ALL-05/13) is a PR concept applied to a release with no single head. Staging acceptance happens on a moving image. MAIN/CLAUDE.md:131; ACD:49-50, :53, :61; IP:962-963; exact-image precedent at MAIN/docs/features/material-3-migration/technical-plan.md:310-318 Classify AC-ALL into merge criteria (01, 02, 05, 06, 10 plus new 14–16) enforced on every PR, and activation criteria evaluated on a recorded release-candidate commit and image SHA on staging. Re-run the affected automated rows when later merges touch the release's paths. Adopt the DoR/DoD in §3.2 and a per-release acceptance record.
AC-04 Major ACD:171, :183, :197, :220, :238, :332, :351, :359, :365, :373; §7 :462-476 PRISMA fixtures have no oracle before R5b. Each fixture is a compound one-sentence description marked "proposed meaningful tests". "Variants 2 and 4 for single-stage forms" are defined nowhere. No PRISMA computation exists before R5b, so "fixture N passes" means nothing in R2a, R2b, R2c, R3a, R3b, C1, O1 or P1. R2a binds one stage per form, so fixture 2's two-stage subject can't even be set up there. PLN/review-prisma-integration-2026-10-03.md:141-165; "variant" occurs only at ACD:171, :470, :472 and IP:437; IP:389-391 Make each fixture a versioned data file FX-PRISMA-01…08 with: (a) inputs; (b) per-release evidence assertions (R2a: one qualifying contribution per reviewer and form; R3a: one current ScreeningOutcome per profile, PoolEntryEvent count and versions, no Included from a personal Include; P1: one downstream source column; C1/O1: Study count and metaAnalysisIncluded unchanged); © R5b box values including reported-external parts. Write 02/03/04 before F3, 01/07 before F-P, and 05/06/08 before F6b.
AC-05 Major IP:560-562, :573-574; ACD:244-260 R3c and R3d point at fixture lists that ACD only partly carries. Missing lifecycle cases: approval spanning two Completed stages (all-or-nothing); approval invalidated when inputs change; approval never granting the requester authority; approval expiry. Missing access-policy cases: advanced personal cross-stage access opens; self-cycle rejected; changed published policy rechecked at transactional admission. Missing setup cases: no duplicate project creation; an empty project creates no PRISMA records; extraction-off adds no unused types; living-search and ML stay disabled. PLN/review-lifecycle-gold-settings-proposal-2026-10-03.md:11, :36-59; PLN/review-stage-step-access-policy-proposal-2026-10-03.md:100-105; PLN/review-guided-setup-template-plan-2026-10-03.md:126-134 Make every fixture a criterion row (AC-R3a-14, AC-R3c-08 to -10, AC-R3d-06 and -07 in §3.1), marked pending Q-01/Q-02 where it encodes the proposal. Replace IP:560 and IP:573's references with the IDs.
AC-06 Major IP:501-524 vs ACD:212-225 R3a's core scope has no criteria. Missing: atomic Complete-and-Include (research A8); Skip records nothing (amendment A, RC6); AND/OR dependency groups; compulsory, handoff and terminal scope; the default profile reproducing legacy threshold maths (A2); consumption of the R2a capture log; the per-profile outcome shape with authority values and structured reasons (amendment H; FEAT-011 phases 13 and 15); "entering screening" fixtures for early-stopped and batched reviews; regression of the eligibility test plan that C6 extends. IP:503-524; rg finds no skip/terminal/compulsory/AND/OR in ACD; PLN/screening-specialised-annotation-research.md:1247, :1253; PKG/prisma-amendments.md:61-66, :139-153; MAIN/docs/features/prisma-specification/prisma-constraint-annotations.md:225-228, :243-245, :260-264; PLN/review-eligibility-policy.md:1204-1241 Add AC-R3a-11 to -20 (§3.1).
AC-07 Major IP:382-418 vs ACD:154-172 R2a's engine and definitions scope has no criteria. Missing: per-project commit sequence (E25, which AC-R5a-02 depends on); server-minted IDs (E27; research A5 forged IDs); maximum form size (E28); system-question snapshots (E24); automatic ancestor inclusion (QM-09); one stage per form until R2b (A-21); presentation state kept outside versions; PV1 provenance; restore (E31); erasure (E32); target sourced from the form. IP:383-397, :405-412; PKG/open-questions-and-assumptions.md:126-134, :203; PKG/contracts.md:265-266, :273-276; PLN/screening-specialised-annotation-research.md:1250 Add AC-R2a-20 to -31 (§3.1).
AC-08 Major ACD:266-278; IP:582-604 R4a criteria state only the positive half. AC-R4a-09 ("every qualifying candidate takes part") also passes if a saved-incomplete-after-Complete, incompatible, withdrawn or ineligible session is used. The ledger's delivery list says those are excluded. Also missing: no gold from agreement or majority before final submission; RE5 negatives (no case or whitespace normalisation; gold and saved drafts survive input changes); RE1; NT1; RE4's single task across stages; three-state exposure capture (VS2); input drift never retracting gold; Q-10 conversation rules; assignment races and ineligible assignees (RA1/RA2, E6). LEDGER:436-448, :690-697, :58, :365-372, :655-669; PKG/contracts.md:372-374; PKG/notifications-integration.md:111-120 Add AC-R4a-14 to -22 (§3.1).
AC-09 Major ACD:30-43, :55, :57, :79-80; IP:247 (L17), :936 (W0) Assumed verification tooling doesn't exist, and L17 has no deliverables in any window. Missing: axe-core integration (listed as future work); visual capture or regression (listed as a non-goal); Firefox and WebKit runs (Playwright runs Desktop Chrome only, while FEAT-023 requires Chromium, Firefox and WebKit); a harness that runs an older binary against canonical data (e2e starts with a fresh database every run). "The Bramble benchmark" is undefined, although FEAT-024 already has an env-gated harness and named datasets. MAIN/docs/platform/e2e-testing-infrastructure/README.md:50, :581-583; MAIN/e2e/package.json (no axe); MAIN/e2e/playwright.config.ts:60-120; MAIN/docs/features/material-3-migration/technical-plan.md:270-277; MAIN/e2e/docker-compose.yml:32; MAIN/src/libs/project-management/SyRF.ProjectManagement.Mongo.Data.Tests/ProjectStatistics/BenchmarkEnvironment.cs Add L17 W0 deliverables with owners and "needed by" releases: axe checks in journey specs and screenshot capture (R1b); multi-persona auth (R1b); Firefox and WebKit smoke projects (R2a); concurrency barriers (M0); mixed-version harness on a persistent database (R0); canonical datasets on FEAT-024's harness pattern (M0). A row whose method needs missing tooling can't be marked passed.
AC-10 Major ACD:51 (AC-ALL-03); §6 :447-448 Permission and blinding tests would run as the wrong actors. The hermetic stack has three users, and "admin" is an application administrator (SyrfAdmin), whose bypasses distort project-permission positives and negatives. The journeys need at least ten personas (four candidates, a reconciler who isn't a candidate, an admin who isn't the owner, an observer, an extra reviewer, a query reviewer). The new users are listed as seed data, not as e2e users. MAIN/e2e/setup/auth.setup.ts:15-43 (roles at :23); MAIN/e2e/helpers/constants.ts:9-25; PKG/notifications-integration.md:75-76 Adopt the §3.5 persona set in auth.setup.ts and SeedDataConstants in W0. AC-ALL-03 tests may use SyrfAdmin only for application-admin paths, and each blinded surface gets one SyrfAdmin negative test.
AC-11 Major ACD §6 :427-460, rule :452-453 The seed rule can't be implemented. The "existing seed reconciliation" only transfers ownership and memberships of the five seed projects. New seed data reaches staging only through an authorised reseed (backup, drop, redeploy), which would wipe tester-created pilot projects, and the seeding doc says new fixtures aren't retrofitted into staging. The seeder writes through legacy services, and R0 makes preview seeding a legacy writer that refuses canonical scopes. MAIN/docs/platform/enhanced-database-seeding.md:45-63, :104-126, :143-156, :410; IP:283-287; MAIN/docs/features/materialized-project-statistics/phase0-benchmark-and-capacity-baseline.md:68 Add an additive, idempotent "seed-if-absent" job keyed by fixed GUIDs. It creates canonical seed projects through canonical commands (admitted via R0), never drops or edits existing data, and runs on preview (/reseed-db) and staging, separately from ownership reconciliation. Adopt the three data tiers in §3.5.
AC-12 Major ACD:159-160, :193, :245, :269; IP:1004-1005 Concurrency criteria name outcomes but no way to force the race. Sequential or timing-dependent tests would pass AC-R2a-06, AC-R2c-05 and AC-R4a-04 without exercising CAS, the fence or the recheck. The repository's own guidance says negative assertions must first prove the operation ran. PLN/review-eligibility-policy.md:1236; MAIN/docs/platform/e2e-testing-infrastructure/performance-plan.md:166-167; MAIN/src/libs/testing/SyRF.Testing.Common/Fixtures/MongoDbReplicaSetTestFixture.cs:37; MAIN/e2e/tests/materialized-statistics-two-api.spec.ts:11 Use a barrier-injection harness (named points such as "after read, before commit") on the replica-set fixture. Each concurrency row asserts the interleaving happened (the loser saw the conflict and kept its draft), runs fixed seeded schedules, and has a two-API e2e variant for claims, drafts and fences.
AC-13 Major ACD:98, :102; IP:280-282 AC-R0-01 can pass with a read-only test while rollback loses data. Study is written by whole-document FindOneAndReplaceAsync. An older binary whose maps merely ignore extra elements strips canonical fields on its next write; only extra-element capture preserves them. The plan offers "tolerant class maps (or extra-element capture)" as alternatives. "Previous production image" is also the wrong subject for R0 itself: the pre-R0 image throws, which is why R0 exists. MAIN/src/libs/project-management/SyRF.ProjectManagement.Mongo.Data/Repositories/StudyRepository.cs:2054-2059; MAIN/src/libs/kernel/SyRF.SharedKernel/BaseClasses/Entity.cs:21 Add AC-R0-06 (round trip through the normal replace path with unknown elements on every type the R0 ADR enumerates) and AC-R0-07 (repeat at each floor step). Require capture, not ignore, for whole-document aggregates. Reword the subject to "the recorded minimum rollback image (R0 or later)".
AC-14 Major ACD:154, :159 AC-R2a-06's two-tab outcome contradicts the drafts contract. "Both drafts are kept" conflicts with "one draft per session" and a "tab lease", so the observable result is undefined. R2a's promise that work is never lost has no durability criterion either: no maximum autosave loss window, and no draft-changes indicator over an explicit version (an SL3 boundary). PKG/contracts.md:250-252; PKG/domain-model.md:91; PKG/open-questions-and-assumptions.md:123; LEDGER:44 Fix E21's observable outcome before F1: the server keeps one draft; the losing tab keeps its edits locally and is offered copy, audited overwrite or discard. Rewrite AC-R2a-06 to match, and add AC-R2a-28.
AC-15 Major ACD:53, :114, :124 Specs for areas this programme changes are excluded from CI. Both configs exclude project-options/** (which holds the ownership-transfer control), the Members & groups dialogs, data-export/screening and question-management/**. AC-R1b-05's "U" check would silently not run. AC-R1a-07 counts the 17 entries in angular.json, but vitest.config.ts also excludes the whole question-management folder. MAIN/src/services/web/vitest.config.ts:25-47 (:33-37, :44); MAIN/src/services/web/angular.json:175-208; MAIN/src/services/web/src/app/project/project-admin/project-options/project-options.component.html:208-226 Add AC-ALL-15 (merge criterion: no excluded spec for changed code, checked by a script against both lists). Amend AC-R1a-07 to name both lists.
AC-16 Major ACD §5 :401-425; IP:717-725, :1014-1017 Pilot exit can't be measured, and there are no entry criteria. PE-01/02 have no detection mechanism, PE-03 has no severity definitions, PE-04 runs synthetic fixtures on pilot data, and there is no minimum exposure, so an unused pilot passes. The plan calls §5.10's conditions "pilot entry and exit criteria", but ACD has no entry criteria (A-09, A-14, A-19, A-21, the target-1 and R4p waits). ACD:405-409; IP:719-725; PKG/open-questions-and-assumptions.md:191, :196, :201, :203 Add PE-06 (exposure) and PE-07 (severity). Rewrite PE-04 against the invariant monitor (AC-ALL-21). Add AC-ALL-20 (telemetry) and PI--nn entry rows (§3.1).
AC-17 Major ACD:68-70, :76 UI-3 can't be verified. It requires new screens to "render correctly on both paths", but there is no M3 component path. The app emits only M2 component themes (light and dark), with M3 system tokens alongside; the M3 component path arrives with FEAT-023's Wave 5 atomic cutover. Until then, UI1 in practice means M3 tokens on M2 components. MAIN/src/services/web/src/global-styles/syrf-theme.scss:140, :176, :427; MAIN/docs/features/material-3-migration/README.md:69-75; MAIN/docs/features/material-3-migration/technical-plan.md:236-244 Rewrite UI-3 so it can be checked statically (roles and public APIs only; no M3 component islands). Add UI-9 (register routes in FEAT-023's route inventory and baselines). Ask Chris (Q1).
AC-18 Major ACD:57, :90, :172, :194, :278 Performance criteria are undersized, use undefined metrics and lack a baseline policy. AC-M0-02 uses 1,000 studies × 3 reviewers × 200 questions; FEAT-024's datasets reach 25,000 studies and 25 reviewers, and the AF2 gate already measures a 1,000-question form at 1,500 ms to first interactive. AC-R4a-13 uses "first meaningful paint" on an undefined "reference laptop". AC-ALL-09 has no endpoint list, no iteration counts and no baseline run, and existing budgets are kept on pomegranate-01, the host where they were measured. MAIN/docs/features/materialized-project-statistics/phase0-benchmark-and-capacity-baseline.md:144-171, :381-389; MAIN/e2e/tests/perf/annotation-form-perf.spec.ts:26-46; MAIN/docs/platform/e2e-testing-infrastructure/performance-plan.md:154-159 Define named canonical datasets on FEAT-024's generator (RV-DS-02/03 mirroring PS-DS-02/03, plus an E28 max-form case). Benchmark a fixed suite of hot-path endpoints with 20 warm-up and 200 recorded iterations, against a baseline run of main on the same host per release candidate. Reuse AF2's client metrics, and state the host for each budget (Q8).
AC-19 Major ACD:55, :79-80; IP:826-827 Accessibility and browser criteria fall short of FEAT-023's matrix. Missing: live regions (autosave, conflicts, status); focus moving to error and conflict summaries; dialog trap and return; forced colours; 400% reflow and short-height layouts; touch targets; a screen-reader matrix; Firefox and WebKit runs. Drag-pairing has a keyboard alternative but no pointer or touch requirement, although native drag never fires from touch on iOS. The ship checklist's "1440/925 px" contradicts UI-6's widths. MAIN/docs/features/material-3-migration/technical-plan.md:279-308; MAIN/docs/planning/non-chromium-browser-compatibility-analysis.md:35 Extend AC-ALL-07 (§3.1), use one width list, require CDK pointer-based drag, and add Firefox and WebKit smoke journeys (Q6).
AC-20 Major ACD:212-225, :332-360; PKG/migration-adoption-rollback.md:196-210 FEAT-011's binding requirements aren't criteria. Its Approved MUSTs and release checklists ("MUST pass before deployment") are mapped in migration §7 but never written as criteria. Missing: reserved field names; all Citation raw fields and immutability; enum ordinals; unique sparse DOI/PMID indexes; per-profile screeningOutcomes[] with no single Study screening status; structured reasons and authority values; pool membership derivable from filter rules; EXP-05; EXP-06; null sourceType handled. MAIN/docs/features/prisma-specification/prisma-constraint-annotations.md:18, :107-111, :146-147, :197-208, :225-228, :243-245, :260-264, :283-287, :294-354 Add AC-P1-09/10, AC-P2-10/14, AC-R3a-18/19, AC-R5b-08 and AC-C1-06. Reuse FEAT-011's validation procedures (:358-519) as automated checks.
AC-21 Major ACD:352-360 AC-P2-01's parity tolerance can't be computed. The datasets are unnamed, the R version isn't pinned, and "99% agreement" has no metric. P2 also lacks: a performance row (FEAT-012 states processing times); canonical enrichment with provenance; scenario 2 (the reviewed study becomes canonical); scenario 4 (cross-stage link only); "not duplicate" never re-queued; audit entries never deleted. AC-P2-06 omits RemovedByAutomation and RemovedOther. MAIN/docs/features/deduplication/service-specification.md:54-57, :254, :383-421, :431-437, :452-485, :523-531, :603, :646-648 Rewrite AC-P2-01: pinned ASySD commit; named datasets (licences UNVERIFIED); golden outputs committed; identical AutoConfirmed groups; ProbableDuplicate pair-set F1 at or above the agreed value; published sensitivity and specificity. Fix AC-P2-06 and add AC-P2-11 to -14.
AC-22 Major ACD:49-50; PKG/contracts.md:49-51 Flag combinations and "fail closed before editing" are untested. Only "release flags off" and "legacy project" are covered. Nothing tests per-project admission combined with environment-wide flags owned by other programmes. An admitted canonical project opened where AF2 is off or not admitted could render AF1, accept input, and then be refused by R0. PKG/contracts.md:49-51; PKG/source-status-inventory.md:172-175; PKG/open-questions-and-assumptions.md:52 Add AC-ALL-17, and per release a supported-flag matrix (release flags × admission × external flags) listing the cells covered by I and E tests.
AC-23 Major ACD:52; PKG/migration-adoption-rollback.md:133-180 Rollback containment has no criterion, and the rehearsal rule is uniform. No criterion says what users see or can still do under read-only containment after canonical writes (U28 is only a UI validation). AC-ALL-04 requires a staging image-rollback rehearsal for every release, including R1b (no data) and R5c (nothing persisted), on a shared, auto-deployed staging. PKG/migration-adoption-rollback.md:139, :154, :166-180; PKG/open-questions-and-assumptions.md:174 Add AC-ALL-18. Scope AC-ALL-04 to releases that persist data or add fields, using the mixed-version harness. Keep staging rehearsals for R0, the floor steps, R2a, O2 and each R6 wave.
AC-24 Major ACD:51; PKG/contracts.md:424-430 Blinding has no oracle. "No hidden identities or answers" isn't backed by a disclosure matrix (persona × surface × element), and SignalR payloads, inbox detail, impersonation and forged IDs aren't named, although research A5 and A26 require them. PLN/screening-specialised-annotation-research.md:1250, :1271; PKG/source-status-inventory.md:187-190, :381-382 Make the disclosure matrix an F1 artefact. Add AC-ALL-19: a table-driven probe suite covering APIs, exports, SignalR, inbox and email, including SyrfAdmin, impersonation and forged-ID cases.
AC-25 Major ACD:47-61, :81, :409; PKG/open-questions-and-assumptions.md:206 Uniform gates are costly and sometimes vacuous. Every release carries the same heavy gates whatever its risk: staging rehearsal, benchmark, five testers, pilot exit, Chris's preview acceptance per screen, and go/no-go. Over about 30 releases this concentrates on Chris and the testers, and AC-ALL-07/08/11 are vacuous for R0. ACD:47-61; IP:64-88 Tier releases (T1 write path, T2 UI over existing data, T3 derived or read-only) with the AC-ALL subset for each tier (§3.2). Chris accepts per release on one staging build, plus per-PR acceptance for high-risk surfaces (Q2).
AC-26 Minor ACD (31 rows verified by E only); ACD:49 Server rules are verified by UI tests alone. Thirty-one criteria rely only on E, including server rules such as AC-R3c-02's commit recheck, AC-R4a-02's pairing history, AC-R3b-04's no re-invitation and AC-P1-04's "unknown stays unknown". E2E runs only on labels, inside a 45-minute job and a 35-minute Playwright budget on one listener. "Touched flows" names no spec set. rg of rows ending "| E |"; MAIN/CLAUDE.md:109; MAIN/e2e/playwright.config.ts:30-40; MAIN/docs/platform/e2e-testing-infrastructure/performance-plan.md:145-153 Persisted-state rows need an I or C test. E is for UI behaviour: one release-journey spec per release plus targeted negatives. Each release lists its flags-off spec set and the e2e minutes it adds.
AC-27 Minor ACD:74 (UI-1) UI-1's checks don't look for hard-coded colours. check:theme-migration scans M2 APIs, Bootstrap and private Material paths only, and there is no stylelint configuration. MAIN/src/services/web/scripts/check-theme-migration.mjs:23-50 Add UI-10: a guard spec over the programme's new folders, modelled on the existing repository guard specs.
AC-28 Minor ACD:75, :78, :81 UI-2, UI-5 and UI-8 are subjective. UI-2 has no pattern inventory; UI-5's handoff is unlinked; UI-8's "materially changed" is undefined. MAIN/docs/features/material-3-migration/handoffs/project-navigation-drawer/README.md Add a pattern checklist, link the handoff, and define "materially changed" (new route, layout or primary flow).
AC-29 Minor ACD:411-425; PKG/open-questions-and-assumptions.md:139-175 User testing has gaps. No tasks for R2b, R2d, R3c, R4b, R4p, R5c, P1 or P2; no rubric for "explain" tasks; U1–U29 have no pass bar. Same Add the tasks in §3.1, a rubric for each explain task, and a pass bar for each U.
AC-30 Minor IP:118-152, :445, :819; ACD:408 "Invariant fixtures" are referenced but never listed. Same Add INV-01 to INV-12, one executable fixture per invariant in IP §2, run from R2a onward.
AC-31 Minor ACD:156, :217; PKG/contracts.md:205-208 Shared fixtures have no location, format or cross-language harness. This applies to the applicability fixtures and the C6 table; fixed lists also miss combinations. PLN/review-eligibility-policy.md:1067-1069 Put a JSON corpus under src/libs/testing, consumed by xUnit theories and Vitest describe.each, plus a seeded generated corpus diffed across the .NET and AF2 evaluators.
AC-32 Minor ACD:195 and others; PKG/notifications-integration.md:157-179 Notification correctness with flags on is untested. Only flags-off behaviour is checked. One item per occurrence, the two-event lifecycle, payload shaping and email suppression are in C15's test list but not in ACD. Same Add AC-ALL-22.
AC-33 Minor ACD:386-390; IP:95-99 GA can flip the production default without a production pilot. R2–R4 pilots may have run on staging and preview only. Same Q4; add AC-GA-06.
AC-34 Minor Various Precision defects in individual criteria:
• AC-R5a-02 "identical" ignores generation metadata (:313).
• AC-R2a-09 says "or", leaving the behaviour open (:162).
• AC-R2a-10 doesn't define "published" (:163).
• AC-R3a-07 omits D1's missing→Allow (:222).
• AC-R1b-08 can't fail before WP9 (:127).
• AC-R2c-06 has no crash-and-resume check (:194).
• AC-O2-01's "read-only" is unproven (:374).
• AC-R3d-01's parity checklist has no author (:256).
PLN/review-eligibility-policy.md:985 (D1) Rewrites in §3.1 (AC-R5a-02r, AC-R2c-10, AC-O2-01r, etc.); mark AC-R1b-08 as conditional.
AC-35 Minor ACD:89; PKG/contracts.md:127 Research cases are only partly referenced. AC-M0-01 lists A1, A3 and A20–A23, while C1 also lists A7 and A11. A2, A5, A6 and A8 have no criterion; A12–A16 appear only in C6. PLN/screening-specialised-annotation-research.md:1242-1273 Map the cases as in §5.3.

3. Improvements

3.1 Criteria to add or rewrite

ID Criterion Verified by Source
AC-ALL-14 Merge criterion: every PR keeps AC-ALL-01/02 true. Unit and integration suites always run; the release's named flags-off e2e spec set runs (run:e2e-full) when the PR touches a listed flow. U, I, E Flags default off
AC-ALL-15 Merge criterion: no spec for a changed component, route or service is excluded in angular.json or vitest.config.ts (checked by a script). U Repository rule: all code changes need tests
AC-ALL-16 Merge criterion: tests evidencing a criterion carry its ID ([Trait("AC","AC-R2a-03")], @AC-R2a-03, Vitest describe prefix), and the traceability check passes. G, U AC1
AC-ALL-17 An admitted canonical project opened through any surface that can't write canonical data (AF1, legacy grid or editor, AF2 not admitted, a flag missing) shows a typed "not available here" state before accepting input; no legacy write is attempted. I, E Contracts stance 6; Q-25
AC-ALL-18 With the flag off or admission removed after canonical writes: every surface the release added is read-only with an explanation; no edit control is enabled; owners can still retrieve drafts and versions; canonical exports work. E, R Migration §5; U28
AC-ALL-19 A disclosure-matrix probe suite passes for every new endpoint, export column, SignalR message and notification kind, including SyrfAdmin, support edit mode and forged foreign IDs (refused before any mutation). C, I C10; A5, A10, A26
AC-ALL-20 Telemetry for pilot-exit failures (draft-save failures, stale conflicts by type, refused legacy writes, typed errors, projection mismatches) is on a staging dashboard, with thresholds agreed at the freeze gate. R, G PE-01/02
AC-ALL-21 A read-only invariant monitor reports zero violations on admitted projects. Checks: every receipt has its version; every draft's base exists; no lost versions; projection equals recomputation; one qualifying contribution per reviewer, study and form; the gold pointer is valid. I, R Invariants 1, 2, 8
AC-ALL-22 With notificationInbox on: one item per recipient per event occurrence; a second lifecycle event is captured; payloads are generic; detail is reshaped under fresh authority ("Related item unavailable" after revocation); email goes to Mailpit only. I C15
AC-ALL-07 (extended) Adds: polite live regions for autosave, conflicts and status; focus moves to error and conflict summaries; dialogs trap and return focus; forced colours; 400% reflow at 320 CSS px and short-height layouts; touch targets; screen-reader matrix (NVDA with Firefox and Chrome, VoiceOver with Safari) on the five key surfaces. A FEAT-023 matrix
UI-3 (rewritten) Only --mat-sys-*/--syrf-* roles and public component APIs; no M3 component islands; check:theme-migration passes. U UI1, FEAT-023
UI-9 New and changed routes are registered in FEAT-023's route inventory and visual baselines. G, V UI1
UI-10 No literal colours or var(--role, #fallback) in new feature styles or templates (guard spec). U UI1
UI-11 Copy and navigation use the C17 terms: "Design", "question templates", "Library" only for Study Management, Save vs draft, Needs updating vs outdated. U, D Q-13; C17
AC-R0-06 The minimum rollback image loads a Study, Project or SystematicSearch with unknown elements on each type the ADR enumerates, edits an unrelated field, writes through its normal replace path, and every unknown element survives. I B-01; R0
AC-R0-07 Floor steps before R3a, P1, C1 and O1 rerun AC-R0-01/05/06 for their newly extended types. I, R IP:294-295
AC-R0-08 Writer census: a test fails if any path writes review data in a canonical scope without passing the single ownership choke point. U, I Research A27
AC-R1b-09 Transferring ownership to a non-member is refused; the former owner loses owner-reserved activities on their next request. I PM1, SEC1
AC-R2a-20 Concurrent commits get strictly increasing per-project sequence numbers allocated in the transaction; commit order and sequence order agree. C E25
AC-R2a-21 IDs are server-minted or validated; a client ID belonging to another scope is refused before mutation. I E27, A5
AC-R2a-22 Publishing above the maximum form size is refused with a typed error; AC-R2a-19 runs at that maximum. I, B E28
AC-R2a-23 A later code change to system questions doesn't alter a published version's pinned snapshot (form, validation, export); SystemQuestionVersion v0 and v1 pin different structures. C E24
AC-R2a-24 Composed form versions include every ancestor; removing an ancestor while a descendant remains is refused. U, I QM-09; C4
AC-R2a-25 Binding a form version to a second stage is refused, with an explanation, until R2b. I A-21
AC-R2a-26 Each revision records source stage, step, stage-settings version, question version and real actor, unchanged on reuse; each session version records route stage, form version and the revisions submitted together. C PV1
AC-R2a-27 An order-only change creates no version, doesn't remove Complete and doesn't appear in exports. C C5
AC-R2a-28 A draft edit persists within the debounce window (PROPOSAL 2 s) and survives reload, tab close and a second device; a draft-changes indicator over the current explicit version is announced. E, A SL1, SL3
AC-R2a-29 After a backup restore: versions, drafts, receipts, commit sequence and projection agree. R E31
AC-R2a-30 Deleting a reviewer's account applies E32 to their revisions, drafts, exposure events and receipts; AC-ALL-21 still passes. I E32
AC-R2a-31 Canonical sufficiency uses the form target; the per-stage override (#3732) has no effect on canonical forms. I SF2
AC-R2b-07 Concurrent Complete through stages A and B on one session yields one completed version and one contribution; the loser gets a typed conflict and keeps its draft. C SF2, A22
AC-R2b-08 My studies, incomplete studies and the no-work page return identical counts through the API for bound stages (backs the E-only AC-R2b-05). I SF1
AC-R2c-10 Killing phase 2 mid-run and resuming applies each transition exactly once; qualification follows the recorded policy throughout. I E22
AC-R2c-11 Qualifying-contribution counts per version match the expected table for each policy combination in FX-PUB. C FV3
AC-R2c-12 "Why it changed" and "What reviewers need to do differently" are stored per version, shown beside affected questions and kept in history; missing guidance never blocks. I, E VU2
AC-R2c-13 FX-PUB-10K: 10,000 sessions across v1–v3 and two stages, split 60/30/10 completed / saved-incomplete / draft-only, 5% with drafts over explicit versions; phase 1 stays within MongoDB limits and the manifest is paged. B E22
AC-R2d-09 Before changing a shared answer, the reviewer is told a new version will be created and which sessions will show "contains outdated annotations"; the same QuestionId under another entity or branch stays separate; legacy conflicting duplicates show a conflict marker. I, E SF5
AC-R3a-11 Complete-and-Include commit together; an injected failure commits neither; a retry after an unknown commit returns the original receipt; navigation waits for acknowledgement. C, E A8
AC-R3a-12 The default profile's outcome equals the characterised legacy result for single, manual dual, automated dual and custom thresholds, across missing, insufficient, conflict, included, excluded and corrected cases. C A2
AC-R3a-13 Skip records no decision, completion or PRISMA event; AND/OR groups follow their truth table; terminal-on-Exclude ends only its scope; a compulsory step blocks dependent work. C, E Amendment A; RC6
AC-R3a-14 Advanced own-Include admits an own-Include reviewer while the collective decision is Pending, and blocks after collective or personal Exclude; a self-cycle is rejected; changed published policy is rechecked at admission. C, E DP7; access-policy fixtures
AC-R3a-15 VS1: with the step policy off, candidates never receive accepted answers (API or UI); with it on, they see the snapshot; own previous answers are always visible; other candidates' answers never are; stage default applies unless overridden by the step. I, E VS1
AC-R3a-16 The eligibility test-plan layers pass for legacy and migrated combined stages, or each changed expectation cites Q-24. C, I, E Eligibility policy :1204-1241
AC-R3a-17 Decisions captured since admission appear as canonical history with original times, authors and "captured legacy" provenance; none lost or duplicated. I E26
AC-R3a-18 One current outcome per study and profile, with route provenance, an authority from the approved set and a structured reason with coverage; no single Study screening status; no free-text-only reasons. C Amendment H; FEAT-011 phases 13 and 15
AC-R3a-19 Pool entry for early-stopped and batched reviews matches FX-PRISMA-03; re-evaluating the recorded filter and profile versions reproduces membership. C Amendment A; FEAT-011 phase 14
AC-R3a-20 Partial combined-step reservations behave as the C7 contract says. C E5
AC-R3b-09 An ordinary study-fact question and a profile question with identical wording never share answers; two profiles copied from one template never share answers or history. C DP4
AC-R3b-10 Reason collection off with reason reconciliation on behaves as E11 says: no invented reasons, no block. C DP5, E11
AC-R3b-11 A DP2 correction creates a new revision, keeps the Exclude in history, re-evaluates eligibility and creates no invitation. I DP2
AC-R3c-08 A change affecting a shared form bound to two Completed stages needs approval for each and commits to both or neither; changed inputs invalidate the approval; approval never grants the requester authority; approvals expire. I LC1; Q-02 pending
AC-R3c-09 A Completed stage's bindings can't be edited outside the reopen flow. I RX2
AC-R3c-10 A reviewer having no available work never marks a stage complete. I A14
AC-R3d-06 With extraction off, setup adds no cohort, outcome or experiment types; templates use verified legacy identities. I, E TC1
AC-R3d-07 An empty project creates no PRISMA records; a double submit creates one project; living-search and ML stay disabled. I, E Setup fixtures
AC-R4a-14 Incompatible, saved-incomplete-current, withdrawn and ineligible sessions are never candidates and never count; agreeing candidates produce no gold until the reconciler completes. C SF4 delivery list; invariant 7
AC-R4a-15 A study × form task is listed once whether reached from A or B; completing it satisfies only that form's requirement in each stage. I, E RE4
AC-R4a-16 Complete succeeds with no explanation even when gold differs from every candidate. I, E RE1
AC-R4a-17 Candidate notes are unchanged by reconciliation; a copied note keeps its author and source. I, E NT1
AC-R4a-18 Texts differing only in case or whitespace don't prefill; gold and the reconciler's saved draft survive input changes. C RE5
AC-R4a-19 Rendering an accepted revision records one exposure per session version and revision; a lost report leaves the contribution "informed or unknown". C, E VS2, C3
AC-R4a-20 A new candidate, a Save after Complete or a Fix after gold sets "inputs changed" and never retracts gold. I C9
AC-R4a-21 With conversations on: one-to-one threads; a reconciler who reviewed the study can't start one; an extra reviewer is excluded until they return; a questioned session gets an exposure marker; context links are read-only. I, E Q-10
AC-R4a-22 Assigning to an ineligible reconciler is refused; an expiry/start race has one outcome; changed defaults affect only new assignments; two reconcilers never hold one task. C, I RA1, RA2, E6
AC-R4b-07 Rejecting a concern with an empty explanation succeeds. I QY5
AC-R4p-05 Two Excludes with differing must-agree reasons: the collective Exclude vetoes dependent work while reason reconciliation is pending; extra votes never resolve reasons; adjudication adds no vote; a replacement retires the old adjudication. C RX1, A25
AC-R4c-04 Every candidate's series stays reachable with provenance, including a fourth candidate, without truncation. C, E SF4
AC-R5a-02r Two exports at one watermark have identical data files and checksums; manifests differ only in generation metadata. I EX1
AC-R5a-07 Exports carry per-question reconciliation status and version references. I EXP-01/02
AC-R5c-05 The CSV equals the on-screen figures; a Reconcile holder without the agreement capability is refused. I, E AG1; EXP-03
AC-R5b-08 PRISMA JSON/CSV export with manifest; an unknown or null source type is its own group. I EXP-05; FEAT-011 phase 16
AC-P1-09 Architecture test: no Study status/state/lifecycle, no SystematicSearch type/category, no clash with planned PRISMA names. U FEAT-011 release 1
AC-P1-10 Citations carry every raw field; no code path updates a Citation. U, I FEAT-011 phase 12
AC-P2-01r Pinned ASySD commit; named datasets; committed golden outputs; identical AutoConfirmed groups; ProbableDuplicate pair-set F1 at or above the agreed value; published sensitivity and specificity on labelled sets. C L; Q-37
AC-P2-06r The admission service and pool filters exclude Duplicate, Merged, PendingDuplicateReview, PendingDedupCheck, RemovedByAutomation and RemovedOther. C FEAT-012 §12
AC-P2-10 Unique sparse DOI and PMID indexes; lifecycle enum ordinals exactly as specified. I FEAT-011 release 3
AC-P2-11 Enrichment follows FEAT-012 §6 with per-field MetadataProvenance; cross-project data is bibliographic only. C FEAT-012 §6
AC-P2-12 Scenario 2 (reviewed study is canonical); scenario 4 (cross-stage link, reinterpreted for shared forms by L's ADR); "not duplicate" never re-queued; Defer changes nothing; audit entries never deleted. C FEAT-012 §7, §8, §10.3
AC-P2-13 On Bramble: Stage 1 per 1,000-record batch within budget; Stage 2 on the 80k-citation set under 1 hour. B FEAT-012 §2.1
AC-P2-14 Dedup report export with canonical mappings and confidence scores. I EXP-06
AC-C1-06 Classification fields don't use lifecycle enum names. U FEAT-011 release 2
AC-C1-07 TC1 templates use verified identities (E14); system types attach only when an extraction feature is used. I TC1
AC-C1-08 Moving from category tabs to entity types keeps every question, answer and guidance reachable. I, E C1 scope
AC-O1-06 Direction is never derived from numeric type and doesn't act as a validator. C OC2
AC-O1-07 New-shape data round-trips through export and import. I O1 scope
AC-O2-01r The dry-run runs under a read-only database role, so any write attempt fails. I MIG1
AC-R6-07 Adopted questions count as published; no automatic gold promotion; E10's labels appear in manifests and exports. R, C QD1, Q-29, E10
AC-GA-06 At least one production opt-in pilot per family (R2, R3, R4a) meets PE-01 to PE-07 (pending Q4). G Production safety
PE-04r The invariant monitor reports zero violations on the pilot projects. R —
PE-06 Minimum exposure (PROPOSAL): at least 2 projects, at least 3 reviewers each, at least 50 studies with two or more completed contributions, at least 10 working days. S —
PE-07 Severity definitions. Sev-1: lost or silently changed work, a disclosure leak, or a wrong authoritative outcome. Sev-2: a wrong count or status, or a blocked task, with a workaround. Sev-3: cosmetic. G —
PI--nn Pilot entry: forms are target-1 or their reconciliation can wait (R2a–R2c); one stage per form (A-21); one form per entity category (A-19); no proportional shares on shared forms (A-09); Stop on a cross-stage profile accepts waiting for R4p; Q-28's interim rule. G IP §5.10

User-testing tasks to add: - R2b: open the same study from two stages and explain the count. - R2d: Fix an outdated session. - R3c: approve a change to a Completed stage, predicting its effect. - R4b: raise a query and find its outcome. - R4p: adjudicate a decision with conflicting reasons. - R5c: explain independent versus informed figures. - P1: enter external dedup counts and explain the warning. - P2: resolve a duplicate pair.

3.2 Definition of ready and done

PR ready: - The PR lists the criterion IDs it advances. - The contracts it consumes are frozen, or it builds against a published fake. - Its UI validation has passed, and copy comes from the C17 contract. - Its flag is declared in env-mapping, default off. - Its test plan gives a layer and fixture for each criterion; the fixture files exist or ride in the PR. - No behaviour depends on an unanswered question. - Excluded specs in the areas it touches are identified for re-enabling.

PR done (merge): - AC-ALL-14 to -16 pass. - Conformance suites for touched contracts are green. - The SonarCloud new-code gate passes. - No new lint suppressions or test exclusions; check:theme-migration and check:contrast pass. - Docs, user guide and any ADR are in the same PR. - A Claude review has passed on the exact head, with no unresolved threads. - The traceability file is updated.

Release ready (before build): - The freeze gate has passed (ADR, DTOs, fake, suite). - Its PROPOSAL thresholds are confirmed, and its pending rows are re-confirmed from the answers (otherwise that behaviour is descoped). - Its fixture files and benchmark dataset exist. - Its seed project is designed. - Its UI validations have passed against their pass bars. - Pilot entry criteria, the pilot plan (projects, testers, duration) and telemetry are agreed. - The minimum rollback image and rehearsal plan are recorded. - Its tier (T1, T2 or T3) is assigned.

Release done (activation for pilots): - A release-candidate commit and image SHAs are deployed to staging. - All automated rows pass on that commit (CI, label e2e, Bramble benchmark report). - The rehearsals and containment check required by its tier are done. - Seeds are deployed additively. - There is a staging acceptance note (who, when, which SHA). - Chris has done a walkthrough of the new screens. - PE-01 to PE-07 are met, with monitor evidence, and user testing is complete. - X-join go/no-go is recorded before any production pilot. - The acceptance record is committed. - Chris gives go/no-go.

Tiers: - T1 (R0, R2a–R2d, R3a–R3c, R4a, R4p, R4b, R4c, P1, P2, O1, O2): all AC-ALL. - T2 (R1a–R1d, R3d, C1, AL1): no image-rollback rehearsal; three testers. - T3 (R5a, R5b, R5c, C2): automated AC-ALL plus one walkthrough.

3.3 Traceability matrix format

Keep one machine-readable file (later moved into a permanent docs/features/ location), for example:

- id: AC-R2a-02
  release: R2a
  source: [SL2, SL3]            # ledger IDs | RECOVERED:<ref> | PROPOSAL | Q-26 | A-14
  status: confirmed             # confirmed | pending-Q-xx | assumption-A-xx
  traces: {research: [A1], qmv2: [QM-07, FORM-04], feat011: [], feat012: [], conformance: [C5-T01]}
  fixtures: [FX-VERSIONED-01]
  methods: [I, C]
  tests: ["dotnet:FormSessionLifecycleTests.SaveAfterComplete_RemovesQualification", "playwright:@AC-R2a-02"]
  evidence: {image: <sha>, ci_run: <url>, record: R2a-acceptance.md#ac-r2a-02}

A script generates three views from it: decision → criteria (fails when a ledger or §1.11 ID has none); criterion → tests (fails at the ship gate when a row has none); and the per-release acceptance record. A docs CI job also fails when a test tag names an unknown criterion ID.

3.4 Test strategy

Layer Proves Tooling Runs Notes
C conformance Contract invariants, C6 table, applicability, PRISMA evidence assertions xUnit theories plus a shared JSON corpus; Vitest for AF2 parity; seeded generated corpora (no new dependency, or FsCheck/fast-check by ADR) Every PR Differential tests: C# vs TypeScript evaluator; legacy vs canonical on the same scenario
U/I Domain rules; transactions, CAS, fences, permissions Replica-set fixture (mongo:8.0), barrier harness Every PR Each concurrency row asserts the interleaving happened
Compatibility Old binary tolerates new data Mixed-version harness (start N, write, start minimum image, run legacy flows, replace round trip) T1 release candidates, floor steps Replaces most staging rehearsals
E journeys UI-observable behaviour Playwright, persona set, scenario builders Label lane; about 4 min per release Firefox and WebKit smoke for reviewer and reconciler journeys
A, V Accessibility; design axe inside journeys; screenshots at named states (light/dark, 1440/390) With E Screenshots reviewed by a human, not pixel-gated, as FEAT-023 does
B Budgets FEAT-024 harness pattern (env-gated, named datasets, 20 warm-up and 200 recorded iterations); AF2 perf spec Bramble (backend); AF2 client budgets on the host where they were measured Baseline run of main per release candidate
R, S, PE, UT Rehearsal, acceptance, pilots Invariant monitor, telemetry Per tier Monitor reused for R6 parity checks

An efficiency rule: a release's first PR lands its criteria as skipped tests tagged with their IDs, plus the fixture files. Progress then equals rows turning green, and parallel lanes and agents share one target.

3.5 Test-data plan

Tiers: - T1, code fixtures. Builders in MAIN/src/libs/testing/SyRF.Testing.Common plus a JSON corpus: FX-PRISMA-01…08, FX-ACCESS-01…10, FX-LIFE-01…10, FX-SETUP-01…13, FX-RX1, FX-RE5, FX-SF4, FX-SF5, FX-DP4 (the ledger's delivery lists), FX-PUB-10K and FX-LEGACY (cross-stage overwrite, legacy reconciled answers, untouched outcome defaults, missing timestamps, SystemQuestionVersion v0 and v1). - T2, e2e scenario builders. API-driven, with per-test unique IDs as the existing factory does. Canonical states are built through a test-only scenario endpoint in the e2etest environment that calls canonical commands, never raw database writes. - T3, human-acceptance seeds. Created by the additive "seed-if-absent" job (AC-11). - T4, benchmark datasets. RV-DS-01…05 on FEAT-024's generator and seeds (worst case: PS-DS-03's 25,000 studies, 25 reviewers and 8 stages) plus an E28 max-form case. - T5, ASySD golden outputs. Generated once in a pinned R container.

Personas: - owner (not an application admin) - admin who isn't the owner - reviewers A, B, C and D - reconciler (never a candidate) - extra reviewer - observer (view only) - query raiser - query reviewer - SyrfAdmin (application-admin paths only) - support impersonator

4. Questions for Chris

  1. What does "uses Material 3" (UI1) require before FEAT-023's cutover? Today the app emits only Material 2 component themes, with M3 tokens alongside. Recommendation: ship on M3 roles over the current components (no mixing), register every new route for FEAT-023's cutover baselines, and don't make R1/R2 wait for Wave 5.
  2. Should you accept each new screen on its PR preview, or once per release? Recommendation: a per-release walkthrough on one staging build, plus per-PR acceptance only for the publication dialog, reconciliation workspace, stage designer, Members & groups and guided setup.
  3. Do you approve pilot severity definitions and minimum exposure (PE-06, PE-07 values)? Recommendation: approve, with overrides allowed at each freeze gate.
  4. Is a production opt-in pilot required before GA? Recommendation: yes, one each for the R2, R3 and R4a families (AC-GA-06).
  5. Staging seeds: additive "seed-if-absent", or periodic authorised reseeds that wipe tester pilots? Recommendation: additive.
  6. Do Firefox and Safari (WebKit) journeys, and touch-capable drag, belong in acceptance for the reviewer and reconciler screens? Recommendation: yes, as smoke journeys.
  7. Who are the testers, and do you accept a tiered panel? Recommendation: 5 testers for R2a, R2c, R3a, R3d and R4a; 3 elsewhere.
  8. Which scale does SyRF commit to support? Recommendation: use FEAT-024's 25,000-study dataset and AF2's 1,000-question form (or an explicit E28 maximum) as the benchmark sizes, and resize AC-M0-02, AC-R2a-19, AC-R2c-06 and AC-R4a-13.
  9. For Q-37, what does ASySD parity mean? Recommendation: both parity with pinned R package outputs and the published sensitivity and specificity; identical AutoConfirmed groups, ProbableDuplicate pair-set F1 of at least 0.99, and 80k citations in under an hour on Bramble.

5. Coverage gaps

5.1 Confirmed decisions with no acceptance criterion: PV1, VS1, RE1, NT1, TC1, QY5 (within QY4–QY7), RE4's single task across stages, Q-10, Q-13, and amendment H (approved under Q-06a). Amendment A's "entering screening" fixtures have none either. IP1 and AC1 are meta decisions and need none.

5.2 Decisions with only partial criteria (what's missing): - SF2: target comes from the form. - SF4/RE3: the negative cases. - SF5: pre-change warning; entity separation; legacy conflicts. - SL3: draft-changes indicator. - PV2: settings-version pinning. - VS2: capture at R4a. - EW1: step override; preserve-only alternative. - FV1: an explicit "adding a question creates v2" row. - FV3: counting under each policy. - VU2: the two fields. - RA1: no simultaneous reconcilers. - RA2: eligible assignees only; default changes affect new assignments only. - RA5: capability check. - AG1: Reconcile doesn't imply access. - UA1: blank isn't N/A; reconciler decides optional blanks. - DP3: reasoning never exposes others. - DP7: advanced option scene. - RX1: agreeing decisions with conflicting reasons. - RX2: frozen bindings. - OC2: not derived from numeric type. - PM1: ordinary admin defaults. - OPS1: overview and settings DTOs. - QD1: definition of "published"; adopted questions. - Q-09: harvested #2224 tests. - Amendments G and I: per-project adoption and rollback rows.

5.3 Research A1–A28. - Covered: A1, A3, A20–A23 (AC-M0-01); A27 (AC-M0-03, AC-R0-02); A28 (AC-R7-01/02). - Partial: A4 (AC-R3b-02), A9 (AC-R3a-04/07), A10 (AC-R2a-16), A11 (AC-R2b-02; C1 only), A12–A16 (C6 list only), A17/A18 (AC-R6-03/06), A19 (AC-R0-01/05), A24 (AC-R4a-04/09), A25 (AC-R4p-01), A26 (AC-R3a-06). - None: A2, A5, A6, A7 (C1 list only), A8.

5.4 QM v2 tracker groups. - Covered or partial: QM-01..05 (AC-R2a-14, R2c); QM-12/13; FORM-03..05; FORM-08 (AC-ALL-03); PGRP-01..04; RECON-01..08, -12..16; EXP-02. - Missing: ARCH-03; QM-09 (ancestors); QM-11 and QM-14 (question version history, diff, badge); FORM-06 (client budget beyond the server p95); FORM-07 (unsaved-changes guard under autosave); RECON-10 (RE4); RECON-11 (VS1); RECON-17 (RE1); SCR-07 (bypass); FILT-04; MIG-01 and MIG-09 (R6); EXP-01, EXP-03 (CSV), EXP-04, EXP-05, EXP-06; PRISMA-07.

5.5 FEAT-011 MUSTs (35 MUST, 12 MUST NOT). - Covered: no annotation reconciliation writing screeningOutcomes (AC-R4p-01); 34 fields derivable (AC-R5b-07); dedup audit and no auto-merge (AC-P2-03/09). - Partial: nullable source type and name (AC-P1-01); Reconcile covering both kinds (AC-R4p-04); pool entry (AC-R3a-10); pool exclusion (AC-P2-06). - Missing: reserved names (phases 7 and 11); no study-level reconciliation status; pmPublication indexes; all Citation fields and immutability; enum ordinals; per-profile screeningOutcomes[] and no single status; structured reasons and authority values; pool derivability; MIG-13 inference (R6); EXP-05; EXP-06; null sourceType handling. None of the three release checklists appears in ACD.

5.6 FEAT-012. - Covered or partial: invariants 1–4; queue display; box 3 equation; reversal; merge with remap. - Missing: §2.1 performance; §3.5 batching; §6 enrichment and provenance; scenarios 2 and 4; Defer and "not duplicate" behaviour; §10.3 retention; the full §12 exclusion set.

5.7 Other sources. - The eligibility test plan (configuration, capacity, concurrency, UI/API parity, revocation and cross-pod layers) is unreferenced. - From FEAT-023's matrix, these are missing: Firefox and WebKit runs; 400% reflow; forced colours; live regions; and verifying light, dark and system modes (UI-7 asks only for light and dark screenshots). - The steps handoff's worked scenes "two forms with targets 2 and 1", "VS1 accepted-gold display" and "extraction-only step offers annotation or Skip" have no criterion (PLN/review-steps-prototype-handoff.md:567, :573-574, :577-578).