Temporary planning review record. The report below is reproduced verbatim as returned by the independent read-only reviewer (Plan agent, Fable model, launched 3 October 2026 about 14:35 BST). Only this front matter and note were added. Line numbers refer to the package as it stood when the reviewer read it. Resolutions are in the round-2 resolution matrix.
UX review of the integrated review programme plan (dimension: UI and UX across the whole application)¶
Reviewer: independent, adversarial, read-only. 3 October 2026.
Path legend (all absolute):
- PKG = /home/chris/workspace/syrf/pr/pr3617.research-screening-as-specialised-annotation-gxgahs/docs/planning/integrated-review-plan-2026-10/
- PLAN = /home/chris/workspace/syrf/pr/pr3617.research-screening-as-specialised-annotation-gxgahs/docs/planning/
- MAIN = /home/chris/workspace/syrf/main/; WEB = MAIN/src/services/web/src/app/
- V10 = /home/chris/.codex/visualizations/2026/09/23/01a0cbec-0c32-7103-ab5a-bfc05665deb7/syrf-v10-review-2026-10-02/source/design_handoff_syrf_v10/
- RP4 = /home/chris/workspace/syrf/handover/2026-09-21-stage-review-design/design_handoff_stage_review_page/README.md
- M3 = MAIN/docs/features/material-3-migration/
- QM2 = MAIN/.local/qm-designer-handover/
- SM = MAIN/docs/features/study-management/
Current-UI claims were checked in MAIN (web at the state after c59d9d0f1; I did not re-verify the HEAD SHA). Anything I could not check is marked UNVERIFIED.
1. Verdict¶
The plan is strong on data contracts and correctness acceptance, but it treats user experience as a list of prototype checks (U1–U29) plus a per-screen Material 3 gate. It does not contain what a plan that "maximises success and efficiency" needs: a research plan with real systematic reviewers, measurable efficiency and comprehension criteria, validations scheduled before the builds that depend on them, a copy and terminology artefact, a design system of record for the many new patterns, an accessibility harness that exists in the repository, and an information architecture for the new work queues. Several of these are not omissions of detail; they are gaps that will decide whether reviewers accept the new workflow or route around it. The five changes that matter most: (1) add a UX research and metrics plan (baseline study of today's screening and annotation now; prototype tests per freeze gate; summative tests per release; task time, actions per decision, comprehension, SUS) and make the metrics ship-gate criteria; (2) make reviewer efficiency a first-class acceptance target, including the keyboard screening path that is specified in the stage-review design but not built on main and currently parked in an external join; (3) schedule every U validation in a window before its consuming build and make the R3 validations (U2, U5, U7, U12, U17, U18) exit evidence of F3/F5; (4) deliver a copy deck and glossary at F1 with a mechanism (typed message constants, user-guide glossary parity) and decide the Save/Complete/draft verbs now; (5) define the design system of record and a handoff template, decide the FEAT-023 cutover sequence relative to R2a, and replace "Chris accepts every screen on every PR preview" with a two-tier design-QA process that Chris samples at release level.
2. Findings¶
| ID | Severity | Location | Finding | Evidence | Recommended change |
|---|---|---|---|---|---|
| UX-01 | Blocker | PKG/acceptance-criteria.md §1 l.39 (method UT), §5 l.401–425; PKG/integrated-plan.md §5.4 l.443–446, §9 item 8 l.1018–1022 |
There is no usability research plan. "User testing" is a verification method whose evidence is "session notes against the tasks in §5"; the only metric is PE-05 ("4 of 5 testers complete each task without help", PROPOSAL). Nothing says who the testers are beyond "CAMARADES testers", how many per role, whether any are external SyRF users (student projects, multi-site reviews), whether sessions are moderated, what is recorded, or when formative research happens relative to the freezes. There is no baseline measurement of today's screening and annotation, so no release can show it improved anything. Pilot projects are synthetic seeds (acceptance-criteria.md §6 l.450–452: "no real user or participant data"), so screening-speed findings on them will not transfer to real abstracts. |
acceptance-criteria.md l.39, l.409, l.413–425, l.429–460; integrated-plan.md l.443–446 ("staging sessions with CAMARADES testers"), l.1018–1022 |
Add a research plan as an L17 deliverable at G0: (a) a baseline study of the current reviewer page with 6–8 reviewers across two roles before R2a (time per title/abstract decision, time per 30-question form, actions per decision, confusion points); (b) formative sessions on the prototype pack at each freeze (≥5 participants per role affected, including at least two external to CAMARADES); © summative sessions per release on staging with realistic content (a seed project built from an open-licence reference set, not lorem abstracts); (d) a diary or interview round during each pilot. Record protocol, recruitment source, consent and data handling. Replace PE-05 with the metric set in §3 below. |
| UX-02 | Blocker | PKG/acceptance-criteria.md AC-R2a-19 l.172, AC-R4a-13 l.278, AC-ALL-07 l.55; PKG/integrated-plan.md l.372–373, §8 l.976; PKG/ui-coverage-comparison.md §2.3 l.97 |
Reviewer efficiency is never measured and the known efficiency features are outside the plan. The only reviewer-facing performance criteria are server latency (p95 Save/Complete) and first paint of the reconciler. The stage-review design specifies keyboard decisions (I/E/N/J/U/[/]), live completeness and jump-to-next, and quote-into-comment as "Reviewer-efficiency additions". On main the I/E shortcuts do not exist: the decision-card spec asserts "No I/E shortcuts exist yet, so a hint would advertise a dead key", and the template renders no kbd. The plan calls "Reviewer design parity" the AF2 programme's work with a PROPOSAL exit set that must still be agreed, so the plan can reach GA with the new step strip, derived-decision confirmation (DP3), Needs updating, Fix and Complete anyway all adding interaction steps, and no criterion that a screening decision costs no more actions than today. |
RP4 l.84–89; WEB/stage/stage-review/review-decision-card/review-decision-card.component.spec.ts:391–392; WEB/stage/stage-review/review-decision-card/review-decision-card.component.html:107–128; PKG/integrated-plan.md l.372–373, l.976; PKG/acceptance-criteria.md l.172, l.278 |
Add release-independent UX criteria (AC-ALL-14…): a keyboard-only screening path (decide, next, skip) with at most two actions per decision; median time per title/abstract decision and per annotation form on the pilot project not worse than the baseline study (UX-01); input latency under autosave p95 < 50 ms; a visible "N required missing / jump" count on every form host. Move the screening keyboard path and live completeness into R3a/R3b acceptance (they belong to the screening renderer and step strip, which this plan owns at F5), not into the AF2 join. Name the AF2 parity exit set at G0, not later. |
| UX-03 | Major | PKG/integrated-plan.md §7 l.936–938, §6.1 l.803–805, §5.5 l.535; PKG/open-questions-and-assumptions.md Q-12 l.68, §3 l.147–175 |
Most R3 validations are not scheduled anywhere, and the freeze gates do not require them. W0 schedules U13–U15, U19, U26 and U1; W2 mentions U1 again and "F5 work (profiles, screening renderer, Q-26)". U2, U5, U7, U12, U17 and U18 appear in no window. Q-12 says "Prototype both; ship neither until U7 passes", and R3a's critical path runs "stage designer → reviewer step UI (step strip, U7)", yet F3's exit evidence lists no U at all and F5's lists none. The R3a build starts in W2. As written, the step strip and stage designer will be built before they are validated, or the build will stall. The same applies to U6 (publication dialog) for R2c (W2 build) and U20/U21 for R4a/R4b. | integrated-plan.md l.936 (W0 list), l.938 (W2), l.803 (F3 exit), l.805 (F5 exit), l.535; open-questions-and-assumptions.md l.68, l.148–175 |
Put every U in a window at least one window before the build that consumes it: U2, U5, U7, U12 in W1 (F3 work); U17, U18 in W1/W2 (F5 work); U6, U4 in W1; U20, U21, U16 in W2 (F4 work); U10, U11 in W2/W3; U22, U23, U24, U29 in W3. Add "U-validations passed" to the exit evidence of F2 (U6), F3 (U2, U5, U7, U12), F4 (U1, U4, U16, U20), F5 (U17, U18), F6a/F6b (U22, U23), F-C (U3), F-O (U24). Define what "passed" means (UX-01 metrics on the prototype). |
| UX-04 | Major | PKG/contracts.md C5 drafts l.249–256; PKG/open-questions-and-assumptions.md E21 l.123, U13 l.159; PKG/acceptance-criteria.md AC-R2a-01, AC-R2a-06 l.154, l.159 |
Latency, autosave failure and offline tolerance have no UX. E21 is a server-side draft with CAS; the only user-facing case designed is two tabs (typed conflict, both drafts kept). Nothing says what a reviewer sees when autosave fails, when the connection drops mid-form, or how a conflict is resolved (which draft, how to compare). AF2 today keeps drafts in memory only (inventory fact 4) and the web rule forbids implicit per-answer server saves; the drafts ADR will change this, so the save-status language is new for every reviewer. QM v2's publishing/versioning UX already designed this (connection-state service, save-status indicator states, offline queue in IndexedDB, beforeunload, tab awareness), and Q-08 "harvest QM v2" covers only the domain code. |
MAIN/src/services/web/CLAUDE.md:343; PKG/source-status-inventory.md l.61–66; QM2/04-versioning-and-publishing/publishing-versioning-ux.md l.845–928; PKG/decision-register.md l.191 (Q-08 scope) |
Extend E21/U13 with a save-status state machine (Saving / Kept / Retrying / Offline, kept on this device / Failed, retry) with copy, a local draft fallback (IndexedDB) that never silently overwrites the server draft, a conflict screen that shows both drafts' timestamps and lets the reviewer choose or merge, and an acceptance criterion "a network loss during a session loses no edit made before the loss" (E test with network throttling). Add "autosave never blocks input" (UX-02 latency). Harvest QM v2 §14 explicitly in the Q-08 list. |
| UX-05 | Major | PKG/contracts.md C17 l.591–593; PKG/integrated-plan.md §9 item 9 l.1023–1026, §6.1 F1 row l.801; PKG/ui-coverage-comparison.md §5 l.217 |
The terminology and copy contract has no artefact, owner, mechanism or decisions. C17 lists four terms. The plan introduces at least twenty user-facing concepts (draft, Save, Complete, current version, Needs updating, contains outdated annotations, Fix, withdraw, step, profile, form, template, Collective Include, held, accepted answer, gold snapshot, query, concern, informed/independent, additional review). The app has no string catalogue (no $localize/translation; strings are inline), although session-copy-review.md already recommended typed per-feature message constants and study-presence/session-messages.ts follows it. The user guide tells reviewers to click "Save" and "Submit" in one section and "Save progress" and "Complete" in another. The plan says v10's "Save draft" and v4's "All changes saved" would mislead under SL2, but does not decide the verbs. |
contracts.md l.591–593; integrated-plan.md l.1023–1026; MAIN/docs/planning/session-copy-review.md l.18–40; WEB/study-presence/session-messages.ts (exists); MAIN/user-guide/annotating.md l.239–249 vs l.375–376; RP4 l.63 |
Make a copy deck and glossary an F1 deliverable owned by L16: one definition per term, the reviewer-facing verb set (recommend: autosave status "Changes kept, not yet saved"; "Save progress" = immutable incomplete version; "Complete"; "Needs updating"; "Outdated answers"; "Fix"), the alias scheme, and the admin vocabulary. Implement as typed message constants per feature with a guard spec that new screens import them; update user-guide/getting-started/glossary.md in the same PR. Add AC-ALL: every new screen's strings come from the deck. |
| UX-06 | Major | PKG/acceptance-criteria.md AC-ALL-07 l.55, UI-4 l.77; PKG/integrated-plan.md §6.2 item 5 l.826–828 |
Accessibility acceptance assumes tooling the repository does not have. AC-ALL-07 requires "axe reports no serious or critical violations"; the e2e package has only Playwright, SignalR, csv and mongodb dependencies, and no axe reference exists under e2e/. There is no screen-reader protocol, no accessibility owner, and no definition of the manual checklist. The navigation and Study Management work set a good precedent (tree semantics, roving tabindex, forced-colours, ≥44 px targets) but the plan does not carry it into new patterns (step strip, candidate pills, agreement icons, prefill outline, drag-pairing). |
MAIN/e2e/package.json dependencies and devDependencies (no axe); M3/navigation-plan.md l.117–129; SM/design-decision-log.md l.62–63 (D16, D17) |
Add an L17 "accessibility harness" deliverable in W0: @axe-core/playwright in the e2e stack with per-route checks and a committed baseline; a manual checklist (keyboard-only journey, NVDA and VoiceOver on the release's main tasks, 200 %/400 % zoom, forced colours, reduced motion) run at staging acceptance by a named person; a pattern-level a11y spec for each new shared pattern (roles, names, keyboard model) in the handoff template (UX-07). |
| UX-07 | Major | PKG/acceptance-criteria.md UI-2 l.75, UI-5 l.78, UI-8 l.81; PKG/ui-coverage-comparison.md §5 l.206–211 |
No design system of record for the new patterns, and four incompatible prototype stacks. New patterns (step strip, candidate pills, agreement icon, dashed prefill outline, held banner, gold timeline, impact dialog, queue lists, status chips for versions) are each "retain/revise" from a different asset: v10 (Claude Design React _ds kit, which was reverse-engineered from the Angular app, "No Figma was provided"), QM v2 (standalone HTML), redesign v7 (React), #2621 (HTML). The navigation handoff shows the cost: its token file cites a _design-tokens.scss that does not exist and the plan had to record deviations. The plan says older assets are "rebuilt with Material 3 components and roles" but names no pattern inventory, no handoff template and no place where a pattern is specified once. The stage-review spec's ALL-CAPS 13 px buttons and 4 px radii are Material 2 idioms, while the rebuilt rail follows M3 (sentence case) and the _ds dialog uses pill actions; "retain v4 parity" and "Material 3" will produce two button languages on one page. |
V10/prototype/_ds/syrf-design-system-…/README.md ("No Figma was provided"; React components); M3/handoffs/project-navigation-drawer/README.md l.13–20; M3/navigation-plan.md l.69, l.143–191 (deviations); RP4 l.49, l.132; V10/README.md l.50–57 |
Declare the design system of record = FEAT-023's emitted --mat-sys-*/--syrf-* contract plus the shared Angular components (shared/page-shell, shared/side-nav, StatusView chips, overlay-scroll). Add an L16 deliverable "pattern inventory": every new pattern specified once (states, tokens, keyboard model, narrow behaviour, copy) as a shared component with a spec gallery route behind a flag. Adopt a handoff template modelled on the navigation handoff (geometry, tokens, data model, behaviour, a11y, acceptance checks, known gaps, deviations log). Decide one button and type language (M3 sentence case) and re-audit the v4 spec against it before more parity work (question for Chris, §4). |
| UX-08 | Major | PKG/notifications-integration.md §4.1 l.181–192; PKG/ui-coverage-comparison.md §4 l.186–202; PKG/contracts.md C17 l.570–575; PKG/open-questions-and-assumptions.md U25 l.171 |
Work is scattered across five surfaces with no "my work" home. A reviewer or reconciler will find actionable items in: the stage Review entry and My studies, the Reconcile pool with KPI cards, four feature-owned queues ("my concerns and outcomes", "changes awaiting approval", "assigned reconciliation work", "requested reviews"), Needs-updating/outdated flags inside sessions (no list of affected studies), and the notification inbox. The IA in C17/§4 places none of the queues. The #2621 dashboard prototype already concluded "one canonical work queue … replaces the triplication of the same data"; the plan calls it reference material outside scope. | notifications-integration.md l.181–192; ui-coverage-comparison.md l.90 ("C's dashboard is reference material … outside this plan's scope"), l.186–202; /home/chris/workspace/syrf/pr/pr2621.screening-profile-versioning-rationale/docs/features/landing-page/README.md l.103–109, l.122–123 |
Add one project-level "My work" surface (R3c/R4a) that lists every actionable item by role (studies to review by step, sessions needing updating, reconciliation tasks assigned or available, queries and outcomes, change requests), with the four queues as filters and the inbox as history; put it in C17's IA. Extend U25 to cover it. Plan the cross-project landing (new-user empty state with three doors, per #2621) as a post-GA item with an owner. |
| UX-09 | Major | PKG/integrated-plan.md §7 l.947–954; PKG/acceptance-criteria.md AC-GA-04 l.389; PKG/integrated-plan.md §9 item 10 l.1027–1029 |
Reviewers will experience eight successive changes to the same workspace, and "what changed" arrives only at GA. The L5 order (R2a → R2b → R2c → R3a → C1 → O1 → R2d → R3b) is a serialisation rule for code, not a change plan for people; nothing bundles reviewer-visible changes or budgets them. The in-product "what changed" page is a GA criterion, although R2a changes the core mental model (drafts vs versions, history, no hard delete) and pilots begin there. Today the app has a one-shot first-run hint ("GOT IT") and a marketing "What's new" on the public home page; the v4 anchored tour is unbuilt. | integrated-plan.md l.947–954, l.1027–1029; acceptance-criteria.md l.389; WEB/stage/stage-review/review-first-run-hint/review-first-run-hint.component.html; WEB/info/home/whats-new/ (public page); RP4 l.89 |
Separate "reviewer-visible UI change" from "backend/flag" releases and bundle visible changes into at most three reviewer-facing steps (for example: R2a+R2c reviewer surface; R3a+R3b step strip and screening renderer; R2d+C1+O1 workspace additions), with the rest dark until bundled. Build a reusable "What changed" panel keyed by release with per-user dismissal in R2a and make it AC-ALL. Build the anchored tour component once (R3a) and reuse it. Add contextual help links on every new screen through the existing userGuideUrl pipe (AC-ALL-06 requires guide updates but not in-product links). |
| UX-10 | Major | PKG/ui-coverage-comparison.md §2.1 l.72; PKG/open-questions-and-assumptions.md U6 l.152; PKG/acceptance-criteria.md PE task R2c l.417 |
The publication impact dialog is the riskiest admin interaction and is specified as "one dialog combining" five concerns. Per-question impact choice, per-category treatment for three session categories across all prior versions and all bound stages, usage-evidence freshness, missing-reason warnings and within-session conflicts in one dialog is a wall of options; the PE task "predict a publication's impact from the dialog before publishing" has no correctness threshold. | ui-coverage-comparison.md l.72; open-questions-and-assumptions.md l.152; acceptance-criteria.md l.417 |
Specify a staged flow (Review changes → Impact summary → Choices with defaults → Confirm) with a "what reviewers will see" preview and a dry-run summary table; prototype U6 on a realistic fixture (10 changed questions, 3 stages, 200 sessions in all three categories); set a comprehension criterion (≥80 % of admin testers predict the effect on each category correctly). Reuse GuardedReviewSettings conflict patterns as the plan says, but state the information hierarchy. |
| UX-11 | Major | PKG/contracts.md C17 l.576–577; PKG/open-questions-and-assumptions.md U27 l.173; PKG/acceptance-criteria.md §3 preamble l.65–70; PKG/migration-adoption-rollback.md §2 l.76 |
Coexistence will produce a visibly two-generation product for years. Legacy projects stay legacy "indefinitely"; the UI standard applies only to new and updated screens; a user who belongs to a legacy and a canonical project will see different navigation sections ("Studies"/"Screening" today vs "Design"/"Stages"/"Members & groups"), different editors (legacy "Question design" vs "Design") and different settings pages, with no mode indicator, no explainer and no decision on whether legacy screens get an M3 restyle. | WEB/project/project-nav/project-nav.component.ts:438–544 (Studies and Screening sections); contracts.md l.576–577; migration-adoption-rollback.md l.76; acceptance-criteria.md l.65–70 |
Define the minimum shared chrome that is identical in both modes (rail, overview, members, data export) and a per-project "workflow version" badge in the overview and rail; write the legacy explainer copy; make U27 a cross-project consistency check with a tester who holds both kinds of project; add a scoped "legacy screens visual refresh" item (M3 restyle, no behaviour change) to FEAT-023 or L16 so that the atomic cutover covers legacy pages too. |
| UX-12 | Major | PKG/acceptance-criteria.md UI-3 l.76, UI-4 l.77, UI-7 l.80; M3/README.md l.43–58; M3/technical-plan.md l.37–47, l.236–245 |
The Material 3 standard doubles verification for every screen until a cutover with no date, and demands dark-mode evidence for a mode users cannot reach. UI-3 requires correct rendering on the M2 bridge and on the M3 path; FEAT-023's atomic light cutover (Wave 5) has no schedule and Waves 0–4 are still open. UI-4/UI-7 require dark-theme correctness and dark screenshots for every new screen, while FEAT-023 says dark mode is a follow-on milestone and themeToggle is default-off. |
acceptance-criteria.md l.76–80; M3/README.md l.55–58; M3/technical-plan.md l.236–245; MAIN/src/charts/syrf-common/env-mapping.yaml:1467–1470 |
Ask Chris to sequence the FEAT-023 light cutover before R2a's reviewer UI (or at the latest before GA) and fund it as a join (X-M3); until then limit UI-3 evidence to the path the preview environment actually renders plus the theme-contract checks. Keep dark-mode evidence as check:contrast/token-contract checks, with full dark screenshot sets only once themeToggle is activated. |
| UX-13 | Major | PKG/acceptance-criteria.md UI-8 l.81, AC-ALL-13 l.61; PKG/open-questions-and-assumptions.md A-24 l.206 |
Design review is a single-person bottleneck with no process. UI-8 says Chris accepts every new or materially changed screen on each PR's preview before merge; A-24 names the cost if he cannot ("design review moves to a delegate or a later checkpoint"). Across R1–R5 and the lanes there are more than forty new or changed screens. "Materially changed" is undefined; there is no design-QA checklist, no screenshot-diff baseline and no role for a reviewing agent. | acceptance-criteria.md l.61, l.81; open-questions-and-assumptions.md l.206 |
Two-tier design review: tier 1 automated and agent-run per PR (theme guards, contrast, screenshot matrix at UX-15 widths in light, axe, a design-QA checklist against the handoff with evidence; a second agent reviews against the handoff); tier 2 Chris accepts at release level on staging with the bundle of screens and the tier-1 evidence, plus spot checks on previews. Define "materially changed" (new route, new pattern, changed primary action, changed copy for a core verb). |
| UX-14 | Minor | PKG/open-questions-and-assumptions.md U1 l.147, U16 l.162, U20 l.166; PKG/contracts.md C9 l.374–380; PKG/ui-coverage-comparison.md §2.3 l.101–102, §5 l.218–220 |
Reconciler journey gaps beyond N candidates. No end-to-end reconciler journey (pool → task → screening part → matching → form → Complete → next) is specified; "random start, assigned first" are PROPOSALs with no fallback design; the composition of the RX1 screening part and the form part in one task is only "a separate part of the workspace"; U16 covers VS1 display but not the disclosure to the reviewer that viewing accepted answers labels their work "informed" (VS2); the narrow-screen layout for N candidates is only "needs its own budget". |
open-questions-and-assumptions.md l.147, l.162, l.166; contracts.md l.374–380; ui-coverage-comparison.md l.101–102, l.218–220 |
Add a reconciler journey map and a narrow layout (candidate selector plus single column, disagreeing candidates never hidden) to F4's exit evidence; add VS2 disclosure copy to U16 ("You are viewing accepted answers; your contribution will be recorded as informed"); specify the "next task" rule and its UI as part of U20; give U4's unseen-control warning a list with jump links. |
| UX-15 | Minor | PKG/acceptance-criteria.md UI-6 l.79; PKG/integrated-plan.md §6.2 item 5 l.826–828 |
The width matrix is not the app's breakpoints, and touch and phone screening are absent. UI-6 lists 320, 375/390, 480, 600, 768, 960, 1280, 1440; the app's breakpoints are 599.98, 904.98, 1239.98, 1439.98, and the reviewer workspace has its own content threshold (980 px in the v4 spec). 960 and 1280 miss the 905 and 1240 edges. No touch-target criterion; title/abstract screening on a phone is a real use case that the plan never names. | WEB/shared/layout/break-points.ts:30–117; SM/design-token-mapping.md l.107–124; RP4 l.116–121; acceptance-criteria.md l.79 |
Align UI-6 with break-points.ts (±1 px around each edge) plus 320 and 200 % zoom; add ≥44 px targets below 600 px (Study Management D17 precedent); add a "screen 20 studies on a 390 px phone with the keyboard hidden" journey to R3a acceptance. |
| UX-16 | Minor | PKG/acceptance-criteria.md UI-2 l.75, UI-7 l.80; PKG/contracts.md C4 publication phases l.217–220; PKG/integrated-plan.md P2 row l.690 |
Loading, empty, error and long-running states are required per component but not tied to the existing shells and job language. Study Management established app-page-shell and app-page-state (one shared loading/empty/error component) and FEAT-023 defines the long-running-job visual language. Publication phase 2, ASySD fuzzy matching, as-of export generation, O2 dry-runs and adoption waves are all long-running and have no UI contract. |
SM/design-decision-log.md l.45–48; WEB/shared/page-shell/page-state.component.* (exists); M3/technical-plan.md l.205–221 |
Make app-page-shell/app-page-state and StatusView chips mandatory for every new admin page (add to UI-2); map every long-running operation to the job language (determinate only with real numerators; honest "progress not measurable"; raw status preserved) and give each a place in Processing or a release-owned surface. |
| UX-17 | Minor | PKG/acceptance-criteria.md AC-R2a-06 l.159, AC-R2a-07 l.160; PKG/contracts.md C1 l.112–114; PKG/open-questions-and-assumptions.md U13 l.159 |
Typed conflicts have no user-facing recovery design. Stale base, two-tab conflict, held-for-compatibility, revoked access, admission refused ("feature unavailable") and read-only containment all return typed errors; U13 covers states, not recovery. | acceptance-criteria.md l.49 (AC-ALL-01 "typed feature unavailable"), l.159–160; contracts.md l.112–114 |
Add an error-and-recovery copy section to the copy deck (UX-05) and to U13: for each typed conflict, the message, the user's options, and what is preserved. Test them in the R2a user session (UX-01 metric: 100 % recover without data loss). |
| UX-18 | Minor | PKG/integrated-plan.md §5.2 item 3 l.289–291, §9 item 7 l.1014–1017; PKG/open-questions-and-assumptions.md U28 l.174 |
Admission and pilot rollback have no named UI. R0's admission service has "an audited admin action to admit or remove a project" but no screen; U28 asks what reviewers see after rollback to read-only but no surface or copy is specified. | integrated-plan.md l.289–291, l.1014–1017; open-questions-and-assumptions.md l.174 |
Add a "Workflow version" panel in Project settings (admitted/legacy, read-only containment state, who changed it, when) and a reviewer banner for read-only containment; include both in U27/U28 and the copy deck. |
| UX-19 | Minor | PKG/integrated-plan.md R3d l.564–574, §7 W4–W5 l.940–941; PKG/acceptance-criteria.md AC-R3d-01..05 l.254–260 |
Guided setup is the moment the new mental model is taught, but it lands in W4/W5 with no interim onboarding and no success metric. New admins on R2a–R3b pilots will configure forms, profiles and steps across several pages without the guided route; AC-R3d measures parity and resumability, not time-to-first-screenable-study. | integrated-plan.md l.564–574, l.940–941; acceptance-criteria.md l.254–260 |
Add an interim setup checklist content update at R2a (readiness-based tasks per C17 already proposed) and a metric for R3d: a new admin reaches a screenable stage from templates in ≤20 minutes without help (PROPOSAL), measured in the R3d session. Prototype U11 in W1, not W2. |
| UX-20 | Note | PKG/ui-coverage-comparison.md §2.3 l.102; PKG/acceptance-criteria.md AC-R1a-07 l.114 |
Debug surfaces on reviewer-facing routes are only partly listed. The plan removes the "Focused question" debug text from the designer, but the reconcile route (which R4a replaces) and the review-completed page render app-debugger-group panels, and the reconcile grid uses FlexLayout (gdColumns). |
WEB/stage/stage-reconcile/stage-reconcile.component.html:4–18; WEB/stage/stage-review/review-completed/review-completed.component.html (debugger group); WEB/project/project-admin/question-management/design/design.component.html:14 |
Add AC-ALL: no debug components render on any route reachable by a reviewer or reconciler unless the debug flag is on; include the review-completed page in R2b's "progress lists" revision. |
| UX-21 | Note | PKG/integrated-plan.md §9 l.1000–1034; PKG/acceptance-criteria.md §5 |
UX telemetry exists and is unused by the plan. The app has Sentry, a LogRocket session-recording flag and a Google Analytics token; no pilot monitoring or UX metric uses them, and no privacy decision is recorded. | MAIN/src/services/web/package.json:60–62; MAIN/src/charts/syrf-common/env-mapping.yaml:1563–1572, 2031–2036 |
Decide (question for Chris) whether privacy-safe in-app timing events (study opened → decision; form opened → Complete; actions per decision) may be collected on pilots; if yes, add an instrumentation item to L17 and use it for the UX-02 metrics; if no, measure in moderated sessions only. |
| UX-22 | Note | PKG/contracts.md C17 l.578–580, C6 route status DTO l.292; MAIN/user-guide/stages/screening.md l.37 |
Reviewer-facing progress copy is a known confusion the plan fixes only for DTOs. C17 and C6 separate gate, sufficiency and work status for admin DTOs; the reviewer progress bar's "Unavailable" and "sufficiently screened" language, which the user guide has to explain, is not in the copy contract. | contracts.md l.292, l.578–580; user-guide/stages/screening.md l.37, l.91–97 |
Add reviewer progress vocabulary (your work, available to you, waiting on others, locked by a step) to the copy deck and to R2b's progress-list revision. |
Resolved first-round items I checked and did not re-raise: C-12 (U13–U29 added), C-15 (Dockview amendment), C-21 (IA corrections), C-26 (U1 detail), accessibility in the ship checklist, terminology contract named in C17, pilot-operations U28, inbox U25, narrow screens and budgets in UI comparison §5. Where I re-raise (UX-03, UX-05, UX-06, UX-14, UX-15), it is because the resolution names the need without a schedule, an owner, a mechanism or tooling.
3. Improvements: an end-to-end UX strategy for the programme¶
3.1 Information architecture. Keep C17's project IA (Overview, Review, Reconcile, Design, Stages, Members & groups, Data, Project settings) and add: a project-level "My work" surface (UX-08) as the reviewer's and reconciler's landing inside a project; a "Workflow version" badge and panel (UX-18); one Reconcile entry per study × form task as decided, reached from My work and from the stage; the agreement view under Data. Keep "Library" with Study Management. After GA, the cross-project landing (#2621 dashboard: digest, resume CTA, one work queue, new-user empty state) becomes the global home. Record the IA as a route inventory with owner per route (the M3 programme's "one owner per route group" rule) so navigation and section-shell changes coordinate.
3.2 Design-system use. Source of truth = FEAT-023's emitted token contract and the shared Angular components. Add an L16 pattern inventory and build each new pattern once as a shared component with a spec gallery route: step strip; candidate pills and agreement icon; prefill (autofill) marker; held/needs-updating/outdated chips (as StatusView kinds); history timeline; impact dialog scaffold; queue list; version badge; "What changed" panel; anchored tour. Each has states (hover, focus-visible, pressed, selected, disabled, error, loading, empty), a keyboard model, narrow behaviour and copy keys. Decide one button and type language (M3 sentence case, 40 px, consistent radius) and re-audit the v4 spec. Sequence the FEAT-023 light cutover as join X-M3 before R2a's reviewer UI if Chris agrees (UX-12).
3.3 Prototyping cadence. One prototype pack per freeze gate (F1: reviewer states, forms, history, export disclosure; F2: impact dialog; F3: step strip, stage designer, route-change messages, skip; F4: N-candidate workspace, matching, assignment, queries; F5: screening renderer, DP2 correction; F6: as-of export, PRISMA report; F-C/F-O: cohort chip, schema authoring; R3d: guided setup). Claude Design .dc.html packs with the _ds kit are acceptable for interaction, but every pack ships with a handoff document on the navigation-handoff template, maps tokens to --mat-sys-*/--syrf-*, and includes the fixture data used (realistic abstracts and a 200-question form where relevant). Each pack is tested with users (3.5) before the gate exits.
3.4 Design reviews and the designer/agent working model. Designer (or Chris with a design agent) produces the pack and handoff. The implementing agent audits the handoff against code first (DONE/PARTIAL/NOT DONE with file:line, as the 21 September agent brief already requires), builds in a sole-writer worktree, and produces a design-QA report: screenshot matrix at the UX-15 widths (light; dark via token checks until activated), state coverage list, axe result, keyboard journey transcript, copy-deck compliance, deviations with reasons. A second agent reviews the report against the handoff. Chris accepts at release level on staging (UX-13), with per-PR previews for new patterns only. Deviations go back into the handoff's deviations log, as the navigation plan does today.
3.5 Research plan with real users. (a) Now, before R2a: a baseline study of current screening and annotation (6–8 reviewers, moderated, think-aloud, on staging with a realistic reference set; measures below) and a review of support requests and the FAQ to list today's pain points. (b) At each freeze: 5+ participants per affected role (reviewer, reconciler, admin/designer), at least two outside CAMARADES, on the prototype pack; tasks mirror PE's "explain why" questions with a correctness threshold. © Per release: summative sessions on staging with the seeded pilot projects plus one realistic-content project; (d) during production pilots: a two-week diary or weekly 20-minute interviews per pilot project; (e) a terminology card-sort in W0 to validate the copy deck. Record protocol, consent and data handling in L17.
3.6 UX metrics and acceptance criteria (proposed AC-UX-01…09, thresholds PROPOSAL). 01 Task success ≥ 80 % without help per task (replaces PE-05's 4 of 5). 02 Comprehension ≥ 80 % correct on "which version counts", "why is this study offered or locked", "what does this gold answer rest on", "what will this publication do". 03 Screening throughput: median time per title/abstract decision and actions per decision on the pilot ≤ baseline; keyboard-only path exists. 04 Annotation: time to complete the baseline form ≤ baseline + 10 %; input latency under autosave p95 < 50 ms. 05 Reconciliation: time per three-candidate study ≤ 1.5× the two-candidate baseline. 06 Error recovery: 100 % of testers recover a stale-draft or two-tab conflict without data loss. 07 Setup: a new admin reaches a screenable stage from templates in ≤ 20 minutes. 08 Satisfaction: SEQ ≥ 5.5 per task or SUS ≥ 70 per role per release. 09 Accessibility: zero serious/critical axe findings; keyboard-only completion of the release's main tasks; screen-reader walkthrough passes the checklist. Instrumented timing on pilots if Chris approves (UX-21).
3.7 Change management for reviewers. Bundle reviewer-visible changes into at most three steps (UX-09); ship the "What changed" panel and the tour component early; contextual help on every new screen; user-guide glossary parity enforced by the copy deck; a legacy-project explainer and workflow-version badge (UX-11).
4. Questions for Chris (product decisions only)¶
- FEAT-023 sequencing. Should the atomic Material 3 light cutover be scheduled as a join before R2a's reviewer UI (or before GA), so new screens are verified once rather than on both paths? Recommendation: yes, before R2a's staging pilot; if not, accept the dual-path cost explicitly and limit UI-3 evidence to the preview-rendered path plus token checks.
- Dark-mode evidence. Until
themeToggleis activated, is token-contract contrast checking sufficient for UI-4/UI-7, with full dark screenshot sets deferred to dark-mode activation? Recommendation: yes. - Button and type language. Adopt M3 sentence-case buttons and consistent radii across the reviewer page, replacing the v4 spec's ALL-CAPS 13 px buttons and 4 px radii? Recommendation: yes; re-audit the v4 spec before further parity PRs.
- Reviewer verbs. Confirm the final reviewer verbs now: "Save progress" (immutable incomplete version), "Complete", autosave status "Changes kept, not yet saved", "Needs updating", "Outdated answers", "Fix". Recommendation: keep "Save progress" (already in the app and guide) rather than "Save", and never "Save draft".
- Research participants and privacy. May CAMARADES recruit external SyRF users for sessions, and may pilots collect privacy-safe timing events (study opened to decision) for the efficiency metrics? Recommendation: yes to both, with consent text in the pilot admission step and no content capture.
- Phone screening. Is title/abstract screening on phones a supported target for R3a/R3b? Recommendation: yes for screening steps only; annotation remains tablet/desktop.
- Legacy visual refresh. Should legacy-project screens get an M3 restyle (no behaviour change) under FEAT-023 so the app does not look like two generations for the years R6 takes? Recommendation: yes, scoped to chrome and shared pages.
- "My work" placement. Project-level "My work" in R3c/R4a and a global landing after GA (per #2621)? Recommendation: yes; the global landing gets its own owner and brief.
- Design acceptance cadence. Replace per-PR preview acceptance with release-level staging acceptance plus tier-1 agent design QA (UX-13)? Recommendation: yes, with per-PR acceptance only for new shared patterns.
5. Coverage gaps¶
- I read the prototype HTML files as text and the handoff documents in full, but did not render the v10, v4, QM v2, redesign v7 or #2621 prototypes in a browser; interaction claims about them come from their READMEs and the plan's own inventory.
- The Figma file "SyRF Design v2" and the remote classification site were not available; the
_dsdesign-system bundle README was read, its components were not inspected. - I did not verify the preview-environment acceptance workflow (how Chris reviews screens today) or whether LogRocket/GA are enabled in any environment (flag and token exist; runtime state UNVERIFIED).
- No current user data (support tickets, FAQ analytics, session recordings) was available to me; the baseline-study recommendation stands in for it.
- I did not check the stage-review workspace at narrow widths on
mainor the Dockview narrow fallback beyond the documents; the plan's responsive claims about the current reviewer page are fromRP4and the AF2 docs. - Project navigation line numbers were verified in
project-nav.component.tsat l.438–544; the 'Screening info' group is defined withvisible: falseat l.519 and its children carry their own visibility, which I did not trace further. - I did not review the notification inbox UI in the open PR stack beyond the plan's description; U25 remains the only validation for it.
Critical Files for Implementation¶
- /home/chris/workspace/syrf/pr/pr3617.research-screening-as-specialised-annotation-gxgahs/docs/planning/integrated-review-plan-2026-10/acceptance-criteria.md
- /home/chris/workspace/syrf/pr/pr3617.research-screening-as-specialised-annotation-gxgahs/docs/planning/integrated-review-plan-2026-10/integrated-plan.md
- /home/chris/workspace/syrf/pr/pr3617.research-screening-as-specialised-annotation-gxgahs/docs/planning/integrated-review-plan-2026-10/open-questions-and-assumptions.md
- /home/chris/workspace/syrf/pr/pr3617.research-screening-as-specialised-annotation-gxgahs/docs/planning/integrated-review-plan-2026-10/contracts.md
- /home/chris/workspace/syrf/pr/pr3617.research-screening-as-specialised-annotation-gxgahs/docs/planning/integrated-review-plan-2026-10/ui-coverage-comparison.md