Skip to content

PRISMA and deduplication specification amendments (A–O)

Temporary planning document; planning only. This page is the single register of the changes this programme needs to the Approved PRISMA package (FEAT-011, docs/features/prisma-specification/) and the Approved deduplication service specification (FEAT-012, docs/features/deduplication/service-specification.md). Approved specifications are never changed silently. Each amendment goes through the FEAT-011 change policy (affected box or entity → downstream phase impact → linked specification and query amendments → release checklist update) in a documentation PR opened before the build that depends on it (plan §12, step 3): the write-shaping amendments A and H are frozen at F3, before R3a writes outcomes; C, D, G, J, K, M, N and O at F-P, before the P1 build; L at F-P for its rules and P2 for its parity evidence; B, E and F at F6b. The implementing code PR then updates the specification text it relies on and the release checklist, and cites the amendment.

Where each amendment came from:

  • A–F were raised by the earlier PRISMA compatibility review, review-prisma-integration-2026-10-03.md, section "Concrete conflicts, recommendations and amendment gates". This page summarises them.
  • G–J came from adversarial review B (finding B-14) of this package.
  • K and L are Chris's additions of 3 October 2026: manually reported counts for steps done outside SyRF, and ASySD deduplication inside SyRF. Their detailed rules are confirmed under Q-37.
  • M, N and O came from round-2 reviews SR (SR-02, SR-07, SR-15) and V2 (V2-01): full-text retrieval, the Citation to Publication link, and report-to-study linkage.

Approval status (Chris, 3 October 2026): A, C, D, G, H, I and J are approved (Q-06a). K and L were requested; their rules below are proposals under Q-37. B, E and F are open (Q-06b). M and N are new proposals (M depends on D4-07 for actors and the PDF rule). O is conditional on D4-08. Nothing on this page is implemented.

Summary

ID Amends One-line change Status Lands in
A FEAT-011 constraints Keep authoritative collective profile outcomes for PRISMA; separate within-stage admission and configurable routing; define "entering screening" Approved R3a (write shape), R5b
B FEAT-011 flow mapping Count reports by real report identity, not by Citations Open (Q-06b) R5b
C FEAT-011 mapping and taxonomy One downstream source column from the earliest import Approved P1
D FEAT-011 and FEAT-012 Reviewed duplicates need admin-reviewed mapping; never double count a reviewer's contribution; scenario 2 becomes admin-reviewed Approved P2
E FEAT-011 reasons Report reason coverage honestly; keep screening and ordinary gold apart; FEAT-009 truncation becomes a coverage value Open (Q-06b) R5b
F FEAT-011 time and history Append corrections; freeze reports; no rollback by removing fields Open (Q-06b) R5b
G FEAT-011 Phase 16 (MIG-11 to MIG-14) Per-project adoption replaces platform-wide backfills Approved R6 (P1 for the classification tool)
H FEAT-011 ScreeningOutcome (three placeholders) Outcome per profile with route provenance and complete authority values, including Imported Approved R3a, R3b
I FEAT-011 rollback Canonical-aware rollback replaces $unset Approved All PRISMA-writing releases
J FEAT-011 deletion cascade Withdrawing keeps Citation history; whole-project deletion per D3-12 (recommended: ADR-014 removal with a tombstone) Approved for withdrawal (presentation per Q-33); whole-project scope pending D3-12 P1
K FEAT-011 flow mapping; FEAT-012 §11 Manually reported counts for steps done outside SyRF, with an entry-phase rule and per-box combination Requested; rules proposed (Q-37) P1 (identification and deduplication records), R5b (other step types, reports)
L FEAT-012 ASySD deduplication inside SyRF; merge as an alias; privacy; QC sampling; parity metric Requested; rules proposed (Q-37, D4-21, D2-12) P2
M FEAT-011 taxonomy, mapping and three-level model Retrieval is fullTextStatus only; human retrieval actions; the FullTextNotRetrieved lifecycle precedence is superseded Proposed (D4-07) P1 (events), R3a/R3b (admission), R5b (boxes)
N FEAT-011 three-level model Citation to Publication link held in a link record; linking never rewrites a Citation Proposed P1, P2
O FEAT-011 flow mapping (boxes 10, 16, 17) Report-to-study linkage: several reports, one study; linking never merges evidence Proposed (D4-08) P2 (or R5b with B)

A. Within-stage admission and configurable routing, separate from PRISMA outcomes

  • Today: older screening and profile documents say operational pools use the final outcome, and FEAT-011 treats "screened" as pool entry.
  • Problem: confirmed decisions DP6 and DP7 let a reviewer's own Include open steps within a stage, and make cross-stage routing configurable. That must never change what PRISMA reports.
  • Amendment:
  • PRISMA keeps the authoritative collective outcome per profile.
  • Personal admission and routing are recorded separately and never counted as collective Included.
  • Skipping a step never invents a final screening vote.
  • "Entering screening" (boxes 4 and 8, "made available to screeners") means protocol scope, or actual release including shared batches and personal grants, recorded as a StudyEnteredPool event (FEAT-011's pool-entry event) in StudyPoolLedger with the filter and profile versions used.
  • Fixtures: FX-PRISMA-03a, and FX-PRISMA-03b for early-stopped and batched reviews.

B. Records versus reports

  • Today: FEAT-011's flow mapping counts reports (boxes 10, 16, 17) as a sum of Citations, but a Citation is every import occurrence.
  • Problem: one report imported twice gives two records, one report and one study. Two real reports of one investigation need two report links, not two inferred cohorts.
  • Amendment: count reports by exact report identity, with explicit coverage. Until this is approved, exports label citation totals as records, never as verified reports.
  • Status: open (Q-06b, Q-23).

C. One downstream source column

  • Today: the mapping uses "has a Citation from source X" predicates, so a study can fall into both source columns, while the taxonomy says the earliest import decides one column.
  • Amendment:
  • Apply the earliest-import rule everywhere downstream, with deterministic tie handling.
  • Every import still counts at identification.
  • An unknown legacy source stays "unknown/unclassified"; it is never guessed as Database.

D. Duplicates that already have review data

  • Today: FEAT-012's scenario table allows merging when one study is reviewed. Its invariant says never auto-merge studies with review data, and reviewed records in different stages may stay separate Studies sharing a Publication.
  • Amendment:
  • Use admin-reviewed mapping whenever review evidence exists.
  • Keep both candidate sets, but never double count the same reviewer's contribution to the same form (SF2).
  • Flag conflicts rather than choosing one duplicate automatically.
  • Moving Citations needs an identity and lineage manifest plus dedup reversal.
  • Outcome migration must never deduplicate as a side effect.
  • Engine link: reviewed-record merge and split are a C1/C2 contract operation (E33).

E. Reasons and gold

  • Amendment:
  • Screening rules may produce a final outcome from candidate agreement; ordinary annotation matching never produces gold automatically.
  • DP5 Off (exclusion-reason reconciliation off) stays allowed.
  • When primary reasons are unavailable or still being reconciled, reports state reason coverage instead of claiming a complete breakdown.
  • Screening gold and ordinary gold versions are never conflated.
  • Status: open (Q-06b, Q-22).

F. Time, protocol amendments and frozen reports

  • Amendment:
  • Corrections and protocol amendments append; they never edit old report counts.
  • Reports freeze against exact evidence, and current views are recomputed under explicit profile and filter versions.
  • Where history is missing, reports show unknown coverage, never reconstructed decisions or timestamps.
  • "Rollback by removing fields" is unsafe after canonical writes (see I).
  • Status: open (Q-06b).

G. Per-project adoption replaces platform-wide backfills

  • Today: prisma-constraint-annotations.md Phase 16 requires backfilling lifecycleStatus = Active on all studies (MIG-11, line 279), migrating all screening decisions into screeningOutcomes[] (MIG-12, line 280), populating sourceType from LibraryFileType (MIG-13, line 281) and migrating stage settings to a unified schema (MIG-14, line 282); three-level-data-model.md lines 319–322 list the same Phase 16 backfills; the taxonomy (line 397) requires an admin bulk classification tool.
  • Problem: this programme keeps legacy projects unchanged until each is adopted after a reviewed manifest (MIG1 planning, assumption A-04). A platform-wide rewrite would contradict that, and would fabricate per-profile outcomes from one project-wide decision.
  • Amendment:
  • MIG-11 and MIG-12: lifecycle status and screening outcomes are created per project at adoption (R6), from the approved manifest, with coverage labels and authority = LegacyUnknown where the authority is not recoverable (H). Legacy projects that aren't adopted keep their current-state counts, labelled with their basis.
  • MIG-13: sourceType is never backfilled platform-wide. P1 ships the admin source-classification tool for searches imported before P1 (AC-P1-04); at adoption (R6) the manifest proposes Database only for PubmedXml searches and leaves every other source null ("unknown/unclassified") until an administrator classifies it. An unknown source is never guessed.
  • MIG-14: stage settings are not migrated to a unified schema platform-wide. Canonical projects use StageSettings versions from R2a/R3a; legacy stages keep their booleans until the project is adopted, when the manifest maps them (per-stage placement table, brief §1.12).
  • FEAT-011's release-3 checklist items for these backfills are tested per adopted project (migration §7); the checklist line "ALL existing studies have lifecycleStatus backfilled" is read as "all studies of adopted projects".

H. Screening outcome shape

  • Today: FEAT-011 defines the ScreeningOutcome in three places that differ: study-lifecycle-and-source-taxonomy.md lines 267–294 ({profileId, stageId, result, primaryExclusionReason, resolvedAt, authority} with ScreeningAuthority = CandidateAgreement | Reconciled); three-level-data-model.md lines 221–234 ({profileId, stageId, finalOutcome, primaryExclusionReason, decidedAt, source} with source = Reconciled | CandidateAgreement | Admin); and prisma-constraint-annotations.md, which pins the taxonomy's shape in the Phase 11 MUST NOT (line 180), the Phase 13 MUST (line 226) and the Release 3 checklist (line 331). None has a legacy or unknown authority value; only the second has Admin.
  • Problem: a profile can be shared by several stages, so one stageId misstates where the decision came from. Adopted legacy decisions have no known authority. Imported decisions (CSV mapping, FEAT-004) have an authority of their own.
  • Amendment (one shape, applied to all three places in one change):
  • The outcome is per profile, with route provenance (the stages and steps that contributed, with their settings versions) instead of a single stage ID.
  • authority ∈ {CandidateAgreement, Reconciled (profile adjudication, R4p), Admin (override, with an audit record), Imported (decisions imported from another tool or a CSV column, with the independence declaration), LegacyUnknown (adopted without recoverable authority)}. Ordinals are appended, never reordered.
  • The reason is structured with a coverage status; its shape allows either one primary reason or several counted reasons until Q-22 decides, and carries the primary-reason rule that produced it (D4-13) and FEAT-009's truncation as a coverage value ("primary agreed; sub-reason not agreed").
  • result ∈ {Pending, Included, Excluded, Conflict, Unsure (D4-01)}; a collective Unsure is never Excluded.
  • The Phase 11 MUST NOT, the Phase 13 MUST and the Release 3 checklist cite this shape; the taxonomy and three-level placeholders are replaced by one reference to the frozen C12 write-shaping contract.
  • Timing: frozen at F3 with amendment A, before R3a writes outcomes.

I. Canonical-aware rollback

  • Today: three-level-data-model.md (lines 291–322) describes rollback by $unset of the new fields, and FEAT-006's design decision D18 says the same.
  • Problem: once canonical data exists, removing fields destroys evidence and breaks older readers.
  • Amendment: after canonical writes, roll back by stopping new writes, keeping new data authoritative, and using read-only containment or a verified forward-recovery adapter (contract C16; migration §1). Never down-migrate destructively.

J. Deletion and withdrawn searches keep identification history

  • Today: three-level-data-model.md line 262 says "Delete Study removes its Citations" and line 263 removes all Studies with a project. ADR-014 (product decisions approved 12 August 2026; flag deletionLifecycle on main) schedules project and standalone-search deletion with a 24-hour grace period and then physically deletes Project, Search and Study documents, leaving minimal tombstones (ADR-014-reversible-deletion-and-permanent-tombstones.md lines 59–74, 142–146, 220–250). Search and project deletion fail closed on main until that lifecycle is enabled.
  • Amendment (scope per D3-12):
  • Citations are immutable identification history.
  • Withdrawing a search is a reversible appended event that hides its Studies from pools and from current reports but keeps Citations and all canonical evidence; its external step records (K) are withdrawn with it. Current reports exclude withdrawn searches and say so ("excluded from this report: n records from withdrawn search X", Q-33). Frozen reports never change.
  • Deleting a whole project keeps ADR-014's physical removal with a tombstone; PRISMA snapshots and canonical evidence do not survive their project. The user guide tells administrators to export reports before scheduling deletion, and the grace period allows restore. Which canonical collections join ADR-014's deletion scope is X-DEL (programme integration), not this amendment.
  • Deleting a single study outside a search withdrawal is not offered for admitted projects; RemovedOther (box 3) is the methodological equivalent and keeps the record.
  • Still open: none for scope; presentation copy is Q-33's recommendation.

K. Steps done outside SyRF: manually reported counts (Chris, 3 October)

Problem: many reviews deduplicate, and sometimes screen or retrieve, outside SyRF before or alongside import, for example ASySD's R package or Shiny app, EndNote, or another screening tool. FEAT-011 derives every count from SyRF data. For example, database_results is the sum of SystematicSearch.numberOfCitations (the imported count), so a search deduplicated before import understates identification and leaves box 3 without its duplicates; a review that screened titles and abstracts outside SyRF gets box 6 = 0 and an overstated box 10 (V2-04). FEAT-011 has no way to record these steps.

Amends: FEAT-011 flow mapping (boxes 1–9, 11–15; fields #3–#30); FEAT-012 §11.1 (box 3 formula) and §11.2 (count consistency).

Proposed amendment (rules under Q-37):

  1. External step record. Each record is project-scoped, optionally tied to one systematic search or source, and holds:
  2. the step type: identification at source, deduplication, automation removal, other removal, title/abstract screening, full-text retrieval, full-text assessment, or studies from a previous review version;
  3. the affected PRISMA field(s), restricted to fields #3–#30 (the derived fields #31–#34 are never reported, always computed);
  4. a non-negative count;
  5. timing: before import, or outside SyRF after import;
  6. the search round it belongs to (searchRound, SR-24);
  7. the tool or method, as free text with suggested values such as "ASySD (R)" or "EndNote";
  8. a description and an optional evidence note;
  9. who entered it and when.

Records are append-only: a correction supersedes the previous record and keeps its history. A record tied to a withdrawn search is withdrawn with it (J). 2. Entry phase per search or import. Each search or import carries an entry phase derived from its external records: identified, after deduplication, after title/abstract screening, after retrieval, or after full-text assessment (included elsewhere). Records imported after an outside step count as having passed that step: they create no pool-entry or outcome counts for that phase in SyRF, and the external record supplies the phase's counts with coverage "reported externally". For "included elsewhere", admission sets lifecycleStatus = Included and records the required profiles' outcomes with authority = Imported, coverage "external" (H), so box 10 counts them without inventing SyRF decisions. 3. Per-box combination. For boxes 2–9 and 11–15 each field value is the computed SyRF value plus any applicable reported external value for the same field; the report manifest stores both parts, the diagram shows the total and marks fields that include reported values ("includes 120 duplicates removed outside SyRF before import"), and exports include the breakdown. Boxes 10, 16 and 17 are computed only. Box 1 comes only from "studies from a previous review version" records (rule 7). 4. Identification (K.2 and K.3 agree). For the identification fields (#3–#6, #10–#12, #29–#30) the "reported" part is the "identified at source" count, which replaces the imported count for that search when it exists, because the imported count is the post-processing count; otherwise identification uses the imported count, labelled "as imported; processing before import not reported". The difference between identified at source and imported records is accounted for in box 3 by the reported removals (rule 6). 5. Consistency check. For a search with an "identified at source" count: identified at source − reported removals before import = imported records. A mismatch warns and asks for an explanation; it never blocks entry. A snapshot that still mismatches cannot be frozen until an administrator records the explanation (C12 identities). 6. No double counting. External screening, retrieval or assessment counts are allowed only for records not screened, retrieved or assessed in SyRF at that phase; records with SyRF decisions, including imported decisions (authority = Imported), count as "screened in SyRF" and refuse an external count for the same phase. When SyRF also deduplicates (L), box 3 duplicates = SyRF-detected duplicates (FEAT-012 §11.1) + reported external duplicates; FEAT-012 §11.2's count-consistency equation holds over SyRF-held Citations only, and the manifest keeps both parts. 7. Box 1 (D4-11). A "studies from a previous review version" record supplies previous_studies and previous_reports; when one exists the diagram switches to the updated-review template variant and box 16 = new + previous (FEAT-011 mapping §6). Full updated-review support (importing the previous review's included studies with PreviouslyIncluded) stays deferred. 8. Authority and audit. Entering or correcting external counts needs the PRISMA report capability from the permission matrix (Q-03). Every change is audited, and frozen reports pin the record versions they used. 9. Withdrawn searches. A withdrawn search's records are excluded from current reports with its Citations (J); frozen reports keep them.

Fixtures: a search deduplicated outside SyRF before import; a mismatch warning and its explanation; a correction superseding a record while an earlier frozen report stays unchanged; combined external and SyRF deduplication in box 3; a search imported after external title/abstract screening (box 4 and 5 reported, box 6 from SyRF retrieval); an "included elsewhere" import counted in box 10 with Imported authority; a previous-review record switching the template variant.

Releases: P1 records external identification and deduplication counts per search, with the entry phase; R5b adds entry of the other step types and uses all records in reports.

L. ASySD deduplication inside SyRF (Chris, 3 October)

Today: FEAT-012's Approved service specification already makes ASySD the deduplication algorithm, implemented natively in C#:

  • a synchronous DOI/PMID exact match during import (Stage 1), and asynchronous fuzzy matching after import (Stage 2): four-round blocking, Jaro-Winkler similarity on ten bibliographic fields, 25 classification rules, and grouping;
  • AutoConfirmed and ProbableDuplicate tiers, an admin review queue, a merge wizard and an audit log;
  • retroactive deduplication for older projects, and the box 3 derivation and count-consistency rule (spec §§1–4, 7–11).

The Draft brief (docs/features/deduplication/brief.md) still says ASySD runs as an R subprocess; the Approved specification supersedes it. FEAT-012 calls the surviving record of a merge the "canonical Study" (§7, §9); this programme calls it the primary Study, because "canonical" means the engine here.

Proposed amendment (alignment with this programme; rules under Q-37, D4-21 and D2-12):

  1. P2 implements FEAT-012 as specified, except where amendment D and this amendment change it: native C# ASySD, the two-stage pipeline, confidence tiers, the admin review queue, the merge wizard, the audit log and retroactive deduplication.
  2. ASySD parity (D4-21). A conformance suite runs the C# port against reference outputs from a pinned commit of the ASySD R package on the labelled benchmark datasets published with Hair et al. 2023 (dataset names and licences UNVERIFIED; pinned by commit in the fixture) and on SyRF's seeded PRISMA pilot data. The fixture pins the normalisation table (case, punctuation, Unicode, DOI prefix) so "identical" is testable. Pass conditions: identical AutoConfirmed groups on the pinned fixture; ProbableDuplicate pair-set F1 ≥ 0.99; sensitivity and specificity on the labelled datasets each within 0.5 percentage points of the R package (PROPOSAL); every divergent pair listed for review; 80,000 citations processed in under one hour on Bramble. Differences beyond these block the release.
  3. No hard deletes. FEAT-012 scenario 1's table says "delete secondary Study", while its detailed steps keep the secondary with lifecycleStatus = Duplicate. The detailed steps win: the secondary Study is never deleted, which also keeps box 3 counts derivable.
  4. Merge is an alias, never a re-key. Every natural key in the engine carries studyId and revisions are never edited (C1), so a merge cannot move sessions, gold or outcomes. Instead the secondary Study gets mergedInto and the primary gets a StudyAlias set. Reads, reconciliation candidate selection, statistics and PRISMA resolve aliases; the ContributionQualificationPolicy counts a reviewer once across aliased studies (SF2). When one reviewer reviewed both duplicates, the wizard resolves per reviewer: the current session is chosen (default the later Complete; admin choice), the other is superseded with provenance, and the reviewer is counted once; no candidate count is inflated (amendment D). The primary's gold and outcome histories continue; the secondary's become candidates with lineage, never promoted automatically. Merges and splits are ADR-020 operations that write both Study documents and refuse busy studies; split removes the alias and re-derives the primary's projections (E33 restated).
  5. Scenarios by form and profile, not stage. Sessions belong to forms (SF1), so FEAT-012's "same stage" and "different stages" tests (§7.1 scenarios 3 and 4) are restated:
  6. Scenario 1 (neither reviewed, high confidence): auto-confirmed alias merge, as specified.
  7. Scenario 2 (one reviewed, high confidence): admin-reviewed under amendment D (not auto-confirmed as §7.1 says); the reviewed Study is the primary. §8.1's queue population therefore includes scenario 2.
  8. Scenario 3 (both have evidence on at least one shared form or profile): admin-reviewed alias merge with candidate joining and per-reviewer resolution (rule 4).
  9. Scenario 4 (evidence only on disjoint forms and profiles): admin-reviewed alias merge; no candidate collision arises. The former "link via Publication only" outcome remains for the case where the administrator judges the records to be different studies of one publication (then they are reports, amendment O).
  10. Scenario 5 (probable duplicate): unchanged; "Not duplicate" is never re-queued for the same pair and algorithm version.
  11. Admission. Studies with lifecycleStatus ∈ {PendingDedupCheck, PendingDuplicateReview, Duplicate, Merged, RemovedByAutomation, RemovedOther} are excluded by the single admission service (C6) as well as by pool filters (FEAT-012 §12), so no path can offer them; only Active studies enter pools (taxonomy §4 rule 1).
  12. Privacy and enrichment visibility. Cross-project Publication enrichment shares bibliographic metadata only, never review data. Reading a Publication never exposes project or citation IDs from projects the caller cannot access: linkedProjectIds[] and metadataProvenance[].sourceProjectId and sourceCitationId are internal and served only to platform administrators; a project sees "metadata enriched from another SyRF project (not identified)". Each enrichment is a recorded event with a hybrid-logical-clock stamp, so an as-of export states the Publication metadata version it used, and Citations stay the raw, immutable source so an export is reproducible without the Publication.
  13. QC sample and reviewer flag. A configurable share of AutoConfirmed groups (PROPOSAL default 5%, minimum 20 groups) is shown in the review queue for human confirmation; a reviewer action "Flag as possible duplicate of…" creates a DuplicateReviewItem from the screening or extraction page.
  14. Box 3 and transparency. Box 3 combines SyRF-detected and externally reported duplicates (K), and the manifest keeps both parts, plus the ASySD algorithm version (AlgorithmVersion), the tier-rules version, the auto-confirmed versus reviewed share, reversals and the QC sample result (PRISMA-S item 16).
  15. Audit. Audit entries are never deleted (FEAT-012 §10.3); reversal restores the previous lifecycle status and pool membership and is itself audited.

Fixtures: FX-PRISMA-01 (P2 part) and FX-PRISMA-07a; ASySD parity; a reversed duplicate decision restoring pool membership; a merge where one reviewer reviewed both duplicates (counted once); a split after a merge; cross-project enrichment visible without project identity (FX-PRISMA-08a).

Release: P2, after P1 and amendment D; reviewed-record scenarios after R3b.

M. Full-text retrieval

  • Today: study-lifecycle-and-source-taxonomy.md makes FullTextSought (3) and FullTextNotRetrieved (4) lifecycle states (lines 100–102, 148–166) with transitions T4, T12 and T13 (lines 241, 249–250), calls FullTextNotRetrieved terminal (line 258), derives box 6 from lifecycle ∧ title/abstract Included (line 453) and box 7 from lifecycle (line 461), and gives the lifecycle precedence over an Included outcome (lines 597–601). prisma-flow-diagram-mapping.md already derives boxes 6, 7, 8, 12, 13 and 14 from Study.fullTextStatus (lines 126, 132, 138, 152, 158, 164), and three-level-data-model.md defines FullTextStatus {Pending, Sought, Retrieved, NotRetrieved} (lines 211–219) while offering both derivations for box 7 (line 368).
  • Problem: the two derivations disagree; the precedence rule contradicts "lifecycle = pipeline position" and amendment H's per-profile outcomes; and no human action populates the boxes, so box 7 would always be 0 and box 8 would not equal box 6 (SR-02, SR-15).
  • Amendment:
  • Retrieval is fullTextStatus only. Lifecycle never changes for retrieval. T4, T12 and T13 are removed; ordinals 3 and 4 remain reserved and are never written. The precedence rule at lines 597–601 is deleted; the Included transition (taxonomy rule 6) is unaffected by retrieval.
  • Actions (actors and the PDF rule per D4-07): Sought (date); Retrieved (how: PDF in SyRF, read externally); Not retrieved (reason from a controlled list plus free text, and an author-contact date). Reasons (PROPOSAL): not available from any source; paywalled and not obtainable; author contacted, no response; wrong document supplied; language or format not usable; other.
  • Actors: project administrators and reviewers holding a stage grant (capability placeholder, A-03). Each action is an append-only StudyLifecycleEvent with actor and time.
  • Defaults: a title/abstract collective Include sets Pending → Sought (system actor, recorded). Attaching a PDF (manual link, bulk PDF, study-source upload) only suggests Retrieved; a human confirms. Reading the full text outside SyRF is Retrieved with how = external.
  • Admission: full-text steps admit only Retrieved studies (admin override, audited); a Not retrieved study receives no full-text outcome.
  • Boxes (by source column, amendment C): 6 and 12 = title/abstract Included with fullTextStatus ∈ {Sought, Retrieved, NotRetrieved}; 7 and 13 = title/abstract Included with NotRetrieved, with the reason breakdown exported; 8 and 14 = Retrieved ∧ entered the full-text pool. Adopted legacy projects without retrieval history show "retrieval not recorded" coverage, never an inferred Retrieved.
  • Superseded wording for the decision register §2: taxonomy T4/T12/T13, rule 5's FullTextNotRetrieved terminal state, the precedence edge case at lines 597–601, and the lifecycle-based box 6/7 derivations at lines 453 and 461.
  • Fixtures: a study title/abstract-included, marked Not retrieved and never full-text screened appears in boxes 6 and 7 and not in box 8, with its reason in the export (AC-P1-11, AC-R5b-18); a PDF attached without confirmation leaves fullTextStatus = Sought; FX-PRISMA-05b (retrieval change events reproduced as-of).
  • Releases: P1 (events and actions), R3a/R3b (admission), R5b (boxes).
  • Today: three-level-data-model.md makes Citation.publicationId required ("Nullable: No", line 147) and forbids changing a Citation after creation (line 167); Study.publicationId is set from the Citation (line 249). P1 writes immutable Citations before any Publication exists (Publication arrives in P2).
  • Problem: P2 would have to rewrite P1's Citations to link them (V2-01).
  • Amendment:
  • The link is held in an append-only CitationPublicationLink record (citation, publication, how linked: DOI, PMID, fuzzy group, admin; time). Linking never rewrites a Citation.
  • Citation.publicationId becomes optional and write-once at creation, set only when the identifier is known at import; it is never updated afterwards.
  • Study.publicationId stays a mutable pointer, maintained by the dedup operations.
  • Whether P1 also creates Publications for exact DOI/PMID matches (FEAT-012 Stage 1 brought forward) is decided at F-P (E92); a link record is needed either way for Stage 2 results.
  • FEAT-012 Stage 1 and Stage 2 write link records and Publications, never Citations (consistent with spec §1 invariant 3); the Release 3 checklist's "ALL Citations are immutable" check extends to link creation.
  • Fixtures: a P1 Citation linked at P2 without any change to the Citation document (AC-P2); FX-PRISMA-01 (P2 part).
  • Releases: P1 (record shape), P2 (links written by dedup).

O. Report-to-study linkage

  • Today: prisma-flow-diagram-mapping.md lines 56–70 map "study" to a Study after dedup consolidation and "report" to a Study with fullTextStatus; box 10 counts lifecycle Included Studies and reports as the sum of their Citations (lines 178–179). Amendment B (open) fixes report identity, but nothing groups several reports (papers) into one study (SR-07). Cochrane Handbook chapter 4 requires collating reports of the same study.
  • Amendment (conditional on D4-08):
  • A "Link reports to one study" action for administrators and reconcilers creates an append-only StudyLink group (members, reason, provenance, actor); dissolving a group is an appended event. Groups are shown in the study view and in exports.
  • Linking never merges screening or extraction evidence: each report keeps its sessions, outcomes and gold; extraction stays per report with a group key.
  • Counting: a group counts once as a study when at least one member is Included; included members count as reports (with B's report identity); an excluded member stays in box 9 with its reason. new_studies and total_studies resolve groups once; new_reports ≥ new_studies.
  • The duplicate review queue's pair view is reused with the outcome "same study, different report".
  • Fixtures: FX-PRISMA-09 (if D4-08 is approved): two reports linked → 1 study, 2 reports in box 10; an excluded companion report stays in box 9 and the group still counts once.
  • Releases: P2 (action and group), R5b (counting with B).