Skip to content

Status correction — 3 October 2026: UNAPPROVED PRELIMINARY MATERIAL. Retained as research input only. Chris has limited the current task to collecting future-planning scope, sources, confirmed decisions and unresolved matters. Actual implementation planning is a separate phase he will initiate. Architecture, sequencing and proposed controls here are not approved or an implementation approval request. Separately recorded owner decisions still apply.

Outcome-data migration plan proposal — 3 October 2026

Planning authorized; execution not authorized. This supersedes P9's deferred planning status, not its prohibition on running a migration. Chris requested a concrete plan based on legacy extraction, the prototype and current research. All target representations, cutover mechanics and defaults below are recommendations for approval. See the owner ledger and permission matrix.

MVP boundary, acceptance and shortest path

First usable implementation slice: a read-only inventory and deterministic dry-run adapter for one selected project's legacy-compatible outcome schema, comparing candidate and reconciled rows against current exports. It writes no review answers and changes no form binding. Acceptance: all selected source rows accounted for, IDs/ancestor paths and source provenance preserved, unknown history labeled, no numeric conversion or candidate-to-gold promotion, mismatch report reviewable. Then implement approved staged copy and project-scoped cutover, retaining originals and rollback routing. Event-count schemas and project customization are required first-release capabilities, but existing continuous rows are never silently relabeled as event counts. Separate delivery slices keep schema creation usable independently from historical conversion.

Verified source and gaps

OutcomeData contains StageId, OutcomeId, CohortId, ExperimentId, InvestigatorId, ProjectId, Reconciled, GraphId, Units, AverageType, ErrorType, NumberOfAnimals, GreaterIsWorse and TimePoints; Study/SchemaVersion are reached via extraction information. This is a fixed quantitative record with annotation references, not proof of versioned target-series definitions. TimePoint stores numeric Time/Average/Error; DTO Average/Error are nullable. The dry-run must inspect real persisted null/default behavior; zero must not be guessed to mean missing. Outcome export row writer resolves cohort, experiment and outcome annotation references, exposes InvestigatorId and Reconciled, and uses legacy categories including DiseaseModelInduction and Treatment. The existing export-history investigation finds reserved as-of modes, not verified full runtime history support. Prototype mappings show intended types/relationships, not production migration.

Required discovery before execution: actual stored shape per legacy Study.SchemaVersion; write endpoints/graph extraction and import paths; reference identity and reconciliation rules; existing immutable history coverage; exact export contract and active readers/writers. These gaps are engineering research and block execution, not reasons to fabricate legacy versions.

Confirmed target boundaries

  • Initial schemas: legacy-compatible, event-count, and project-created/customized schemas. Numerator/denominator, range, CI, standard error and variation are configurable examples, not a mandatory bundle. Variation is not automatically variance or SD.
  • Schema versions define series-level fields separately from per-observation fields, data types, validation and cardinality. Meaning used by UI/export requires explicit semantic roles. A reviewer supplies values, not validator code.
  • Cohort/outcome measure/experiment system types are required only when the relevant extraction/ export feature is used. Default templates map existing disease model, disease-model intervention and treatment types; verify catalogue aliases, do not equate Intervention with DiseaseModelInduction solely by wording. Keep custom types/project questions.
  • Entities retain label annotations and child trees, stable lookup identity and population boundaries. Classifier relationship inference does not imply outcome observations or counts.
  • Form sessions/requirements, candidate answers and accepted gold remain separate and pinned. Compatible answer sharing must retain exact repeated-entity/ancestor context. Never merge solely by label, question ID or value.

PRISMA compatibility gate added before mapping/cutover

PRISMA comparison is a required migration dependency. Outcome migration must preserve immutable Citation/import/search source lineage, system Publication links and project Study identity without executing bibliographic dedup. Map legacy screenings to profile outcomes only with evidenced profile/context/authority; legacy thresholds are not invented historical profiles. Keep retrieval/lifecycle, protocol pool entry, per-profile outcome and animal population/cohort concepts separate. No form target, inferred cohort or series count becomes a PRISMA Study/report count. Verify candidate/reconciled evidence and same-reviewer/form dedup before adopting reviewed-duplicate data; conflicting imports/reviews require admin mapping and never auto promotion. Unknown source/report/retrieval history stays unknown; removing new fields is not safe rollback after canonical writes. Apply FEAT-011 affected release checklists, and report-unit/source query amendments must precede complete PRISMA export claims. Extend dry-run manifest with those identities, available-history coverage and before/after PRISMA semantic populations.

Legacy source Target proposal Preservation / failure handling
OutcomeData.Id Stable imported-series identity + source reference Idempotency key includes project/study/source ID, mapping version and target schema version
Outcome/Cohort/Experiment IDs Entity lookup annotations on series Resolve exact source IDs/path; quarantine dangling or ambiguous references, no label matching
InvestigatorId and Reconciled Source author and candidate/legacy-reconciled provenance Retain flags separately; Reconciled alone never invents an accepted snapshot or resolver history
StageId / extraction Study / ProjectId Source context plus target binding references Preserve originating stage; deduplicate shared-form contributions only with proven identity
Units / AverageType / ErrorType Series-level typed values under legacy-compatible roles Preserve exact code/text, define known aliases explicitly; unknown codes remain visible unmapped
NumberOfAnimals Series sample-size value Do not silently overwrite cohort enrolment count; may refer to different outcome/time sample
GreaterIsWorse Original series value + proposed outcome-measure direction mapping Detect conflicting series values; retain per-source value, unresolved measure mapping, never choose majority
TimePoints[index] Stable ordered observation children with time/average/error roles Preserve duplicates/order; index in source path, no grouping by identical time
GraphId Source graph reference Retain relation/availability; do not regenerate extraction evidence
SchemaVersion / recoverable revisions Observed legacy version metadata + migration record Explicitly mark unavailable authoring/submission history; migration time is not original time

Recommended target type retains legacy field permissiveness for imported evidence, recording validation findings separately. Future stricter schema requirements affect new completion policy; do not reject/delete previously stored facts or treat blanks as N/A. No conversion of error type, unit, mean, or event denominator without an approved mapping and recorded transformation. Confirmed single outcome-measure direction and conflict-handling recommendation are in the settings proposal.

Worked mapping example — synthetic, not a recovered paper

A legacy candidate series has source ID series-17, cohort cohort-A, outcome pain-1, experiment exp-1, units score, AverageType mean, ErrorType SE, NumberOfAnimals 12, GreaterIsWorse true, and ordered rows (time 0, average 4, error 0.5) and (time 7, average 3, error 0.4). Target shape below is an illustrative mapping record, not the final JSON Schema:

{
  "schemaRef": {"id": "legacy-continuous", "version": 1},
  "source": {"id": "series-17", "kind": "candidate", "historyCoverage": "observed-record-only"},
  "references": {"cohort": "cohort-A", "outcomeMeasure": "pain-1", "experiment": "exp-1"},
  "seriesValues": {"unit": "score", "averageType": "mean", "errorType": "SE", "sampleSize": 12},
  "observations": [
    {"sourcePath": "TimePoints[0]", "time": 0, "average": 4, "error": 0.5},
    {"sourcePath": "TimePoints[1]", "time": 7, "average": 3, "error": 0.4}
  ],
  "directionMapping": {"proposedMeasureValue": "higher-is-worse", "sourceValue": true},
  "migration": {"mappingVersion": 1, "state": "dry-run"}
}

Carry real author/project/study/stage/graph IDs when present; omitted here for brevity. Do not assert two completed submissions or a historical gold snapshot from this single row. A separate reconciled row remains a separate source. If one legacy series says GreaterIsWorse false for the same outcome measure, block automatic measure consolidation and present both values. If legacy sample size is 12 while the cohort enrolment count is 20, retain both; they are not a migration arithmetic conflict without evidence they measure the same sample.

Phased execution proposal and review gates

  1. Inventory/dry-run: enumerate by project/study/legacy version and author/reconciled status; capture source hashes and available-history coverage; compare current exports. Produce counts, unresolved mappings, semantic conflicts and candidate/gold separation report. No writes.
  2. Approve mapping manifest: exact target schema/question versions, field semantic roles, legacy aliases, reference map, preserved unknowns and explicitly excluded records with reasons. Admin chooses unresolved-record treatment; recommendation retain them on legacy read route, never silently omit from exports. Obtain owner/delegated migration approval, not just Design.
  3. Staged copy: deterministic idempotent batches into non-authoritative target storage; checkpoint source watermark/hash, lineage and mapping version. Rerun creates no duplicates. New target schema validation findings stay attached, never rewrite source evidence.
  4. Verify/read shadow: compare semantic values, counts, reference paths, candidate/reconciled flags and export results, not byte-identical reformatted JSON. Check references/children and permissions. Validate all pinned sessions/gold read original exact versions. No gold promotion.
  5. Scoped cutover: recommendation project-scoped write fence at actual cutover only, not a programme-wide pause. Finish or safely reject in-flight writes before watermark; refresh delta, verify source hashes, then atomically switch approved read/write binding. Stale clients receive explicit retry/update message retaining drafts. Do not assume unchecked dual writes are atomic. Projects outside scope remain unchanged.
  6. Monitor/recover: resume checkpoints after failure; quarantine mismatches with source IDs. Retain originals, manifests and mapping adapters. Do not delete source records after success. Read current and all recoverable pre-migration history under EX2; unavailable old events remain unavailable, never synthetic original submissions.

Publication choices autoUpdate/requireReanswer/doNothing are version-adoption decisions, not migration commands. If migration also changes active schema/form bindings, perform its separate usage/impact publication gate and lifecycle approvals. A migration record cannot silently change reviewer completion, shared target sufficiency, accepted gold or how informed exposure is counted.

Invariants, recovery and rollback proposal

  • Every source row maps once or has a visible unresolved disposition; every observation path retained; no duplicate counting from shared sessions, history, inherited answers or repeated times.
  • Source candidate and accepted/reconciled provenance never collapse; legacy unknown authoring/ approval histories explicitly unknown. Exact known immutable session/gold refs stay unchanged.
  • Schema/question/mapping versions immutable; counts, units, time and missingness preserved. Transformation, if explicitly approved, retains originals and operation/rationale.
  • No cross-population instance links, inferred partition/count promotion or identity merging by name.
  • Dry-run and rerun produce stable mapping keys/hashes; source changes invalidate approval/preview.
  • Permission tests cover legacy and target read/export paths; authorization is not copied blindly from an obsolete stored grant. Historical manifests identify recoverable coverage and as-of basis.

Rollback recommendation: before new target writes, switch routing back to intact legacy data. After target writes, retain target as authoritative for those records until a verified reverse adapter or forward recovery is approved; blind routing rollback would lose new semantics. Keep both stores/read paths and journal the cutover boundary. No destructive down-migration or fabricated reverse conversion of custom/event fields. Rehearse failure between copy, validation, fence, delta and pointer commit in a non-production fixture before any authorized execution.

Approval items and ordered follow-up work

Approve: read-only legacy adapter MVP; mapping invariants/unknown-history treatment; retaining unresolved records on legacy route; explicit execution grant; staged-copy rather than blanket rewrite; project-scoped write fence and manifest revalidation; rollback boundary after new writes. Approve linked PRISMA A–F amendments and count-unit/source/retrieval coverage where affected; outcome mapping does not independently settle bibliographic/report identity. Approve semantic mappings only after dry-run evidence, especially missing/default numeric values, GreaterIsWorse conflicts and legacy type aliases.

Follow-ups in order: real persisted-shape/import/writer audit; exact typed schema and validation contract; dry-run/export parity; approved execution/cutover and recovery; event/custom schema migration adapters where evidence supports them; optional performance/operational enhancements. No cutover date, migration run, live-data access or implementation approval is implied here.