Return an immutable evidence graph: study bundle to sample to run to raw data to processing and analysis to observations and reports to bounded human review, with missing evidence and corrections preserved.
Return an evidence graph, not only a final report
A final PDF can communicate a conclusion, but it rarely carries enough structure to reconstruct how each value came to exist. A useful result return connects the exact study-bundle version to biological resources, samples, conditions, runs, observations, source files, transformations, summaries and review events. Each object needs a stable identifier and each relationship needs an explicit direction. The result is a navigable evidence graph rather than a folder whose meaning depends on filenames and memory.
NIST's Research Data Framework describes provenance, authoritative copies, derivatives, versions, timestamps, instruments, software and responsible parties as connected research-data concerns. W3C PROV supplies a vocabulary for entities, activities, agents, derivation, revision and invalidation. SingularCell can use those principles to preserve asserted lineage. A complete graph still does not prove that a result is correct, a method is suitable or a conclusion is scientifically justified.
- Identify the package, schema, study bundle and result versions
- Give every sample, run, observation, file and activity a stable ID
- Link every returned object to the exact study context it belongs to
- Preserve the returning legal entity, site, generation time and manifest
- Use neutral states such as returned for review rather than passed
Bind every observation to a sample, run and measurand
A result row is interpretable only when its identity and measurement context are explicit. Record the exact biological resource and material, sample or aliquot, condition, experimental unit, declared replicate relationship, run, measurand, readout, value state, unit and timepoint. Include the event from which the timepoint is measured. Do not infer biological or technical replication from repeated rows; the returning laboratory should declare the relationship and a qualified reviewer should assess whether the classification is scientifically and statistically appropriate.
NIH reporting principles call for distinguishing biological from technical replicates and for reporting relevant cell-line source, authentication and contamination information. NIH also notes that key biological resources can differ across laboratories or over time. Those records make identity questions inspectable, but authentication evidence is bounded to the method and context used. It cannot by itself establish phenotype, purity, viability, function, suitability or equivalence.
- Resource, material, sample, parent-sample and condition identifiers
- Experimental-unit, biological-replicate and technical-replicate declarations
- Run ID, status, site, operator role and acquisition timestamps
- Measurand, readout, value state, unit and timepoint reference event
- Links from each observation to its source or processed data asset
Keep raw, processed and reported data separate
Raw or source data, processed data, analysis outputs and reports are different object classes. Each file-manifest entry should record its logical role, format and version, byte size, hash algorithm and value, source, creation or acquisition time and access classification. The package should name the original authoritative copy and must never silently replace it with a reintegrated, normalized or otherwise processed derivative.
Every processing or analysis activity should name its exact input IDs, versions and hashes; code, pipeline, notebook and configuration versions; software or runtime environment; responsible role; start and end times; warnings or logs; and generated outputs. A hash can detect whether bytes match a recorded value. It cannot prove authorship, custody, completeness, truth or scientific validity, so SingularCell should never use file integrity as a substitute for scientific review.
- Classify each asset as raw source, processed, analysis output or report
- Record format, format version, size, source and cryptographic hash
- Map files to exact samples and acquisition runs
- Bind every derivative to exact inputs and its generating activity
- Preserve original, superseded and invalidated versions
Record method, instrument and software context
A result package should identify the method and protocol references, effective versions, measurement type, instrument and model, configuration, acquisition settings, acquisition software and relevant operating conditions. Where the laboratory supplies them, it can also link instrument-status or calibration evidence, critical reagents, reference materials and their lots. These fields help reviewers understand which configuration produced a returned observation and whether important context is absent.
NIST's cell-characterisation programme highlights measurement context, variability, reproducibility, uncertainty, detection range and limitations. The schema can preserve the laboratory-supplied evidence for those topics, but it should not invent a universal assay profile or threshold. Recording an instrument status, software version or calibration reference does not establish that it was adequate, validated or fit for the study's intended use.
- Method, procedure and protocol identifiers with effective versions
- Instrument, model, configuration and acquisition-settings references
- Acquisition software and version
- Critical reagent and reference-material records where declared relevant
- Measurement range, uncertainty and limitation evidence where supplied
Preserve declared QC, deviations and exclusions without deciding them
Quality-control records should retain the exact laboratory-supplied rule, rule version, owner, observed value or state, units, evidence links and declared status. Use qualified labels such as laboratory declared within rule, laboratory declared outside rule, not assessed, rule not supplied or conflict. A package-level pass or fail would collapse the laboratory's declaration, deterministic calculation and scientific decision into one misleading word.
Deviations, exclusions, failures and reprocessing events require the same discipline. Keep the original value or asset; identify the affected samples, runs, data and reports; record whether an exclusion was prespecified or post hoc; and preserve the actor, time, reason, investigation, impact statement and disposition supplied. FDA data-integrity guidance illustrates retention of originals, invalidated data and reprocessing reasons in its CGMP scope. Borrowing that record pattern does not make a research package CGMP or Part 11 compliant.
- Exact QC rule, version, owner, observed state and supporting evidence
- Laboratory-qualified QC status rather than an unqualified pass or fail
- Deviation type, affected objects, reason and supplied impact statement
- Prespecified or post hoc exclusion status with retained original
- Reprocessing reason and links to both prior and successor outputs
Make uncertainty, limitations and missing evidence explicit
For each uncertainty statement, record what quantity it applies to, the reported numeric or qualitative value, units, estimation method, confidence or coverage information where supplied, responsible party, applicable range and limitations. Not every qualitative or quantitative cell assay uses the same uncertainty representation. The schema should carry what the laboratory supplied and expose what it did not supply rather than manufacture a precision estimate.
Missingness needs its own controlled vocabulary. Distinguish an object not returned from a field not reported, a measurement not acquired, a statistic not calculated, raw data not returned, incomplete metadata, missing sample-to-run mapping, an undeclared replicate relationship, a missing QC rule, a broken provenance link and conflicting records. None of these states should be silently converted to zero, normal, negative, acceptable or failed.
- Scope, value, units and method for every supplied uncertainty statement
- Applicability range, matrix, conditions and known limitations
- Explicit missing-evidence state instead of a blank field
- No silent zero substitution, bound substitution or imputation
- Visible conflicts and unresolved records in summaries and exports
Bind review and corrections to exact versions
A review record should identify the reviewer and role, organisation, exact package and object versions reviewed, evidence considered, review scope, decision label, rationale, time and limitations. The decision vocabulary should remain bounded: reviewed for declared scope is safer than an unqualified approved or validated. If any dependency changes, an earlier review does not automatically migrate to the corrected package.
A correction should create a linked revision or invalidation event that preserves both the prior and successor objects, actor, time, reason, old and new values where relevant, and downstream reports or reviews requiring reconsideration. W3C PROV can represent revision and invalidation relationships, but those are asserted lineage claims rather than proof. The final interface must repeat the central boundary: traceability supports review; it does not determine biological correctness, acceptance, laboratory quality, equivalence or compliance.
- Reviewer identity, role, organisation and exact review scope
- Exact package, object and evidence versions reviewed
- Bounded decision label, rationale, time and limitations
- Correction link, reason, actor and affected downstream outputs
- New review required whenever a relevant version changes
Primary sources
Material claims were checked against the organisations responsible for the guidance or measurement work.
- NIH — Principles and Guidelines for Reporting Preclinical Research ↗National Institutes of Health · Reporting principles for replicates, statistics context, exclusions, data sharing and cell-line identity information.
- NIH — Authentication of Key Biological and/or Chemical Resources ↗National Institutes of Health · Study-specific identity and authentication planning for key resources that may differ by laboratory or over time.
- NIST — Research Data Framework, Version 2.0 ↗National Institute of Standards and Technology · Research-data lifecycle, authoritative copies, metadata, provenance, versions, derivatives, instruments, software, timestamps and responsible parties.
- NIST — Building Measurement Confidence in Cell Characterization ↗National Institute of Standards and Technology · Measurement context, variability, reproducibility, uncertainty, detection range and limitations for cell characterization.
- FDA — Data Integrity and Compliance With Drug CGMP ↗United States Food and Drug Administration · Regulated-context principles for metadata, originals, audit trails, invalidated data, corrections and reprocessing records.
- W3C — PROV-O: The PROV Ontology ↗World Wide Web Consortium · Standard relationships for entities, activities, agents, generation, use, derivation, revision, attribution and invalidation.
Limitations
- This guide proposes a research-use documentation schema; it does not prescribe an assay, protocol, threshold, QC rule, analysis or scientific acceptance criterion.
- Complete metadata, provenance and file hashes improve reconstruction but do not prove data accuracy, biological validity, authenticity, authorship or reproducibility.
- NIH, NIST and W3C sources provide reporting, framework and provenance principles; FDA examples are bounded to CGMP and do not make this schema compliant with any regulation or standard.
- Only qualified reviewers using context-specific evidence can assess method suitability, laboratory quality, site or method equivalence, deviations, exclusions and biological conclusions.
See the handoff as a working system.
Explore one synthetic study from research question through capability comparison, returned results and review-required evidence.