Reproducibility Manifest¶
Every Portiere pipeline run emits a versioned manifest.lock.json that captures exactly what produced the output — model identity, vocabulary fingerprints, source-data hash, threshold snapshot, per-stage entries — so a run can be reproduced by command months later.
Reproducibility in Portiere means the same inputs and configuration produce the same pipeline state. It does not promise byte-identical outputs: clinical pipelines have legitimate nondeterminism (LLM sampling, BM25 tie-breaks, FAISS thread effects). Outputs reproduce within a small tolerance.
What's recorded¶
The manifest is a strict Pydantic-validated JSON document with a versioned schema (manifest_version: "1"). Every field is either present or explicitly null.
Always recorded¶
| Field | Example | Notes |
|---|---|---|
manifest_version |
"1" |
Bumped on breaking schema changes |
run.run_id |
"abc123def456" |
12-hex-char identifier per run |
run.started_at / finished_at / duration_seconds |
ISO-8601 / float | UTC timestamps |
portiere_version / python_version / os_string |
"0.2.0" / "3.12.1" / "Darwin-25.3.0" |
Environment fingerprint |
git_sha / git_dirty |
hex / bool, or null |
null when not a git repo |
project_name / target_model / task |
strings | task is "standardize" or "cross_map" |
source_standard |
"omop_cdm_v5.4" or null |
Set when task == "cross_map" |
vocabularies_requested |
["SNOMED", "LOINC"] |
The vocab list passed at init() |
embedding |
{name, hf_revision, dimension, sha256_of_config} |
Identity tuple — see below |
knowledge_backend |
{type, index_hash} |
"bm25s" \| "faiss" \| ... |
vocabularies |
[{name, version_date, sha256_of_source_file, path}] |
Per-vocab fingerprint |
prompt_templates |
[{name, sha256}] |
Hash only — template text never stored |
thresholds |
full ThresholdsConfig snapshot | |
source_data |
file: {path, sha256} / DB: {connection_string_redacted, table_or_query} |
Credentials never enter the schema |
stages |
per-stage {stage, started_at, finished_at, inputs, outputs, metrics} |
One entry per pipeline op |
Never recorded¶
API keys (OpenAI / Anthropic / AWS credentials), unredacted database connection strings, prompt template text, embedding model weight bytes (only the identity tuple — see below).
A regression test (tests/test_recorder.py::TestRecorderNoSecretLeak) sets OPENAI_API_KEY / ANTHROPIC_API_KEY / AWS_SECRET_ACCESS_KEY in the environment and asserts none of those values appear in the emitted manifest. Belt-and-suspenders.
Why the embedding is an identity tuple, not a hash of weights¶
SapBERT is ~440 MB. Hashing it on every run would add seconds and convey no more identity than the HF revision SHA already does. The manifest records {name, hf_revision, dimension} — three small strings — and trusts Hugging Face to deliver the same weights for the same revision.
When a model is local-only (no HF metadata), sha256_of_config is an optional fallback that hashes config.json (a few KB).
Where it lives¶
<config.local_project_dir>/<project_name>/runs/<run_id>/manifest.lock.json
A run begins on the first pipeline op (profile, map_schema, map_concepts, run_etl, or validate) and is finalized when:
Projectis used as a context manager and__exit__fires (recommended), orproject.finalize_run()is called explicitly.
import portiere
from portiere.engines import PolarsEngine
with portiere.init(name="hospital-migration", engine=PolarsEngine()) as project:
source = project.add_source("patients.csv")
schema_map = project.map_schema(source)
concept_map = project.map_concepts(source=source)
project.run_etl(source, output_dir="out/")
project.validate(output_path="out/")
# manifest finalized here, written to runs/<run_id>/manifest.lock.json
Replay¶
portiere replay <project>/runs/<run_id>/manifest.lock.json
Replay does three things:
-
Validate referenced artifacts. Every recorded vocab path and the source data path must exist; sha256 hashes must match. If anything's missing or different, replay raises
ManifestReplayErrorwith a clear message:Replay failed: source data sha256 mismatch: /path/to/patients.csv (manifest says 'a4f3...'; current file differs) -
Reconstruct the project. A new
portiere.init()is called with the manifest's recordedtarget_model,task,source_standard, andvocabularies_requested. The new project gets areplay-prefixed name to avoid clobbering the original's storage. -
Re-attach the source. If
source_data.pathpoints to an existing file, it's added to the new project. From there, the caller can re-invoke pipeline ops as needed.
Auto-replay (--auto-replay, v0.3.1+)¶
Pass --auto-replay to additionally re-run each recorded stage and compare outputs to the manifest within per-stage tolerance bands:
| Stage | Tolerance |
|---|---|
ingest |
Source-data sha256 verified upstream; format must match the recorded value. |
schema |
n_mappings identical. |
concept |
n_mappings must not drop; auto_rate within ±1% absolute. |
etl |
row_count and column set identical (paths may differ). |
validate |
all_passed flag identical (binary). |
Stages whose dependencies are not available in the current environment (e.g., concept requires the same LLM + knowledge layer as the original run; validate requires the recorded output directory to still exist) are recorded as UNAVAILABLE and do not fail the report. v0.3.1's auto_replay() is honest about what it can verify without a live re-execution environment; the full re-execution path for the LLM-bound stages is a v0.3.x roadmap item once we settle on how to seed BYO-LLM config from a manifest.
Exit codes: 0 = all attempted stages passed (or were UNAVAILABLE); 1 = at least one stage failed comparison; 2 = artifact-verification failed (same as plain replay).
Common replay errors¶
| Error | Cause | Fix |
|---|---|---|
source data missing |
Source file moved or deleted | Restore the file or update the manifest path |
vocabulary missing |
Athena vocab moved or deleted | Restore the vocab directory |
vocabulary sha256 mismatch |
Vocab file content changed | Re-export from Athena to the same release date |
source data sha256 mismatch |
Source CSV was overwritten | Restore from backup |
What replay does NOT promise¶
- Bit-identical outputs. LLM sampling, BM25 tie-breaks, and FAISS thread effects all introduce small nondeterminism. The same input, same model revision, same vocab will produce very similar mappings, but the per-row decisions can flicker at the edge of confidence thresholds.
- Perfect environment recreation. Replay validates artifact identity but doesn't reinstall Python packages or pin transitive deps. Pin your environment with
pip freeze > requirements.lock.txtfor that level of fidelity. - Full stage replay for LLM-bound stages. v0.3.1's
--auto-replayre-runs cheap deterministic stages (ingest,validatewhen the output directory still exists) and records LLM-dependent stages (schema,concept,etl) asUNAVAILABLE. End-to-end re-execution of those — including BYO-LLM config rehydration from the manifest — is a v0.3.x roadmap item.
Audit / compliance use¶
The manifest is the right artifact to attach to an IRB submission, a publication's data-availability statement, or an internal change-management ticket. It answers:
- Which model was used?
- Which vocabulary version?
- Which source data (by hash, not just path)?
- Which thresholds?
- When was each stage run, with what inputs and outputs?
It does not answer "what were the per-row mapping decisions" — those live in the concept_mapping.csv / schema_mapping.csv artifacts, which the manifest's stage entries can point to via outputs.
See also¶
- Spec §4.2 (v0.2.0 release design, maintainer archive) — full design rationale
portiere replayCLI — invocation reference- Quickstart manifest example — interactive walkthrough