Context
Evaluation runs preserve the visible answer in runs/*/results.jsonl and the full provider payload/reasoning trace in the content-addressed data/response_cache/. Those paths are intentionally Git-ignored. During the Nemotron 3 Ultra evaluation tracked in #6, we decided the traces are scientifically valuable because reproducing them can be expensive in both wall time and API cost.
This issue is for the storage decision only. It should remain separate from model evaluation and response-viewer implementation.
Current evidence
An early sample of three Nemotron cache entries measured:
- 78,875 bytes uncompressed
- 11,474 bytes as a gzip-compressed tar archive
- 14.5% compressed/raw ratio
- straight-line 54-response projection: ~1.35 MiB raw / ~0.20 MiB compressed
Later responses include reasoning traces around 25k characters and individual calls taking roughly nine minutes, so storage size is currently modest while regeneration cost is not.
Storage properties we need
- Preserve the exact successful provider payload and reasoning trace.
- Bind traces to the run config, dataset fingerprint, cache key, and visible result row.
- Verify integrity with SHA-256 checksums.
- Avoid committing the mutable shared cache, locks, partial state, unrelated models, or regenerated gold answers.
- Let collaborators inspect traces from another clone/session.
- Ideally allow exact cache restoration so a preserved response prevents another provider call.
- Keep API credentials and secrets out of every artifact.
Options to compare
-
Compressed immutable artifacts in Git
- One
.tar.gz/.tar.zst per completed run plus a readable manifest.
- Lowest operational overhead and likely appropriate during exploration.
- Binary artifacts do not diff and will eventually bloat repository history.
-
Git LFS
- Keeps large blobs outside ordinary Git objects while retaining repository references.
- Adds quota, tooling, and clone/setup considerations.
-
GitHub Release assets
- Durable downloadable bundles without bloating Git history.
- Awkward lifecycle/tag semantics for frequent exploratory runs.
-
Content-addressed object storage
- Store each compressed cache entry once under its cache key; run manifests reference shared blobs.
- Best deduplication and long-term scale, but introduces credentials, lifecycle policy, availability, and implementation work.
Proposed exploratory default
Until measurements show otherwise, check in one immutable compressed bundle per completed run under experiments/model-traces/, accompanied by a readable manifest and checksums. Do not check in data/response_cache/ itself. Revisit LFS/object storage when a single bundle or the accumulated trace history becomes inconvenient enough to affect normal clones and reviews.
Questions / decision criteria
- Expected number of models, reasoning-effort variants, repeats, and dataset versions.
- Actual compressed size distribution after complete runs, especially 64k-token traces.
- Whether traces need private access or have provider redistribution constraints.
- Desired one-command export, verify, inspect, and restore workflow.
- Whether the response visualizer should read archives directly or only extracted/local cache entries.
- Thresholds that trigger migration away from ordinary Git.
Completion criteria
Context
Evaluation runs preserve the visible answer in
runs/*/results.jsonland the full provider payload/reasoning trace in the content-addresseddata/response_cache/. Those paths are intentionally Git-ignored. During the Nemotron 3 Ultra evaluation tracked in #6, we decided the traces are scientifically valuable because reproducing them can be expensive in both wall time and API cost.This issue is for the storage decision only. It should remain separate from model evaluation and response-viewer implementation.
Current evidence
An early sample of three Nemotron cache entries measured:
Later responses include reasoning traces around 25k characters and individual calls taking roughly nine minutes, so storage size is currently modest while regeneration cost is not.
Storage properties we need
Options to compare
Compressed immutable artifacts in Git
.tar.gz/.tar.zstper completed run plus a readable manifest.Git LFS
GitHub Release assets
Content-addressed object storage
Proposed exploratory default
Until measurements show otherwise, check in one immutable compressed bundle per completed run under
experiments/model-traces/, accompanied by a readable manifest and checksums. Do not check indata/response_cache/itself. Revisit LFS/object storage when a single bundle or the accumulated trace history becomes inconvenient enough to affect normal clones and reviews.Questions / decision criteria
Completion criteria