ostia is a profiling and benchmarking toolkit for Bun. It times subprocess commands
(like hyperfine) and in-process functions (like mitata), optionally captures CPU
profiles, heap snapshots, JIT tiers, retained heap and peak memory, and writes everything
to one schema-versioned JSON document (ProfileDocument). Two documents compare with a
bootstrap confidence interval and a Mann-Whitney test, with the regression threshold
widened to the machine's measured noise floor. ostia ab pairs the working tree against
a git ref in one process, which holds up on machines too noisy for that. ostia ci gates
a config file of workloads against a saved baseline, and --format minimal gives scripts
and LLM agents a compact JSON line protocol. The CLI is a thin wrapper over the library, so anything
ostia time/ostia bench do, time()/bench() do too.
Zero runtime dependencies. Requires Bun ≥ 1.4.
bun add ostiaostia time --samples 10 "bun fixtures/fast.ts" "bun fixtures/slow.ts"Apple M2 · 8 cores · load 3.6 · noise floor 0.5%
Task Median Spread Range User/Sys Relative
---------------------------------------------------------------------------------------------------
bun fixtures/fast.ts 9.38 ms 9.52 ms…9.74 ms 9.10 ms…9.76 ms 6.98 ms/3.06 ms 1.00×
bun fixtures/slow.ts 23.2 ms 23.3 ms…23.8 ms 23.0 ms…23.9 ms 20.9 ms/2.89 ms 2.47× slower
! outliers-detected
Warnings:
bun fixtures/slow.ts: 1 outlier(s) detected (1 severe, 0 mild).
The header line shows the machine, its load average, and the noise floor from a ~200ms
reference measurement taken once per run (--no-noise-check skips it). Spread is
p75…p99; User/Sys is the median user/system CPU time per trial.
// suite.ts
import { group, task } from "ostia"
const input = Array.from({ length: 2_000 }, (_, i) => i % 500)
group("dedupe", () => {
task("naive (indexOf scan, O(n²))", () => dedupeNaive(input))
task("Set-based (O(n))", () => [...new Set(input)])
})ostia bench suite.tsApple M2 · 8 cores · load 3.5 · noise floor 0.4%
Task Median Spread Range Relative
-----------------------------------------------------------------------------------------
dedupe:
naive (indexOf scan, O(n²)) 173.9 µs 180.4 µs…205.8 µs 170.1 µs…635.5 µs 7.40× slower
Set-based (O(n)) 23.5 µs 24.9 µs…72.5 µs 18.7 µs…208.3 µs 1.00×
Commit the suite, then change the code: say the Set-based dedupe becomes
input.filter((x, i) => input.indexOf(x) === i).
ostia ab suite.ts # every task: working tree vs HEAD, paired in one processA/B: working tree vs HEAD (3945768) · 15 rounds · threshold 10% · geomean threshold 1.5%
Task Base Candidate Change p25…p75 Verdict
------------------------------------------------------------------------------------------
dedupe:
naive (indexOf scan, O(n²)) 183.5 µs 178.5 µs -1.9% -3.5%…-0.3%
Set-based (O(n)) 25.2 µs 181.2 µs +618.1% +554.2%…+659.2% regressed, confirmed (repeats: +630.6%, +584.5%)
Geomean +165.4% (threshold 1.5%) · 1 regressed, 0 improved, 1 unchanged of 2 · fail
Base and candidate run in alternating ~10ms batches, so machine drift cancels within each round; a flagged task counts only if two fresh processes agree. Exit 1 on a regression.
// ostia.config.json
{
"samples": 10,
"workloads": [
{ "label": "work", "command": ["bun", "fixtures/work.ts"], "inputs": ["fixtures/**"] }
]
}ostia baseline save # on known-good code: writes .ostia/baselines/main.json
ostia ci # on your change: exit 1 on a regression1 workloads
0 cached
1 executed
0 passed 1 regressed (+44.1% median on work)
Profile CI: ✗
...
✗ work
timing: +44.1% median, 95% CI [+41.4%, +45.6%], p<0.001 (regressed)
Scratch output (cache, artifacts) goes to node_modules/.cache/ostia. Baselines go to
.ostia/baselines/ so they survive reinstalls; add .ostia/ to .gitignore.
--format minimal (on time, bench, ab, compare, report, ci) prints one JSON
object per line on stdout and nothing else. Every line has event and
protocolVersion: 2. Timing values are in nanoseconds.
ostia time --samples 10 "bun a.ts" --format minimal
ostia compare before.json after.json --format minimal
ostia ci --format minimal; echo $?event |
When | Key fields |
|---|---|---|
run |
One per timing measurement, every command | workloadId, task, group?, params?, skipped?, unit, samples (0 when no trial produced one), batch, mean/median/stddev/stddevPct/min/max/p75/p99/mad, userNs/systemNs (subprocess only), retainedBytesPerOp?/peakBytes? (--alloc/--peak-mem), relative?, noiseFloorPct?, warnings[], threw? (ab, a task that threw), on compare/ci: delta: { medianPct, meanPct, verdict, pass, ci95?, pValue?, effectiveTimingPct, matched }, and on ab: paired: { baseMedian, medianRatio, ratioP25, ratioP75, rounds, verdict, flagged?, confirmed?, repeats?, sameOutput, suiteChanged?, retained?, peak? } |
unmatched |
One per workload on only one side of compare/ci/ab |
workloadId, task, side: "base" | "cand" |
summary |
Last line of compare/ci/ab only |
command, matched/regressed/improved/unchanged/unmatched, cached/executed/failed/missingBaseline (ci), geomeanPct, effectiveTimingPct, noiseFloorPct?, baseline? (ci), base?/geomeanThresholdPct?/unconfirmed?/outputDiffers?/notComparable?/threw?/newSuites?/memory? (ab), git?, exportedTo?, verdict, exitCode |
{"event":"run","protocolVersion":2,"schemaVersion":2,"workloadId":"wl_11e8562f3622d528","task":"work","unit":"ns","samples":10,"batch":1,"mean":21012800,"median":20999900,"stddev":231456,"stddevPct":1.1015,"min":20664000,"max":21552300,"warnings":[{"code":"outliers-detected","data":{"mild":1,"severe":0}}],"p75":21086100,"p99":21517600,"mad":126625,"userNs":15519000,"systemNs":6015500,"noiseFloorPct":2.09286,"delta":{"medianPct":44.0989,"meanPct":43.9626,"verdict":"regressed","pass":false,"effectiveTimingPct":10,"matched":true,"ci95":[41.4394,45.5841],"pValue":0.000157103}}
{"event":"summary","protocolVersion":2,"command":"ci","matched":1,"regressed":1,"improved":0,"unchanged":0,"unmatched":0,"geomeanPct":44.098920968212305,"effectiveTimingPct":10,"verdict":"fail","exitCode":1,"cached":1,"executed":0,"failed":0,"missingBaseline":0,"baseline":{"name":"main","path":".ostia/baselines/main.json"},"noiseFloorPct":2.09286}
Within protocolVersion: 2, keys are only ever added, never renamed or removed.
Exit codes, the same for every command:
| Code | Meaning |
|---|---|
0 |
Pass |
1 |
At least one workload regressed (compare/ci/ab only; time/bench never return 1) |
2 |
Harness error: a command exited non-zero or produced no samples, a suite failed, nothing matched, a bad flag, a missing/invalid config or baseline |
130 |
Cancelled with Ctrl-C (time/bench/ab/ci; partial results are still exported) |
On exit 2, stderr's last line is {"event":"error","protocolVersion":2,"code":...,"message":...,"data"?:...}
when stderr is not a TTY or a machine format (minimal/json/jsonl) was requested.
A person at a terminal sees only the prose message. code is one of invalid-flag,
config-missing, config-invalid, baseline-missing, no-matches, spawn-failed, suite-failed,
command-failed, timeout, time-source-no-match, document-load-failed,
no-cpu-evidence, internal. Full reference: docs/agent-protocol.md.
Every command takes --help. Per-flag detail is in docs/cli.md.
Times commands as subprocesses. --cpu/--heap add one separate instrumented trial each;
the profiler never runs during timing trials.
ostia time "bun a.ts" "bun b.ts"
ostia time --samples 25 --cpu --heap "bun src/server.ts"
ostia time --prepare "rm -rf dist" "bun build.ts"
ostia time --time-source "built in (\d+)ms" "bun build.ts"- Each command string is whitespace-split into argv, with no shell. Everything after
--is one more command's argv, verbatim:ostia time -- bun -e "console.log('a b')". - Default sampling: 3 warmup trials, then trials until ~3s have elapsed and at least 10
ran.
--samples Ngives an exact count per command;--budget MS/--min-samples Ntune the loop. - With 2+ commands, trials round-robin across commands (
--no-interleaveto run them one after another). - A command stops at its first non-ignored non-zero exit, and
ostia timeexits 2.--ignore-failure[=CODE,...]treats the listed codes (bare: all) as success.
Runs in-process group()/task() suites. Each suite file runs in its own child process.
ostia bench bench/*.ts
ostia bench bench/*.ts --filter parse --cpu --alloc
ostia bench bench/*.ts --filter large --peak-mem
ostia bench bench/*.ts --isolate
ostia bench --preload ./bench/dom-setup.ts --bun-flags="--conditions=browser" bench/*.ts- Each task samples for
--budgetms (default 500). Fast calls are batched so one trial spans at least 1µs; the budget-driven loop stops at 20,000 trials. --isolateruns every task in its own process, isolating JIT state, builtin call-site feedback (e.g.Array.prototype.map) and GC heap from other tasks. Use it when you need the most comparable numbers.--jobs N|autoruns suite files in parallel. Faster, but noisier; keep the default of 1 for anything youcompareor gate inci.--cpuprofiles each task at 100µs for about 2,000 samples (--cpu-intervalchanges the interval). Inlined helpers count as their callers' self time.--allocreports the heap each call retains after a full GC: a leak check, not an allocation count.--peak-memreports how far the task's first call raises RSS, garbage included, in 3 fresh processes (OSTIA_PEAK_MEM=1is set there, so a suite can skip heavy setup that would peak first).- With no files,
ostia benchuses the config'sbenchsection. Each flag overrides its config field;--no-gc/--no-cpu/--no-alloc/--no-peak-mem/--no-isolateoverride a configtrue.
Runs suites on a git ref's committed tree and on the working tree in one process, alternating short batches, and gates on the per-round time ratio.
ostia ab bench/*.ts # vs HEAD
ostia ab bench/*.ts --base origin/main
ostia ab bench/*.ts --threshold 5 --rounds 21- The same suite files as
ostia bench, unchanged. The ref's tree is extracted once per commit undernode_modules/.cache/ostia/ab/; relative imports resolve within each tree, package imports to the project'snode_modules. - Files git doesn't hold (generated or gitignored sources) aren't in the base tree.
--base-setup "bun scripts/generate.ts"(orab.setupin the config) builds them once per tree, with the project'snode_moduleslinked in. - A task is flagged when its median ratio moves past
--threshold(default 10%) in at least three quarters of rounds, and counts only if--confirm(default 2) fresh processes agree. Each suite runs once with each side first, since the side that goes first can run faster. The run also fails when the geometric mean of all ratios is more than--geomean-threshold(default 1.5%) slower. - Tasks whose first call returns different values on each side are listed (not a
failure). The base side runs the base's copy of each suite, so a task whose suite file
changed is marked
suite-changed; if its output changed too, it readsnot comparableand stays out of the verdict. A task that throws readsbase threw,candidate threworboth threwand isn't timed; one that throws on the candidate side only fails the run. Exit:0pass,1regression (time or memory),2nothing paired or a harness error. --allocand--peak-memcompare memory too: retained heap per call, and how far the first call raises RSS. A reading that grows past--mem-threshold(default 10%) and past its noise floor fails the run.- Progress goes to stderr on a terminal;
--progressturns it on in logs and pipes too. - The 5 most recently used base trees stay cached (
--keep-trees);ostia ab --cleanremoves them all.
Matches two documents' workloads by id and reports a verdict per workload.
ostia compare before.json after.json
ostia compare after.json --baseline .ostia/baselines/main.json
ostia compare before.json after.json --format markdown✗ bun fixtures/work.ts
timing: +23.8% median, 95% CI [+18.3%, +30.1%], p<0.001 (regressed)
A regression needs the whole 95% CI above the threshold and a Mann-Whitney p-value below
alpha (default 0.01). Thresholds come from the config file when one exists, otherwise
the defaults (timingPct: 5). The bootstrap is seeded from the samples, so the same two
documents always give the same verdict. See docs/statistics.md.
Exit: 0 pass, 1 regression, 2 nothing matched or a load error.
Renders a saved document without re-running anything.
ostia report doc.json --format markdown
ostia report doc.json --format minimal
ostia report doc.json --format speedscope --out-dir viz/
ostia report doc.json --format collapsed | flamegraph.pl > flame.svgFormats: table (default), json, jsonl, markdown, minimal, and, for documents
with CPU evidence, collapsed, mermaid, speedscope, cpuprofile. time, bench,
compare and ci accept only the first five; export a document and use report for the
visualization formats.
Runs the config's workloads, compares them against a named baseline, and exits 1 on a regression.
ostia ci
ostia ci --full # ignore the cache
ostia ci --baseline release
ostia ci --save-baseline # after a pass, make this run the new baseline- Command workloads are cached by their declared
inputs: noinputsfield always reruns;inputs: []means "depends on nothing" and caches; otherwise the run is reused while the matched files' contents are unchanged.suitesworkloads always run. suitesworkloads run with the config'sbenchsection, the same wayostia benchreads it.- Exit 2 if any command workload exits non-zero (not ignored) or produces no samples, or
if the baseline file is missing. A baseline that matches none of the configured
workloads is also an error; one missing only some lists them and carries on
(
onMissingBaselinein the config changes this).
ostia baseline save # measure the configured workloads -> .ostia/baselines/main.json
ostia baseline save my-feature
ostia baseline list
ostia baseline show main --format markdownsave uses the same measurement code path as ci. show accepts report's flags.
import {
time, bench, ab, group, task, sweep, range, run, profile, keep,
compareDocuments, defineConfig, createDocument, loadDocument, saveDocument, renderers,
} from "ostia"
import type { ProfileDocument, MinimalEvent } from "ostia"| Export | Does |
|---|---|
time(opts) |
Subprocess timing, same as ostia time. Returns a ProfileDocument. |
bench(opts) |
Runs suite files, same as ostia bench. |
ab(opts) |
Paired A/B of suite files against a git ref, same as ostia ab. |
group(name, fn, opts?) / task(name, fn, opts?) |
Register in-process tasks; .skip/.only variants. |
sweep(dims, fn) / range(start, end, mult?) |
Parameter sweeps; tasks inherit the point as params. |
run(opts?) |
Runs the tasks registered in the current file, in this process (bun suite.ts). |
profile(fn, opts?) |
In-process CPU capture; origin: "jsc" adds JIT tier data. |
keep(value) |
Pins an intermediate value against dead-code elimination. |
compareDocuments(base, cand, thresholds?) |
Same comparison as ostia compare. |
defineConfig(config) |
Typing helper for ostia.config.ts. |
createDocument / loadDocument / saveDocument |
Build, read (schema v2 only), and write documents. |
renderers |
table, markdown, json, jsonl, minimal, collapsed, mermaid, speedscope, cpuprofile. |
const doc = await time({
commands: ["bun a.ts", { command: "bun b.ts", label: "b", prepare: "rm -rf dist" }],
samples: 20,
cpu: true,
})
group("parse", () => {
sweep({ size: range(100, 10_000) }, ({ size }) => {
const input = buildInput(size) // unmeasured setup, once per point
task("parse", () => parse(input), { isolate: true })
})
})
const result = compareDocuments(await loadDocument("before.json"), doc)
if (result.summary.verdict === "fail") process.exitCode = 1
const { text } = await renderers.markdown.render(doc, {})Full reference, including task options, hooks, and a mitata/hyperfine migration table: docs/library.md.
ostia.config.ts (checked first) or ostia.config.json, in the current directory.
// ostia.config.ts
import { defineConfig } from "ostia"
export default defineConfig({
baseline: "main",
samples: 15, // command workloads; or budgetMs/minSamples
warmup: 3,
thresholds: { timingPct: 5 },
workloads: [
{ label: "cold-start", command: ["bun", "src/cli.ts", "--help"], inputs: ["src/**/*.ts"] },
{ label: "spawn", command: ["bun", "-e", "1"], inputs: [] },
{ label: "build:cold", command: ["bun", "build.ts"], prepare: "rm -rf dist" },
{ label: "suites", suites: ["bench/*.ts"] },
],
bench: { budgetMs: 500, isolate: true, preload: ["bench/setup.ts"] },
})A config with a wrong-typed value, or the old runs field, fails to load with a message
naming the key (error code config-invalid). All fields: docs/config.md.
- docs/cli.md: every command and flag
- docs/config.md: config file reference
- docs/library.md: library API reference
- docs/agent-protocol.md:
--format minimal, exit codes, error codes - docs/statistics.md: sampling, the comparison test, noise floor
- docs/document-schema.md:
ProfileDocument, workload ids, warnings - docs/preload-recipes.md: jsdom, happy-dom and
Bun.plugin()preloads
examples/ has runnable recipes (they use ../../src directly, no install):
compare-two-commands,
find-a-hotspot,
heap-usage,
gate-a-regression,
profile-in-process,
benchmark-a-function.
cd examples/find-a-hotspot && bun run demo
bun run examples # all of them, from the repo root