Covers the harness's safety, refusal, and portability behaviour — the parts a third party hits before they ever produce a number:
doctornaming every unmet prerequisite and its remedy, while changing nothing.- The disposable per-run scratch home: the runner's real application profiles are never read or written.
- Refusing to start when a subject is already running, keyed on bundle identifier.
seed/restoreround-tripping a profile byte-for-byte.runrefusing on an unquiesced machine and measuring nothing.- Per-subject bundle-path overrides and version-drift reporting.
Deliberately excluded. A full measurement sweep is not verifiable from a testing guide: it needs a rebooted machine with nothing else running, for hours. See Not verifiable here. This guide proves the harness is safe to hand a stranger; it does not prove any published number.
Risk being managed. This harness seeds application state and launches real applications. Its predecessor wrote directly into the operator's live TermTree, Collaborator and CodeNomad profiles. Every scenario below exists to prove that is no longer possible.
backend — the harness is a command-line binary with no UI. Every assertion
is made at its CLI boundary: exit code, stdout, and the filesystem state it
did or did not change. There is no browser surface. The one native concern
(launching a real .app with an isolated HOME) is covered under Not
verifiable here, because it requires quitting applications the operator is
using.
- Repository: this repo (
documentnode/termtree). The harness is a standalone[workspace]crate atbenchmark/with no path dependencies — it needs no other repo checked out. - Dependencies: Rust stable, plus the nightly toolchain for the format check.
The macOS tools it shells out to —
lsappinfo,footprint,vm_stat,sysctl,pmset,notifyutil,open— ship with a stock macOS and none needsudo. - macOS only.
open --envmust be supported;doctorprobes for it.
cd benchmark
cargo build --release
B=./target/release/resource-benchmarkYou no longer need to export a scratch HOME. Earlier versions of this
harness read your real $HOME and required you to remember to override it.
It now creates a fresh disposable home under the OS temp directory for every
invocation, and refuses an explicit --home that points at your real home
directory. Scenario 3 proves both halves.
Run before the manual pass. They complement it; they do not replace it.
cd benchmark
cargo test # expect: all pass, exit 0
cargo clippy --all-targets -- -D warnings # expect: exit 0
cargo +nightly fmt -- --check # expect: no output, exit 0Nightly is required for the format check — rustfmt.toml sets
unstable_features = true.
- Run
$B doctor. - Read every line. For each unmet prerequisite, confirm the output says what
to do about it, not only what is wrong:
- a not-installed subject names the expected path and the
--bundle-path <id>=<path>override, e.g.subject not installed: Collaborator (/Applications/Collaborator.app) -- install it there, or point the harness at an existing install with--bundle-path collaborator=<path to the .app>`` - a failing quiesce gate is followed by one indented
to clear <signal>:line per failing signal, each naming a concrete command or action.
- a not-installed subject names the expected path and the
- Confirm the exit code is
1when any problem was printed:$B doctor; echo "exit=$?". With no problems it printsdoctor: all checks passed.and exits0. - Confirm it changed nothing.
doctorcreates and deletes a scratch-home probe directory. Afterwards,ls $TMPDIR | grep resource-benchmark-doctor-probemust return nothing.
- Launch TermTree normally (or any subject in the registry).
- Run
$B doctor. - Expect a line naming the app, its bundle identifier, and its pid:
TermTree (com.termtree.desktop) is already running (pid NNNNN) -- quit it before running the benchmark. - Confirm the pid matches:
lsappinfo list | grep -A5 TermTree | grep 'pid ='.
Bundle identifier rather than app name is the point. Two differently named
bundles can declare the same identifier; the second launch is then handed off
to the running instance by the single-instance plugin and exits within
seconds, having measured nothing. open -n does not bypass that.
This is the guide's most important scenario.
-
Record your real profile's state before anything:
REAL="$HOME/Library/Application Support/DocumentNode/TermTree" ls "$REAL" | grep -c before-resource-benchmark # expect: 0
-
Seed with no
--home:$B seed --subject termtree --sessions 3 -
Expect exit
0and a message naming the scratch home it chose, e.g.seeded termtree with 3 sessions via production state.json pre-write (scratch home: /var/folders/.../resource-benchmark-home-<pid>-<nanos>). -
Confirm the seeded fixture landed there, not in your profile:
S=<the scratch home from step 3> python3 -c "import json;print(json.load(open('$S/Library/Application Support/DocumentNode/TermTree/state.json'))['tree']['label'])"
Expect
resource-benchmark-root. -
Confirm your real profile was not written:
ls "$REAL" | grep -c before-resource-benchmark # expect: still 0
The seeder always writes a
state.json.before-resource-benchmark.jsonbackup before touching a state file, so the absence of that file is proof it never wrote there.Do not use the real
state.json's checksum for this check. If TermTree is running it rewrites its own state continuously, so the hash changes for reasons that have nothing to do with the harness. Check for the backup file, and check the tree's root label is still yours. -
Now point
--homeat your real home. It must refuse:$B seed --subject termtree --sessions 3 --home "$HOME"; echo "exit=$?"
Expect exit
1andRefusing to use /Users/<you> as the scratch home: it is the real home directory (...) or contains it, so seeding would overwrite the runner's own application profiles. -
Same via the environment variable:
RESOURCE_BENCHMARK_HOME="$HOME" $B seed --subject termtree --sessions 3; echo "exit=$?"
Expect exit
1. -
Confirm it is not over-blocking — an explicit scratch directory still works:
S=$(mktemp -d); $B seed --subject termtree --sessions 2 --home "$S"; echo "exit=$?" find "$S" -name state.json
Expect exit
0and onestate.jsonunder$S.
- Build a scratch home with a state file standing in for a real one:
S=$(mktemp -d); D="$S/Library/Application Support/DocumentNode/TermTree" mkdir -p "$D" printf '{"tree":{"id":"original-root","label":"my real work","children":[]},"themeKey":"dark"}' > "$D/state.json" ORIG=$(shasum -a 256 "$D/state.json" | cut -d' ' -f1)
$B seed --subject termtree --sessions 4 --home "$S"- Confirm both files now exist:
ls "$D"showsstate.jsonandstate.json.before-resource-benchmark.json. - Confirm the live file is the fixture: its
tree.labelisresource-benchmark-root. $B restore --subject termtree --home "$S"— expectrestored termtree.- Confirm byte-identical restoration:
[ "$ORIG" = "$(shasum -a 256 "$D/state.json" | cut -d' ' -f1)" ] && echo PASS || echo FAIL
- Confirm the backup was consumed:
ls "$D" | grep -c before-resource-benchmarkreturns0. rm -rf "$S".
Covered by cargo test's refuses_a_directory_outside_the_scratch_root and
refuses_a_termtreedev_directory_even_under_the_scratch_root. There is no
safe manual equivalent: the manual version would require pointing the seeder
at a real profile, which Scenario 3 step 6 now refuses outright.
- On an ordinary working machine (browser open, apps running), run
$B run. - Expect exit
1andrefused to start: quiesce gate failed, refusing to start: <signals>. - Confirm nothing was produced:
ls results/is unchanged. - Confirm no subject was launched or quit — any app that was running
before is still running:
lsappinfo list | grep -c com.termtree.desktopis unchanged.
The preflight order is quiesce gate → open --env support → already-running
check, and every one of them returns before any seeding, launching, or
teardown. Teardown issues a graceful quit, so it must never be reachable for a
process the harness did not itself launch.
- Point a subject at an app that exists but is the wrong one, to exercise all
three behaviours at once:
$B doctor --bundle-path 'codenomad-electron=/Applications/<some installed>.app'
- Expect the "subject not installed" line for that subject to disappear — the override was honoured.
- Expect a version-drift line naming both versions, e.g.
CodeNomad (Electron): version drift, expected 0.18.0 found 2.2.0. - Expect a
doctor note:line stating that subject's seeder has not been verified against a real install and that its N-session/sustained-use samples will reportinvalidReason=seed-format-unverified.
Note the notes are only reachable for an installed subject; a not-installed subject short-circuits before them, which is why this scenario uses an override to make one reachable.
Collaborator, CodeNomad and diri have seeders whose formats have never been checked against a real install. Confirm the harness says so rather than producing a number that looks valid:
grep -rn 'seed_format_verified' src/subject.rs— onlytermtreeistrue.- Their N-session and sustained-use samples carry
invalidReason: "seed-format-unverified", asserted bycargo test. $B doctoremits the note from Scenario 7 for each installed one.
- A real measurement sweep. Requires a rebooted machine at nominal memory
and thermal pressure, no swap in use, on AC power, with exclusive use for
hours.
$B doctormust pass first. On a 16 GB host, never run two subjects concurrently. - Launch isolation against a real application — that
open -n -F --env HOME=<scratch> -a <bundle>genuinely redirects an app's data directory. Verifying it means launching a subject, which requires that subject to be quit first. On a working machine the operator's own TermTree and MarkNode are typically running, and launching either triggers the single-instance handoff described in Scenario 2. Precondition to verify: a quiesced machine with the subject quit — i.e. the same sitting as the sweep. The mechanism was confirmed manually against MarkNode on 2026-08-25:lsappinfoattribution still resolved under a scratchHOME, the app used the scratch data directory, and the real profile's mtime was unchanged. - Cold-start log-mark self-validation (
termtree-log-marks-unrecognized) andapp-data-dir-not-created. Both fire only after a real TermTree launch, so they share the precondition above. Their pure logic is unit tested. - The three unverified seeders against real installs. Collaborator, CodeNomad (Electron and Tauri) and diri must be installed at the pinned versions first.
- The spawn-and-wait orchestration path
(
build_envelope/measure_one/measure_cold_start/teardown) has never executed end to end. Treat the first sweep as its verification and budget it as debugging.
- Scratch homes are not removed automatically. They accumulate under the
OS temp directory as
resource-benchmark-home-<pid>-<nanos>until the OS reclaims them. To clear them now:rm -rf "$TMPDIR"/resource-benchmark-home-* - Remove any scratch directory you created explicitly with
--home. - If a seeding scenario was interrupted between
seedandrestore, run$B restore --subject <id> --home <that home>.$B doctor --home <that home>reports a leftover backup; it only performs that check for an explicitly supplied home, since a fresh default home can never have one. - Never run
restoreagainst your real home —--home "$HOME"is refused.
- Capture exit codes directly (
cmd; echo "exit=$?"), never through a pipe —cmd | tailreports the exit code oftail, so a failure reads as0. - Do not assert on the real
state.json's checksum while TermTree is running; it rewrites its own state. Assert on the absence ofstate.json.before-resource-benchmark.jsonand on the tree's root label. - Do not launch a subject to test isolation while the operator is working. Check
lsappinfo list | grep bundleID=first; if the subject or anything sharing its bundle identifier is running, record the scenario as blocked rather than quitting the operator's application. $B runis safe to invoke on an unquiesced machine — it refuses in preflight before touching anything — but do not add--allow-*flags to force past a refusal during verification.