Skip to main content
Xum’s microbenchmarks work like Go’s testing.B and benchstat. A *.bench.ts file registers mitata benchmarks, make bench runs them, and scripts/perf/benchCompare.ts compares two revisions. Use them to back optimization PRs with numbers. To find out why the running app is slow, see Profiling Xum instead.

Write a benchmark

Put the file next to the code it measures (src/node/services/agentTaskIndex.bench.ts), or under scripts/perf/bench/ for scenarios that span modules. A bench file only registers benchmarks with mitata (bench, group, summary). It never calls run(), and it must not import bun:test or test harnesses: it also runs on Node. src/node/services/agentTaskIndex.bench.ts is the example. It shows four features:
  1. Setup outside the timing. With the generator form, mitata times only the yielded function:
  2. Sizes. .args("workspaces", [1000, 5000]) runs the benchmark once per value. $workspaces in the name becomes the value.
  3. Async code. An async function* can await in setup, and mitata awaits the promise that the yielded function returns.
  4. Heap and GC. .gc("inner") collects garbage before every iteration. The table then shows the GC time and the heap bytes allocated per iteration. Node runs with --expose-gc for this.
Benchmarks use real implementations with synthetic data in temp dirs. Test lanes never run *.bench.ts files, and the app build excludes them.

Run benchmarks

BENCH matches a substring of the file path, or a glob when it contains *, ?, [ or {. Each file runs in a fresh process. Node is the default because production runs on V8: the desktop app runs Electron (Node 24) and xum server runs Node 22. Bun runs on JavaScriptCore, so its numbers differ, sometimes a lot (for example Intl.Segmenter). With Node, the runner bundles the bench with esbuild into a temporary directory under build/bench/ and keeps npm packages external. The runner prints the runtime, CPU and load average before the mitata table, and the load average and JSON path after it. It exits nonzero when a benchmark throws.

JSON output

Each run writes artifacts/bench/<name>-<runtime>-<sha>.json, where <name> is the bench file name plus a short hash of its path. A run where any benchmark throws writes no JSON. Pass JSON=<path> (with a BENCH filter that selects one file) to choose the path.

Compare two revisions

The head is your working tree, uncommitted changes included. To compare two commits instead, run bun scripts/perf/benchCompare.ts --bench <filter> --base <ref> --head <ref>. The script checks out each ref into a temporary detached worktree in the OS temp directory and removes it on exit. It does not touch your checkout, index or stash. Both sides run your working-tree *.bench.ts files, so a new bench can measure a base that does not have it. Everything a bench file imports, helpers included, comes from each side’s own tree: keep bench-only helpers inside the bench file. Both sides also share your node_modules, so the script compares source changes, not dependency upgrades. Each round runs base and head once each as fresh processes, and the first side alternates every round. Host load drift then hits both runs of a round alike. For every benchmark, the script takes the mean time of each process run and reports:
  • the median of those per-round means for base and head,
  • the delta: the geometric mean of the per-round head/base ratios,
  • a 95% confidence interval of that ratio (paired t-interval over the rounds), so drift shared by both runs of a round cancels out,
  • a verdict: slower or faster when the whole interval is above or below zero, otherwise ~ (no significant change).
It also prints the load average at the start and end of the session and writes artifacts/bench/compare-<sha>-<timestamp>.json.
Shared hosts swing in load. A wide interval means noise, not a result: add rounds or rerun on a quieter host. Each row is a 95% interval, so about one row in 20 gets a wrong verdict on identical code, usually a delta of a few percent. Rerun before you claim a change below about 5%.

Report numbers in a PR

Paste the full compare table into the PR description, with the rounds, runtime, base and head SHAs and the load average lines. Never report a single make bench run: one process says nothing about noise.