Measure a Rust program with interleaved runs and output checks
An optimization experiment needs two answers: did the program get faster, and did it keep the
required output? takt-harness is an open-source Linux measurement tool. It records wall time,
user-mode hardware counters, exit status and a hash of stdout, then compares repeated runs. It
does not optimize code. This walkthrough needs no account and sends no results anywhere.
Make a small deterministic Rust workload
Save this as bench.rs. Wrapping arithmetic makes overflow part of the defined workload, and the
printed value lets the tool compare the builds' output. black_box limits compile-time
elimination; this loop does not represent every production workload.
use std::hint::black_box;
fn main() {
let mut x = black_box(0x1234_5678_9abc_def0_u64);
for i in 0..black_box(20_000_000_u64) {
x = x.rotate_left(7).wrapping_add(i).wrapping_mul(0x9e37_79b9);
}
println!("{x:016x}");
}
Build two variants with the same Rust toolchain:
rustc -C opt-level=3 -o base bench.rs
rustc -C opt-level=3 -C target-cpu=native -o tuned bench.rs
target-cpu=native changes the available instructions. It can help, do nothing, or regress;
measure rather than assuming a gain. This compares build settings, not TAKT's optimizer.
Install and run the measurement tool
Download the Linux x86-64 binary and checksum file from the
takt-harness release.
The source and README explain the record format and
counter setup. Run takt-harness --help after putting the binary on your PATH.
Choose one CPU from the mask printed by taskset -pc $$; replace CPU below with its number.
Keep the same CPU and settings for both builds.
taskset -c CPU takt-harness run --reps 9 --warmup 1 --out readings.json \
--label base --label tuned -- ./base ::: ./tuned
takt-harness compare readings.json
takt-harness show readings.json
Interleaving the builds reduces the effect of gradual changes in temperature and host load. The bundle stores every measured round and describes the CPU and protocol. It reports medians, spread and bootstrap intervals; a single favorable run is not a reliable comparison.
If your VM or container does not expose counters, add --no-counters and compare wall time.
There is no need to change kernel settings for this first experiment. Wall time includes startup
and waiting, whereas user-mode cycles cover CPU execution; neither explains every bottleneck.
Check more than a stdout hash
Equal stdout hashes check what the benchmark actually printed. They do not establish equivalence of unprinted arrays, files, internal state or every possible input. Define the observable result and add direct checks over it, including held-back inputs. Keep the compiler, floating-point settings, libraries and platform in the measurement record when their behavior matters.
Try a published example
The TA-Lib demo runs the original, compiler-tuned and TAKT builds on the same generated series. Its full package includes detailed measurements and bitwise checks. It illustrates the measurement process; it does not predict a gain for another program. TAKT delivers optimized binaries and keeps its optimization methods private.
Disclosure: this walkthrough was drafted by an AI assistant for TAKT. The Rust example and measurement commands were run on Linux; TAKT welcomes corrections through [email protected].