We've just launched. The site and the service are still being polished — we apologize in advance for any inaccuracies or rough edges. Spotted a mistake or something that doesn't work as expected? Tell us and we'll fix it quickly.

EN
Sign in Free estimate
Backtests · model training · parameter sweeps

Faster repeated computation, the same result

Send the hot computational core of a backtest, simulator or model. Rust crates use our self-service dashboard; C, C++ and Python start with an individual engineering assessment. Published builds let you check the results yourself. Your own gain and acceptance criteria are agreed for your workload and platform.

1 h → 11 minindicator parameter sweep at the measured 5.7×, level 2 vs free level 0
1 h → 4 minspiking-network training sweep at the measured 14.2×, level 2 vs free level 0
0 bitsof difference from the original code
How it works free initial assessment
  1. Send the hot code

    Only the computational core, without business logic. Not ready to share code? Start with a readings bundle from our free harness.

  2. Agree on the evidence

    For Rust, build tuning and automatic optimization are measured before purchase. Individual level-2 work has a quoted scope, a measurable minimum and an acceptance procedure.

  3. Choose the paid work

    Level 1 is not charged below a 10% measured speedup. Level 2 pricing, any advance payment and acceptance are agreed in a separate order. Output equivalence is checked on the agreed platform.

Measured cases you can download and run are just below ↓
Already measured

Standard benchmarks, output identical to the original down to the bit

Public code and widely used datasets: the Silesia compression corpus, Xiph video clips, the ETT time-series forecasting sets, technical-analysis indicators. Every measurement is a real sequential run of the whole set against the unmodified original, and all outputs are compared byte by byte. The correctness proofs for the codecs passed an independent review.

All three levels · binaries published

ADX, ATR and PSAR indicators from Python ta

A Rust port of the ta 0.11 library that matches Python bit for bit. A parameter sweep over one million bars: ATR, ADX and ±DI for five windows and Parabolic SAR in four configurations. Every level is built, measured on the rig and published; the output of all levels is identical to the byte, including on inputs where the program panics.

5.7×level 2 vs free level 0, full-run time
18×vs Numba on the computation
< 48 hof engineer's work on level 2

The same algorithm in Numba (also bit for bit): 335 ms of computation and 201 ms of pandas.read_csv. Level 2: 18 ms of computation and 20 ms of CSV parsing, 81 ms for the whole program. Python ta: 20.5 s per 100k bars.

Baseline, cycles per run (M1, median) 1,376.9 M 503 ms
Level 0LTO, codegen-units, target-cpu · free 1,248.6 M463 ms · the reference
Level 1automatic proven optimization 791.4 M342 ms · 1.58× cycles · 1.35× time
Level 2an engineer: a new computation scheme 143.3 M81 ms · 8.7× cycles · 5.7× time
Measured on the rig: Zen 3, vPMU cycles, 9 interleaved rounds; time is the whole program, including reading the data. Levels 1 and 2 are compared with the free level 0. Exactly the published files were measured. In this case automatic optimization alone gives 1.58× in cycles; the largest gain comes from level 2.
Download · MIT · Linux x86-64 full package1 MB: programs for every level, baseline source, data generator, measurements baseline level 0 level 1 level 2 how to verify SHA256SUMS
All three levels · binaries published

Training sweep of a spiking neural network (a BindsNET model)

A reservoir of 128 leaky integrate-and-fire neurons with recurrent connections; the input weights are trained online by STDP (BindsNET PostPre). The sweep trains 24 networks over threshold, membrane time constant and learning rate on a synthetic market-like stream of 5,000 steps. The Rust program prints exactly what the BindsNET script prints, to the bit, on the main input and on 54 generated datasets.

15.8×level 2 vs free level 0, cycles (14.2× time)
395×vs the same sweep in BindsNET, one core
0 bitsof difference from BindsNET

BindsNET 0.3.3 on PyTorch (CPU, one thread, the same core): 24.0 s for the sweep, mostly per-step Python and dispatch overhead. Level 2: 60.9 ms on that core. Build tuning alone is 6% slower than the plain build here; automatic optimization gives 3.89× over level 0 in cycles, and level 2 gives 15.8×.

Baseline, cycles per run (M1, median) 3,525.0 M 805 ms
Level 0LTO, codegen-units, target-cpu · free 3,742.6 M865 ms · the reference
Level 1automatic proven optimization 962.4 M222 ms · 3.89× cycles · 3.90× time
Level 2an engineer: a new computation scheme 237.0 M60.7 ms · 15.8× cycles · 14.2× time
Measured on the rig: Zen 3, vPMU cycles in an isolated VM, 9 interleaved rounds; the CCX was reserved for the measurement (2 of its 16 logical CPUs were outside our control). Levels 1 and 2 are compared with the free level 0. Exactly the published files were measured.
Download · MIT · Linux x86-64 full package0.8 MB: programs for every level, baseline source, BindsNET script, data generator, measurements baseline level 0 level 1 level 2 how to verify SHA256SUMS
16–20×vs a competent native port
≈12,000×vs vn.py’s own optimizer
Strategy optimization · vn.py · binaries published

vn.py parameter optimization: the same top 10, an order of magnitude faster

The brute-force optimizer of vn.py over 600 parameter sets of its stock moving-average strategy. Our program prints the same ten best sets with the same Sharpe ratios, returns, drawdowns and trade counts, bit for bit on AVX2 machines. It is 16–20× faster than a competent native port that computes each moving average once, and about 12,000× faster than vn.py itself, most of which comes from leaving Python. Measured on AMD Zen 3 and Zen 4.

Download · MIT · Linux x86-64 full package60 KB: our program, both native ports, the vn.py driver, data generator, measurements on two platforms what was measured
2.3–7.4×seven indicators, C
2.0–6.0×from Python, unchanged
C library · drop-in · binaries published

TA-Lib technical indicators: a faster drop-in library

A replacement for the TA-Lib 0.8.1 C library with the same ABI: C programs and the TA-Lib Python wrapper use it without changes, and every output matches the official library bit for bit (4.59 million calls checked). In a parameter sweep over a million bars, seven indicators (CCI, SAR, DX, ADX, EMA, STOCH, RSI) are 2.3–7.4× faster than the fastest of the official and the -march=native builds, measured on AMD Zen 3 and Zen 4. Through the unchanged Python wrapper the same sweep is 2.0–6.0× faster.

Download · BSD-3 · Linux x86-64 full package3.2 MB: the library, benchmark drivers, the bit-exactness checker, measurements on two platforms what was measured
5.6–5.9×vs the original Rust
2.3–3.6×vs PyTorch + Inductor on all cores

Time-series forecasting in Rust

DLinear with the officially trained weights on the four ETT sets, every window of the test split, against the original Rust implementation. On all cores it is 2.3–3.6× faster than PyTorch with the Inductor compiler. PatchTST: 2.2–2.5×.

1.29–1.36×

rav1e 0.8: AV1 encoding

Five Xiph clips picked at random before measuring: Elephants Dream, Park Joy, Johnny, Bus, Stefan. The full encoding cycle at speed 6 and 10. Every clip is faster, and the bitstreams match byte for byte.

Download · BSD-2 · Linux x86-64 full package baseline TAKT how to verify SHA256SUMS
3.4×vs the original Rust
4.0×vs libbz2 in C

bzip2 decompression

The pure-Rust bzip2-rs crate, all 12 files of the Silesia corpus, 212 MB. Every file is faster, by 2.8× to 4.2×. On 400,000 test streams, 360,000 of them corrupted, both the data and the error messages match.

Download · MIT · Linux x86-64 full package baseline TAKT how to verify SHA256SUMS

Rig: AMD Threadripper PRO 5975WX (Zen 3, AVX2). Ratio of medians of real runs over the whole set: seven rounds for codecs, five repeats for models. Single-threaded, PyTorch is still faster on these models: it reorders additions, while we keep the output bit-identical to the original code. These are results on specific code, not a promise for yours: for your crate we name numbers after measuring.

Pricing

Pay for a speedup you have already seen

The level 1 price depends on the size of the hot code and the speedup achieved. You see the result before paying, so you pay only for what you get.

Level 0
$0
  • Build profiles: LTO, codegen-units, target-cpu, PGO
  • A flag recipe you apply yourself
  • Hot code up to 2k lines
  • Ready in minutes

The number is measured on the rig, not forecast.

Level 2 · engineer
quoted 30% upfront
  • First and foremost a change of computational paradigm, then data layout, SIMD and everything else
  • From a couple of hours to a couple of days
  • The rest after acceptance against the metrics

Guaranteed minimum not reached: the money goes back to your balance according to the table in the contract.

Level 1 prices

Speedup over level 0up to 1k lines1k–3k lines3k–10k linesover 10k lines
10–25%$290$490$890on request
25–100%$790$1,190$2,390on request
100–200%$1,190$1,790$3,590on request
over 200%$1,990$2,990$5,990on request

Speedup is how many times faster the new build is than level 0, minus one: 100% means twice as fast, 200% three times; in cycles that is −20% at 25%, −50% at 100% and −67% at 200%. Free level 0 covers hot code up to 2k lines; above that it is included in level 1. An additional platform: +30%. Lines are counted with llvm-cov: only code executed by the passport benchmark (details in the question “How are lines counted?”).

What it saves you

Enter how long one run takes now and how often you run it. Take the speedup from your free estimate; 5.7× is the measured ta case.

Waiting time saved per month—
Compute saved per month—

Plain arithmetic on your inputs: saved time = runs × duration × (1 − 1/speedup). The speedup for your code is measured before you pay.

CI subscription: the speedup survives new releases

PlanPer month, billed yearlyWhat’s included
Crate$149Hot code up to 2k lines, up to 4 releases a month, one platform
Team$399Up to 10k lines, up to 20 releases a month, two platforms
Enterprisefrom $1,500Several crates, SLA, acceptance on your hardware

The subscription runs for a year; the annual fee ($1,788, $4,788 or from $18,000) is debited from your balance at the start. Every new release is optimized again and checked for equivalence. If a release could not be brought to the contract level, you get the share of the month such releases make up, as you choose: that share of the monthly fee on your balance, or the subscription extended by that share of the month. One of four releases in a month failed: a quarter of the monthly fee, or a quarter of a month.

All payments go through the balance in the dashboard. Top it up by card through Stripe (cards, Apple Pay, Google Pay and local payment methods), by bank transfer or with a promo code; each service is debited when you order it. Unused paid funds are paid back on request; bonus funds from promo codes and bonuses pay for services.

Who it is for

Where the speedup is largest

Today we work on x86-64 (AMD Zen 3, AVX2). The biggest wins come from dense iterative computations with feedback, where each step depends on the previous one: indicator recurrences, strategy simulation along a history, training of spiking and other recurrent models, and sweeps over many parameter sets. Code that waits on memory for long stretches or parses data bit by bit also gains a lot.

Backtesting and strategy simulation

Trading engines, market data replay, parameter search and evolution. A private client case: 2.5–10.5× automatically, up to 33.7× with an engineer. A public case: the ta indicators, 5.7× in full-run time over the free level 0.

Model training and parameter search

Models with strong feedback, such as spiking networks (in the style of BindsNET) and other recurrent models, trained and swept over many parameter sets. At level 2 we port a Python model to Rust bit for bit. Public case: an STDP training sweep of a BindsNET model, 15.8× over the free level 0 and 395× over BindsNET itself, with identical output.

Archives and data pipelines

Archive decompression, Wikipedia and OpenStreetMap dumps, scientific and genomic data, log storage. bzip2: 3.4× vs the original Rust and 4.0× vs C.

Video and AV1

Video platforms and cloud transcoding on rav1e. 1.29–1.36× on full encoding: on a cluster that is about a quarter of the servers.

ML inference on CPU

Time-series models in Rust without Python: load, demand and energy forecasting. 5.6–5.9× vs the original Rust, faster than PyTorch on all cores.

Class of codeMeasured speedupWhere measured
Stateful simulations, searching candidates along one history2.5–34×trading engine
Spiking networks: STDP training, parameter sweeps3.9–15.8×BindsNET: LIF + STDP
Technical-analysis indicators, parameter sweeps1.6–8.7×ta: ADX, ATR, PSAR
Strategy optimizers: parameter grid search16–20×vn.py: DoubleMaStrategy
Indicator libraries in C, called from C or Python2.0–7.4×TA-Lib 0.8.1
Decoders with table walks and bit-level parsing2.5–4×bzip2, Silesia
Inference of small models on all cores2.2–5.9×DLinear, PatchTST, ETT
Encoders with mode search1.3–1.4×rav1e, Xiph
Crypto primitives, hand-written SIMD codecs, formats with a strong C librarynoneSHA-256, ChaCha20, x264, LZ4

For ta and the spiking network: against free level 0, in cycles. Other rows: against the original program, or the reference C library where the case says so.

How it works

From the first measurement to a signed build

Measure it yourself

The free harness records cycles, instructions, cache and branch misses and wall time. You can send us the readings bundle without source code, and we give a preliminary range.

Send the hot part

A real measurement, as opposed to an estimate, needs the source. We help you extract just the computational core without business logic and take its baseline on the reference rig.

We fix the passport

Platform, workload, tool versions and protocol are recorded before work starts. The result is counted from these numbers.

See three numbers

Level 0 and level 1 are measured on the rig in advance. Level 2 comes as a range, a timeline and a guaranteed minimum.

Receive the build

After payment: a signed library, a header, an SBOM, a report on measurements and the equivalence check.

For engineers

The evidence

What we sign, how we measure and what you can check yourself: the measurement passport, the open harness, the isolation of your code and a comparison with the usual options.

Guarantee

We promise specific numbers on specific hardware

Before work starts we fix the measurement passport: the CPU model and frequency, the workload with hashes of the input data, the rustc and tool versions, the measurement protocol. The contract speaks only about these metrics.

M1Cycles on a real CPUThe PMU counter on the reference rig at a fixed frequency, baseline and new version interleaved. This is the promise.
M2Deterministic cycles and instructionsSimulation of a pinned version. The same number for us and for you, for cross-checks and CI.
M3Time p50 / p99For latency-sensitive work the promise can be given on p99.
M4Peak memoryConstraint: no worse than the baseline by more than an agreed percentage.
M5CorrectnessDifferential tests and fuzzing, including on hidden inputs the optimizer has never seen.
Wording in the contract

“On platform P, under the measurement passport protocol, the median of M1 will drop by at least X% relative to the baseline, provided M4 and M5 hold.”

Why cycles on a real CPU. A simulator gives a repeatable number, but not always the right one: it does not see the frequency drop under wide SIMD, the prefetcher or multithreading. So the contract uses real hardware, and the simulation serves for cross-checks.

Hidden inputs. Part of your data is held back until acceptance, so the speedup cannot be fitted to the benchmark. You send us this data encrypted and give the key once the binary is ready.

Harness

Measure it yourself

A free open-source command-line tool (MIT) takes the same kind of readings as our rig: user-mode cycles and instructions from the hardware counters. Send us the readings without code, and accept the work yourself.

run

Cycles, instructions, cache and branch misses, wall time and memory. Median, spread, 95% confidence interval and a quality grade.

:::

Several builds in interleaved rounds, so drift in temperature and load affects them all alike.

compare

Before and after: the speed-up factor with its interval, as defined in our Terms, and a check that the outputs match. The exit code can gate a CI job.

show

Review the bundle before you send it: no code, no output, no file contents, only hashes and machine facts. --redact hashes paths and arguments too.

$ TAKT_INPUTS=data taskset -c 44 takt-harness run --input data \
    -- bin/ta-bench-base ::: bin/ta-bench-level2
machine   AMD Ryzen Threadripper PRO 5975WX 32-Cores · x86-64-v3
checks    pinned to CPU 44 · governor performance · SMT on
protocol  9 rounds after 1 warm-up · CI: percentile bootstrap

metric          baseline       new   factor   95% CI of factor
cycles:u         1.370 G   141.3 M   9.695×   9.582 – 9.927
instructions:u   3.156 G   391.7 M   8.057×   8.057 – 8.057
wall time      484.7 ms  66.90 ms   7.245×   6.983 – 7.346
stdout    match · sha256 ada682c61e5e… · 973 B
quality   good · 9 runs, typical deviation 0.06% / 0.73% (MAD)

verdict   outputs identical · cycles:u speed-up factor 9.695×
bundle    readings.json  # can be sent without source code

Shortened real output on our ta-indicators showcase (1 million bars): the baseline build against the level 2 build. The readings bundle holds the machine facts, hashes of the programs, inputs and output, and the metrics with their spread and a quality grade.

Source code and a ready Linux binary: github.com/rdmitry0911/takt-harness (MIT).

Security

Your code runs only inside an isolated perimeter with no network

Isolation

Every build is a disposable microVM with no network access. A benchmark node serves one client at a time.

Encryption and deletion

A separate key per project. After 30 days, or at the press of a button, the key is destroyed together with the data.

Your code stays yours

Code and data are not shared with third parties, not used in work for other clients, and deleted on schedule. General mathematical methods we discover while speeding up your algorithm extend our library without your code, data or names. These are clauses of the contract and of the NDA signed before upload.

Signed package

SHA-256 of every file, signed with our published key; the SBOM of components; the measurement report and the toolchain the build was made with.

Verifiable cleanliness

The build makes no network calls, starts no other programs and writes no files. A script in the package shows it on your own inputs: no network at all, plus strace.

Only the hot module

You can send just the computational core without business logic. The documentation explains how to extract it.

We deliver an optimized binary and evidence of its measurements and equivalence; the method remains our know-how. Source escrow is available on written request before payment through an order-specific escrow box under the escrow rider. Its platform and provisioning/handover procedure must qualify and be recorded for the order. Shadow mode needs replayable isolated state; you may test the build’s security yourself or with contractors.

How delivery works and what you can check yourself

Comparison

How this differs from the usual options

An AI agent on your ownCI benchmarking servicesA consultantTAKT
Speeds up codeyesno, catches regressionsyesyes
Result known before paymentno—noyes, levels 0–1
Guarantee in the contractnonousually notyes, on M1
Protection against fitting to the benchmarkup to you—depends on the personhidden inputs
Proof of correctness for every changeno, only your tests—rarelyyes, with independent review
Own optimization methodsthe model’s general knowledge—one person’s experiencea library of proven methods that grows with every project
An honest refusal when there is nothing to speed upan agent always changes something—not alwaysyes, before payment
Reference hardwareyoursyesyoursyes
Pricing modeltokensper userper hourper result
FAQ

What people usually ask

Why do you deliver a build and not the source?

The optimization method is our know-how, so we deliver a static or dynamic library with a C header and a thin Rust wrapper. Hooking it up takes one line in Cargo.toml. Correctness is confirmed by equivalence tests, origin by the signature and provenance.

What happens when I change my code?

Rust integration and switchback are agreed for the delivered library. If the optimized core changes, send a new version for another measurement; CI work requires a separate accepted order. For business continuity, source escrow can be requested before payment through a qualified escrow box under the rider.

Do I have to send you my strategy?

Not for a preliminary range: the free harness produces a readings bundle without source code. For the build we need the code that runs in the hot loop, and only that, usually the simulator, the indicators and the portfolio state. It is covered by the NDA, is not used in work for other clients and is deleted on schedule. Part of your data can stay with you until acceptance: hidden inputs are sent encrypted, and you give the key once the build is ready.

Why not just add more cores?

Do both. Independent runs of a sweep spread across cores well, and our build makes each of those runs faster on every core you add. More cores do not help a single simulation along one history, where each step depends on the previous one, and that is exactly the code we speed up. Fewer core-hours for the same sweep also means a smaller cloud bill.

What if the promised speedup does not materialize?

On level 1 you pay for a result that is already measured, so this cannot happen. On level 2 we name a guaranteed minimum; if it is not reached, the money goes back to your balance according to the table in the contract, and paid funds on the balance are paid back on request.

How are lines counted?

We build your code with -C instrument-coverage in the same profile and for the same platform as in the passport, and run the passport benchmark on the open inputs. We count the Rust lines that executed at least once according to llvm-cov. Not counted: blank lines and comments, code that never ran, tests and benchmark scaffolding, third-party dependencies outside the scope of work, and assembly; a generic function counts once. The number goes into the passport before payment. Data does not count, only your code. For reference: decompressing the Silesia corpus with bzip2-rs executes 598 lines of its code, and encoding five Xiph clips with rav1e executes 12,809 lines out of 55,419.

Which platforms are supported?

The public demos target Linux x86-64 with AVX2/FMA. TA-Lib and vn.py measurements include AMD Zen 3 and Zen 4. For a paid build, CPU features, OS, ABI and acceptance hardware are fixed for the order; Intel, ARM and other environments need their own assessment.

What code speeds up poorly?

Decoders of formats that already have long-optimized C libraries (LZ4, zstd, deflate): we have not beaten them yet. Mature codecs with hand-written SIMD and assembly (x264, x265), cryptographic primitives such as SHA-256 and ChaCha20, processing of random incompressible data, and very short calls where all the time goes to overhead. In such cases the free estimate honestly shows a small gain, and you will not have to pay.

My cycle counts differ from yours. Why?

Cycles depend on the CPU model, frequency, turbo, memory and input data. Contract numbers are taken on the rig under the measurement passport. The harness reads the same counters, so on the same CPU model and settings your numbers should be close to ours. The instruction count depends far less on the machine: compare it first.

What about floating point?

By default, outputs match bit for bit, with a limited exception for NaN payloads produced by floating-point arithmetic where Rust permits different payloads. Explicitly preserved NaN encodings and other outputs remain exact unless you agree otherwise in the passport. The passport can also require exact NaN payloads. Any other agreed deviation is defined in the passport and checked by tests.

Will the build crash on an old server?

Check the CPU requirements of each package. Public showcase binaries may require AVX2/FMA and will not run on an older CPU. A paid build’s supported features and any fallback are agreed in its passport; a fallback is not assumed for every binary.

Will it cut the latency of my live trading?

Only if profiling shows that the computation is the bottleneck. For a live path we measure the whole path from the market event to the order, for example its p99 on your server, not one function: often most of the time goes to the network and the exchange. Our strongest results are batch computations: backtests, training and parameter sweeps.

Do you work with C, C++ and Python?

Yes, as individual orders. For C and C++ we deliver a drop-in replacement library with the same ABI, so your programs, and Python code that calls the library, work unchanged (see the TA-Lib case). For Python code we deliver a native program or module with the same output (see the vn.py case). This is level-2 engineering work, priced individually; self-service in the dashboard currently accepts Rust crates.

Beyond TAKT

Our other projects

Systems software from the same team: unlocking an encrypted ZFS root at boot, reshaping RAIDZ in place, a browser console for Proxmox VE and Thunderbolt hotplug in QEMU. Each card opens a short description; the code is on GitHub. Also included: conditional secret escrow in a TPM virtual machine.

escrow-box

A prototype for conditional secret escrow in a TPM virtual machine. Source deposits can be verified by rebuilding before release.

  • TPM 2.0
  • UKI
  • Clevis

zbm-openwrt-clevis

A measured OpenWrt boot runtime that unlocks an encrypted ZFS root with Clevis and the TPM, then boots the system through ZFSBootMenu.

  • OpenWrt
  • ZFS
  • TPM 2.0
  • Clevis

zfs

A fork of OpenZFS with a prototype of in-place RAIDZ reshaping: change parity between raidz1, raidz2 and raidz3 without exporting the pool.

  • OpenZFS
  • RAIDZ
  • C

qsm-rd

QSM Direct: a WebRTC graphical console for Proxmox VE virtual machines and LXC containers, hardware-encoded where the node has an encoder.

  • Proxmox VE
  • WebRTC
  • LXC

qemu-thunderbolt

A Thunderbolt-flavoured PCIe hotplug layer for QEMU: add and remove devices in guests such as macOS that cannot hotplug them directly.

  • QEMU
  • Thunderbolt
  • PCIe
  • macOS
Get started

Check one workload

Rust crate? Open the dashboard now. For C, C++ or Python, describe one slow run below. Start with its duration, frequency and platform; no code or harness installation is needed at this step.

Company and technical details (optional)
No code is needed at this step. By sending the request you agree to the processing of your contact details.