← back to the showcase

TAKT showcase: vn.py parameter optimization, same top results bit for bit

vn.py (MIT) is an open-source Python trading platform. Its CTA backtester ships with a built-in brute-force parameter optimizer, BacktestingEngine.run_bf_optimization. It runs one backtest per parameter set in a pool of worker processes and returns every parameter set sorted by the target metric.

We ran this optimizer over DoubleMaStrategy, the moving-average crossover strategy that ships with vnpy_ctastrategy, on 200,000 one-minute bars and 600 parameter sets. bin/vnpy-sweep-takt prints the same top-10 table, identical byte for byte, in 19.6 ms instead of 4 minutes 2 seconds (median of 9 rounds). Both used the same 4 CPU cores. The plain native port of the same run (B) takes 6.4 s, so C is 326× faster than translation alone.

What was measured

Four programs print the same ten lines. Each line is a parameter set with its Sharpe ratio, total return, maximum percentage drawdown and trade count, printed with %.17g:

Label Program What it is
A tools/vnpy_sweep.py vn.py's own optimizer: run_bf_optimization(setting, max_workers=4), unchanged. A thin driver calls the public API and prints the top 10.
B bin/vnpy-sweep-port A plain native port of the same run: one backtest per parameter set, with the same bar loop, order handling, daily results and statistics as vn.py, and the official TA-Lib 0.8.1 SMA. 4 threads.
B2 bin/vnpy-sweep-port-shared B plus one obvious restructuring: each moving average is computed once and shared by all parameter sets. A sensitivity check of the B baseline. 4 threads.
C bin/vnpy-sweep-takt The TAKT build. 4 threads.

B shows what translating the tool into a compiled language gives by itself. C over B is the part we optimized.

Workload. - Strategy: DoubleMaStrategy from vnpy_ctastrategy 1.4.1, unchanged. - Engine settings (vn.py's CTA demo values for CSI 300 index futures): rate 0.3/10000, slippage 0.2, contract size 300, price tick 0.2, capital 1,000,000, 1-minute bars. Target: sharpe_ratio. - Grid: fast_window 2..31 step 1 × slow_window 20..96 step 4, 600 parameter sets. - Data: synthetic one-minute OHLCV from a fixed seed (tools/gen_data.py, shape 1, seed 1). It has 200,000 bars, 834 trading days of 240 bars, and trending regimes. All prices are on the 0.2 tick grid. The same bars are stored in vn.py's own sqlite database, which A reads, and in a flat binary file, which B, B2 and C read. - Why 200,000 bars and not 1M: vn.py's engine takes about 8 µs per bar per backtest. At 1M bars, one 600-set optimizer run takes about 21 minutes on 4 cores, so 10 measured rounds would take 3.5 hours. 200,000 bars is about 3.3 years of minute bars of one futures contract. B, B2 and C were also measured on 1M bars (below).

Stand. AMD Ryzen Threadripper PRO 5975WX (Zen 3), Ubuntu 24.04, Linux 7.0.14, performance governor, boost on. All four programs were pinned to the same 4 CPUs with taskset -c 33-36. vn.py uses 4 worker processes, and the native programs use 4 threads. The tool was takt-harness 0.2.1: 9 interleaved rounds after one warm-up, with the median taken. It counts user-mode cycles:u of the process and all its children and measures the wall time of the whole process. For A, that includes the Python start-up, the worker processes and their database reads. The 95% CIs are bootstrap intervals of the factor. Raw readings are in results/.

Results: 200,000 bars, 600 parameter sets

Median of 9 interleaved rounds per program. All four printed the same output in every round (sha256 31e99c2d1713…).

Program Wall time cycles:u Peak RSS (largest process)
A, vn.py optimizer 242.4 s 4,202 G 279.0 MiB
B, plain native port 6.398 s 111.5 G 10.6 MiB
B2, port with shared moving averages 320.4 ms 4.89 G 154.1 MiB
C, TAKT build 19.6 ms 0.222 G 29.6 MiB
Factor Wall time (95% CI) cycles:u (95% CI)
C over A: what a vn.py user sees 12,340× (12,054–12,545) 18,962× (18,520–19,088)
C over B: beyond translation 326× (318–331) 503× (493–506)
B over A: translation alone 37.9× (37.6–38.2) 37.7× (37.5–37.8)
C over B2 16.3× (15.8–17.9) 22.1× (21.2–22.2)
B2 over B 20.0× (18.3–20.2) 22.8× (22.8–23.4)

cycles:u sums the user-mode cycles of all threads and processes. For A it includes the four worker processes. The wall-time factor is what a user waits for.

B, B2 and C on 1,000,000 bars (same grid)

Median of 9 interleaved rounds. vn.py itself was not timed at this size; one optimizer run at 1M bars was used for the identity check (see Equivalence). All three printed the same output in every round (sha256 4e28d2259f2c…).

Program Wall time cycles:u
B, plain native port 32.14 s 559.7 G
B2, port with shared moving averages 1.883 s 29.85 G
C, TAKT build 92.6 ms 1.138 G
Factor Wall time (95% CI) cycles:u (95% CI)
C over B 347× (337–350) 492× (487–498)
C over B2 20.3× (19.8–20.6) 26.2× (26.0–26.6)
B2 over B 17.1× (16.8–17.2) 18.8× (18.7–18.8)

Equivalence

Where the result applies

Second platform: AMD Ryzen 9 7950X3D (Zen 4, AVX-512)

The same binaries, unchanged, on a second machine (Debian 13, numpy 2.5.3 with its AVX-512 code paths active, Python 3.13), CPUs 0-3, user-mode counters, 9 interleaved rounds, 1M bars (results/zen4/):

Pair Wall factor (95% CI) Cycles factor (95% CI)
C over B2 18.33× (16.89–18.62) 23.17× (22.92–23.56)
C over B 359.69× (332.93–371.00) 503.99× (499.53–515.39)
B2 over B 19.62× (19.44–20.31) 21.75× (21.50–22.15)

B, B2 and C print the same output in every round, and the same output as on the AVX2 machine.

Identity with vn.py on that machine (results/zen4/identity.txt): identical on the 30,000-bar gappy series and on the main 200,000-bar case. On a 100,000-bar series (shape 2, seed 7), vn.py's own output differs from its output on the AVX2 machine: the Sharpe ratio of two of the ten rows differs in the last digits (0.61673083389321259 against 0.61673083389321237), because numpy computes the logarithm differently with AVX-512. The parameters, their order, the trade counts, returns and drawdowns are the same. B, B2 and C print the AVX2 result on both machines. Equivalence is therefore stated for a platform: on AVX2 machines the output equals vn.py's bit for bit; on AVX-512 machines the decisions are the same and the last digits of a Sharpe ratio can differ, as they do for vn.py itself between the two machines.

Reproducing

python3.12 -m venv venv && venv/bin/pip install -r tools/requirements.txt
export TZ=Asia/Shanghai          # vn.py's sqlite layer converts times through the system time zone
venv/bin/python tools/gen_data.py data/m1_1_200k 1 1 200000      # bars.bin + vn.py database, ~10 s

# one run of each by hand (the same 10 lines)
taskset -c 33-36 venv/bin/python tools/vnpy_sweep.py data/m1_1_200k 4 2 31 1 20 96 4
taskset -c 33-36 bin/vnpy-sweep-takt data/m1_1_200k/bars.bin 4 2 31 1 20 96 4

# the measurement (A runs for about 4 minutes per round)
TAKT_HARNESS=/path/to/takt-harness tools/run_harness.sh results-local/main data/m1_1_200k 9 1
python3 tools/factors.py results-local/main.json C/A C/B B/A C/B2 B2/B
# B, B2 and C only (last argument 0 = without vn.py), e.g. on 1M bars
venv/bin/python tools/gen_data.py data/L1_1_1m 1 1 1000000
TAKT_HARNESS=/path/to/takt-harness tools/run_harness.sh results-local/1m data/L1_1_1m 9 0

# identity of A, B, B2 and C on another series and grid (here: gappy series, seed 12, 30,000 bars)
venv/bin/python tools/gen_data.py data/i3_12_30k 3 12 30000
tools/identity.sh data/i3_12_30k 2 100 7 2 100 7          # prints "... IDENTICAL ..."

The arguments are <data> <workers> <fast_lo> <fast_hi> <fast_step> <slow_lo> <slow_hi> <slow_step>, the same for all four programs.

Contents

Licence

vn.py, vnpy_ctastrategy and vnpy_sqlite are © 2015-present Xiaoyou Chen and distributed under the MIT licence (LICENSE). They are not included; tools/requirements.txt installs them from PyPI. B and B2 link the official TA-Lib 0.8.1 C library statically (BSD 3-clause, THIRD_PARTY_LICENSES.txt). The TAKT build contains no third-party code. We deliver builds; the source of the TAKT build is not published.

Source text: README.md

Telegram