Performance reports
Every backtest run with a strategy phase ends with a performance: block after the report’s phases: line — a TradeStation-style summary computed from the run’s own fills and per-bar marks:
performance: net profit -14600.00 (-14.60%) | closed equity 85400.00
max drawdown 31700.00 (30.61%) peak 2016-06-24 trough 2018-05-07 | sharpe -0.07 | cagr -3.11%
trades 29 closed | win rate 41.38% (12W/8L/9E gross) | profit factor 0.75
gross profit 43750.00 | gross loss -58200.00 | avg trade -498.28 | avg win 3645.83 | avg loss -7275.00
largest win 9000.00 | largest loss -13700.00 | max consec 4W/1L | avg bars held 4.2
The definitions that matter when reading it:
- Open vs closed equity. The engine tracks two curves per bar. OPEN equity is mark-to-market net liquidation (cash plus open positions at the bar close — the same single source the
simulator:equity line uses); CLOSED equity counts only realized results (capital + realized PnL - commission). The headline net profit, drawdown, Sharpe and CAGR are computed on the open-equity curve; both curves are persisted and charted. - Trades are FIFO round trips paired from fills (one trade per entry-lot/exit-fill overlap; a through-zero flip closes the old side and opens the new). This pairing is validated against the legacy engine’s exported books trade-for-trade (
internal/goldentest/perfreplay.go). - Gross basis. Winners and losers are classified on gross trade P&L; commission and slippage are reported separately (the legacy
profitCurrency/commissionPointssplit). - Continuous futures pair across contract rolls as one economic round trip by default (
Rollscounts boundaries survived);--roll-trades splitgives per-contract legs instead. Thepairing residualpublished in result.json is the (small) roll-seam execution gap between trade-level P&L and the equity curve — visible, never hidden. - Sharpe uses daily (last mark per UTC date) returns with the sample standard deviation and needs at least 20 daily returns;
n/aotherwise.
When the run has closed trades, the block continues with the extended sections: a trade stats triptych — every statistic in All/Long/Short columns, including Sortino, Calmar, Ulcer index, the linear-fit R-squared/Pearson pair, CAR, percent of months/quarters/years profitable, and percent time in market — a per-symbol breakdown (multi- symbol runs), and a Year x Month monthly pnl grid whose grand total reconciles exactly to net profit. Triptych curve statistics are computed on the subset’s CLOSED-trade curve (realized PnL bucketed by exit date), so the All column is deliberately a closed-basis view beside the mark-to-market headline numbers.
With --results-dir DIR, every backtest run also persists artefacts under DIR/<run-id>/: result.json (header, metrics, trade list), equity.csv (ts,open_equity,closed_equity, end-of-day sampled by default; equity-sampling: "bar" for every bar), trades.csv, and report.html — a self-contained interactive page built on TradingView Lightweight Charts (embedded, Apache-2.0; no network access): open/closed equity and a synced drawdown pane under one time scale with a crosshair legend, the monthly-PnL heat grid, the metric triptych, sortable symbol/trade tables, and a light/dark theme toggle. Currency reads as $#,### throughout, green above zero and red below; profit factor colours around 1.0 rather than zero, drawdowns are always red, and costs and unsigned rates (win rate, time in market) stay uncoloured so the colour that is there means something. A multi-run config additionally writes a real portfolio combination under DIR/portfolio/ — joint equity curve (capital committed from portfolio start, forward-filled, conserved after a run ends), joint drawdown and Sharpe, a contribution table, and the correlation matrix of daily returns — plus DIR/index.json, the run index that makes the set self-describing.
algo report DIR re-renders all of it offline from the artefacts alone — identical per-run blocks, recomputed portfolio section — with --runs to subset and --partial to combine an incomplete set. Reporting is on by default and --report off restores the pre-report output byte-for-byte; --risk-free-rate feeds the Sharpe computation. In a config file the global-only report object carries the same knobs plus fills (write the per-run fill blotter beside the artefacts) and equity-sampling. TRADES schema runs and live mode do not record performance in v1.
Benchmarks
Every run is compared against one or more benchmarks you choose. Out of the box there are three:
| Name | What it is |
|---|---|
Long ES | Buy and hold the S&P 500 E-mini, rolled, from the first bar to the last. The primary. |
ES/NQ 50:50 | Half the capital in ES, half in NQ, both rolled. |
Cash (3M T-Bill) | The capital left in cash, earning the published 3-month bill rate, compounded monthly. |
Name your own with report.benchmarks:
Inside the run configuration’s global object:
"report": {
"benchmarks": [
{"name": "Long ES", "hold": {"fut:XCME:ES": 1.0}, "primary": true},
{"name": "60/40", "hold": {"fut:XCME:ES": 0.6, "fut:XCBT:ZN": 0.4}},
{"name": "Cash", "rate": "macro:FRED:DGS3MO"}
]
}Weights are relative — they are renormalised over the sleeves that actually priced — and a negative weight is a short sleeve, so a market-neutral pair is a legal basket. Symbols are ordinary wire symbols, class:venue:ticker; FRED/DGS3MO is FRED’s own notation and will not resolve. Exactly one benchmark is primary: it drives the headline comparison, the chart overlay and the benchmark column in equity.csv.
A benchmark does not have to be something the strategy traded, and usually shouldn’t be. Holding what a strategy traded stops meaning anything as soon as the strategy goes short — “you made less than buying what you were selling” compares nothing. Any named instrument the run never trades is fetched separately, purely for the report; it never reaches the strategy or the simulator, so adding a benchmark cannot change a backtest’s results. It can currently stop one: on marketfeed, a futures sleeve the adapter refuses fails the whole run, and the default set’s futures sleeves are refused whenever the run’s wire is DBN (see Report benchmark findings).
The comparison appears as a block of lines in the stdout report, a table in report.html with a value and a difference column per benchmark, a legs table and transaction ledger for each, a dotted reference line on the equity chart for the primary (click the legend chip to hide it), and the primary’s benchmark column in equity.csv. Turn the whole feature off with report.benchmark: false.
A sleeve that could not be priced is named on the page rather than dropped, because dropping it would silently reweight the basket you asked for.
Under the per-instrument table is a folded trades accordion listing every buy-and-hold round trip — the opening buy to the first roll, each roll to the next, and the last roll to the closing sell — with the dated contract, quantity, entry and exit prices, P&L, slippage, commission and net. Costs follow the same convention as the strategy’s trades: slippage is inside the fill prices, so each trade’s P&L already carries it and the Slippage column is a memo, while commission is charged on top and Net is P&L minus commission alone. The totals row foots to the buy-and-hold figures printed above it, so the comparison can be checked rather than taken on trust. It stays folded for printing unless you open it first; the full list is also in result.json.
The benchmark pays the run’s own commission and slippage on every entry, exit and roll leg, and the report shows the roll count, the leg count and the cost separately. This matters more than it sounds: a held position rolls on every boundary, while a strategy pays a roll only when it happens to have a position on. On an 11-year fut:XCME:ES run that is 46 rolls (92 legs) for buy & hold against four for a strategy that is usually flat.
Because comparing on capital alone is fair only when both sides risk the same, the report also states exposure — what the benchmark holds every day against the strategy’s average and peak simultaneous notional, and how much of the time it was in the market — and adds a net profit at equal drawdown row, which solves for the leverage at which buy-and-hold’s percentage drawdown matches the strategy’s. That row is the like-for-like answer. It is omitted entirely when no leverage reaches the strategy’s drawdown — which is itself worth knowing, because it means the strategy took a risk holding could not take at any size. Note the direction is not always the one you expect: a strategy that is levered but rarely on takes more drawdown than holding, so equalising risk scales the benchmark up.
Two things the benchmark deliberately is not, both stated on the page: cash earns nothing — neither the benchmark’s collateral nor the strategy’s idle cash — so a fully collateralised futures benchmark is understated by whatever the collateral would have yielded, which over a long run is usually far larger than the roll bill; and a corporate-action-adjusted instrument reinvests its dividends, making that sleeve a total-return figure while the strategy receives them as cash. Continuous futures are held as a constant contract count, sized and marked in the raw contract frame, so the figure does not depend on the back-adjustment method.
A benchmark and the strategy roll on the same schedule. A futures sleeve the run does not trade is fetched as marketfeed’s :cont composite, so it rolls on marketfeed’s served roll schedule: ES and NQ 12 calendar days before expiry. A continuous strategy follows the same served schedule (see Continuous run conventions), so a fut:XCME:ES strategy and its Long ES benchmark roll together. A sleeve the run itself trades is priced from the run’s own marks and rolls with the strategy.
The BENCH line
The last line of every run is machine-readable: space-separated key=value pairs with the same numbers as the human report (records, bars, orders, fills, final equity, throughput). It exists so scripts and benchmark harnesses can parse results without scraping the pretty report; scripts/bench.sh is built on it. You can ignore it entirely.