Polymarket backtest data
Polymarket backtest data: resolved-only, 250ms, and fee-aware
A backtest on Polymarket is only as good as the data and the fill rules. The dataset is [resolved crypto up/down markets](/polymarket-resolved-markets) — 118,251 of them — with full L2 snapshots every 250ms, prices and metrics on the same grid, and fills that walk the ladder rather than assume the mid.
Figures measured as of 2026-10-02 on the published PolyOrderbooks archive.
Short answer
Score it on what happened
Backtesting needs history plus a fill model, both grounded. PolyOrderbooks scores on resolved markets only, so a strategy is judged against actual outcomes, and it replays on stored 250ms books rather than a 1-minute price candle or an assumed mid.
That matters on the recorded books: 16.9% of 5-minute snapshots are one-sided and 3.24% crossed — any model that fills at the mid or last trade systematically overstates fills.
Dataset
The data a backtest reads
Fills
Ladder-walking fills and fees
At each tick the engine walks the historical bid/ask ladder: it reads size at the current level, models a partial fill when a level runs out, and applies slippage when size crosses the book. Exit reasons are logged — take profit, stop loss, window end, or no fill. Across 298 resolved BTC contracts, 90% of final-second books had a single side, a regime no mid-fill model represents.
Costs are regime-aware. Fee semantics changed at the CLOB V2 cutover on April 28, 2026: V1 fee rules before it, V2 taker-only fees after. The migration guide shows tagging each snapshot with the regime active at that moment so cost modeling stays honest across the boundary.
Get it
Where to get the data
- Backtest AI — plain-English strategies on the resolved archive, $29/mo standalone or +$19/mo add-on (overview).
- REST API + Python SDK — the same books, prices, and metrics at 250ms for your own strategy code (backtest API).
- Guide — the backtesting strategy guide covers look-ahead bias, signal-on-current-snapshot rules, and walk-forward design.
- Open sample — 897,192 snapshots across 805 resolved markets on Zenodo to prototype on before paying.
Worked example
The same idea through the record
Backtest data, here, means two things bound together: the corpus and the fill model. The corpus is the resolved crypto up/down set — 118,251 markets — with full L2 snapshots every 250ms, and the model walks that ladder at each tick, reading size per level, modeling partial fills when a level runs out, and applying slippage when an order must cross the book rather than take the mid.
The alternative everyone is comparing against is the mid-fill backtest, and the venue makes the difference measurable: 5-minute books are one-sided 16.9% of the time and 3.24% of snapshots are crossed, so mid-fill scoring quietly buys liquidity that was never there. On the stored books, the same strategy often tells a different, more honest story.
The discipline of a defensible backtest is to publish the corpus definition: windows (5m, 15m, 4h), coins (eight recorded), the resolved-only filter, the fee model, and the exit-reason table. Each of those is a design decision, and each changes the result.
A beginner's first run should be one trivial rule over one contract, reconciled by hand against the stored frames for the same second. It takes a few minutes, it anchors the toolchain, and it makes every later automated run legible.
Save the corpus definition with every run: windows, coins, resolved-only filter, fee and fill model, and the exit table — that paragraph of metadata is what turns a backtest into a citable result.
Signals
What to check before you trust it
Check the fill assumptions before the strategy: a backtest engine that cannot represent a one-sided book cannot represent this venue at all, whatever its speed claims.
Compare the same rule at 5m and 4h windows; the one-sidedness and crossed-book rates differ by window, and a rule tuned on one cadence can be silently worthless on another.
Use the exit-reason column as diagnostics: "window ended" is a structural fact about the book, while "stop loss hit" is a fact about direction, and conflating the two misdirects every fix.
Name your corpus in every published result; an unnamed corpus is an uncheckable claim, and the 897,192-snapshot open dataset exists precisely to make such claims checkable.
FAQ
Why backtest only on resolved markets?
Because the outcome is known, a strategy can be scored against what actually happened. PolyOrderbooks and Backtest AI use the same resolved-only, 118,251-market archive, so results are judged against reality rather than a prediction.
What resolution does backtest data support?
Full L2 snapshots every 250ms with prices and metrics on the same timestamps. A 5-minute market is roughly 1,200 snapshots. Official price history, by comparison, floors at 1 minute and carries no depth.
How should backtests model fees?
Regime-aware: apply V1 fee semantics before April 28, 2026 and V2 taker-only fees after, tagging each snapshot with the fee regime active at that moment so cost modeling stays correct across the CLOB V2 cutover.
Can I try the data before building?
Yes. A free ready-made BTC 5-minute sample is downloadable as CSV/JSON with no signup, 897,192 snapshots across 805 resolved markets are on Zenodo under CC BY 4.0, and Starter includes 3 days of live API history.