Polymarket order book archive

Polymarket order book archive: full L2 ladders captured every 250ms

An order book archive that records the full L2 ladder four times a second and keeps [resolved markets](/polymarket-resolved-markets) queryable. 250ms capture on every plan, roughly 1,200 snapshots per 5-minute market, and an open 897,192-snapshot Zenodo dataset to verify the shape of the data without a key.

Figures measured as of 2026-10-02 on the published PolyOrderbooks archive.

Short answer

Stored books, not reconstructed ones

Polymarket's own CLOB API serves the live book; it does not archive L2 depth. PolyOrderbooks stores the complete book state at regular intervals instead — four full bid/ask ladders per second for crypto up/down markets across 8 coins.

Because the archive is stored rather than reconstructed, you can query the exact book that existed at any point in a market's life — including after resolution, with the winning outcome on every row. Websocket delta reconstruction cannot do that.

Under the hood

The archive in numbers

250msfinest queryable resolution
~1,200snapshots per 5-minute market
897,192open sample snapshots
805resolved markets in the sample

Schema

What a snapshot holds

Two tokens per marketA binary market has an Up token and a Down token, each with its own book. They are separate order books, not two sides of one.
Prices are probabilitiesEvery level sits in [0, 1]. A bid of 0.19 is someone offering 19 cents for a contract that pays 1.00 if that outcome wins.
Sizes are share countsMultiply by price for the dollar value resting at that level.
Ladders are best-price firstBids descend from the highest, asks ascend from the lowest.
Depth varies per snapshotMedian 39 bid levels on 5-minute markets, but it ranges from zero to over 120. Never assume a fixed number.

Get it

Where to get the archive

  • Free BTC sample — one market from open to settlement in CSV or JSON, no signup.
  • Zenodo dataset — 897,192 snapshots across 805 resolved markets and three contract lengths, CC BY 4.0 with DOI 10.5281/zenodo.22084114, mirrored on Hugging Face and Kaggle as Parquet (research data).
  • REST API — books, prices, and metrics at 250ms-to-1d query resolution on every plan, starting free.
  • Resolved markets kept — nothing is deleted at settlement; every past book stays queryable with its winning outcome.

Worked example

The same idea through the record

The whole point of a stored archive is that the book as it stood — full ladder, both sides, sizes, at a recorded UTC second — is queryable after the market closes. The archive captures full L2 books four times a second and keeps resolved markets queryable, so a study that starts three weeks after an event still opens the exact frames that were traded.

Contrast that with the live-only path: the official /book serves the current ladder and the websocket pushes deltas, but neither can answer what the book was at 14:42:50 yesterday. A replay of a public delta archive against stored snapshots diverged at 67.8% of 381,093 checkpoints, with 6.4% coming back crossed — because removal events do not survive a stream.

The archive's rows are the same stripes as live: price levels with sizes, a captured_at timestamp, flags for crossed states, and resolution attached after settlement. Because the whole book is stored whole, every metric can be recomputed from raw rows instead of trusted from a summary.

The practical first pull is one hour of one contract: compare the ledger of full ladders against a chart of the same window. The chart shows the path; the archive shows the book behind every point of it, and that comparison is the fastest way to see what "history of the book" really buys.

The archive's guarantee is stated plainly on every plan: full L2 books at 250ms, resolved markets queryable, and prices and metrics on the same grid — so the same window serves the book, the tape, and the summary in one join.

A useful reflex for anyone new: whenever a claim quotes a past book, ask for the stored frame; if the source cannot produce one, the number is a reconstruction and inherits the 67.8-percent divergence risk this page measures.

Signals

What to check before you trust it

Archive quality is a captured-state question: what matters is whether every second has its own whole book, not whether a summary row exists for the day.

The 897,192-snapshot open dataset — 805 resolved markets on Zenodo under CC BY 4.0 — is the place to audit the capture before paying for a wider window.

Resolved markets being queryable is the feature that matters for research: the entire settlement literature on this site is built on post-resolution queries, and none of it would exist a live feed.

When someone quotes a book at a historical second, ask for the raw frame: a number derived from a reconstruction carries the 67.8% divergence problem, while a stored frame is a fact.

FAQ

How is an archive different from Polymarket's live book endpoint?

The CLOB book endpoint returns the live order book at the current moment. An archive stores the complete book state at regular intervals, so you can query the exact book that existed at any timestamp — including for resolved markets.

What resolution is the archive stored at?

Full L2 snapshots every 250ms — four times a second — for crypto up/down markets. The API lets you query at any coarser grid from 1 day down to 250ms. A 5-minute market is roughly 1,200 snapshots.

What is in the open dataset?

897,192 snapshots across 805 resolved markets and 5-minute, 15-minute, and 4-hour contract lengths, published on Zenodo under CC BY 4.0 with DOI 10.5281/zenodo.22084114 and mirrored on Hugging Face and Kaggle as Parquet.

Can I inspect the data before paying?

Yes. A single BTC 5-minute market from open to settlement is available as a free CSV or JSON download with no signup, at the full 250ms structure.