Polymarket websocket data

Polymarket websocket data: live streams vs stored 250ms books

Polymarket's order-book websocket is a [live event stream](/polymarket-api-data), not a historical archive. We replayed its deltas against stored snapshots to measure the gap: **67.8% of reconstructed books diverged** from the stored snapshots and **6.4% came back crossed** — a book with the bid above the ask.

Figures measured as of 2026-10-02 on the published PolyOrderbooks archive.

Short answer

Live-only streams, no history endpoint

Polymarket exposes CLOB websocket channels — market, book, and price — that push events for markets you subscribe to. They are live-only: there is no endpoint to request yesterday's book deltas.

For anything that needs the historical book — a backtest, a replay, a study of settlement — a stored 250ms archive is the record of what the book actually was.

Measured

The replay experiment

67.8%of reconstructed books diverged
6.4%came back crossed
250msstored snapshot resolution
0websocket history endpoints

Why

Why delta replay fails

We took the archive's own snapshots and tried to rebuild them by applying websocket deltas from the same capture. 67.8% of the reconstructed books diverged from the stored snapshot, and 6.4% came back crossed.

The divergence is almost entirely missed events — a quote removed, a level cleared — that a live stream cannot deliver retroactively. A stored snapshot archive does not have that problem: the book is captured whole, four times a second. The methodology is in the replay-accuracy post.

Alternative

What a stored archive gives you

Worked example

The same idea through the record

The Polymarket websocket is a live event stream: market, book, and price channels that push events for subscribed markets. It is genuinely free and keyless and it moves roughly 24,000 events per second across the venue, but it is forward-only, with no endpoint that can return yesterday's book deltas.

The replay experiment measures the consequence: rebuilding the archive's own snapshots from captured deltas diverged at 67.8% of 381,093 checkpoints, and 6.4% of the reconstructions came back crossed — bid above ask. The missing pieces are removals: a level cleared, a quote pulled, an event that a stream cannot replay retroactively.

That is why the archive stores whole books at 250ms rather than trusting delta reconstruction: the book is captured complete four times per second, so there is nothing to rebuild and no divergence to audit. A stored frame at a past second is a fact; a reconstructed frame is an estimate with a measured failure rate.

The practical workflow for real-time users is to run live and archive in parallel: let the websocket drive current decisions, and pull the stored frames for any second that later matters. Live reads stay free; the archive is the receipt for everything the stream cannot answer after the fact.

For most users the winning setup is hybrid: the websocket for millisecond awareness and a stored-archive query for anything that will be cited later, with the two reconciled on the UTC second.

The decision rule in one line: if a second will matter next week, store it now, because a stream can tell you today what happened, and only an archive can tell you what the book was at a second that has already passed.

Signals

What to check before you trust it

Treat the absence of sequence numbers as the design constraint it is: with roughly 24,000 events per second and no sequencing, a dropped message cannot be detected, so any downstream audit needs the stored snapshot as the reference.

When validation matters, compare the stream against a stored frame for the same second rather than against another stream; two streams can diverge together and still look consistent.

A crossed book in a reconstruction is an alarm, not noise — 6.4% of checkpoints returning bid-above-ask means the replay silently priced an impossibility into the data.

For anything that becomes evidence — a backtest, a settlement study, a published number — pull the stored frame instead of the reconstructed one, because evidence is judged by provenance.

FAQ

Does Polymarket offer historical websocket data?

Not for order books. The websocket streams are live-only, and delta-based reconstruction is unreliable: in our measurement against stored snapshots, 67.8% of rebuilt books diverged and 6.4% came back crossed. A stored 250ms snapshot archive is the reliable way to get historical depth.

Why does delta replay fail?

Reconstruction requires every event; a missed quote withdrawal or level clear cannot be recovered later. Measured against the archive's own snapshots, 67.8% of rebuilt books diverged and 6.4% came back crossed — methodology in the replay-accuracy post.

What resolution does the archive have?

Full L2 snapshots every 250ms for crypto up/down markets, queryable from 1 day down to 250ms. Source data for the archive, prices, and metrics is the same 250ms grid. Free Starter includes 3 days of history.

Can I also get live data?

For live order books and prices, Polymarket's own CLOB websocket works well. Use a stored archive when you need the historical book that the stream cannot give you.