Datasets guide
The Polymarket Datasets Guide
The Polymarket dataset landscape in 2026 has four meaningful layers: the recorded 250ms archive, the open Zenodo release, the official live feeds, and the third-party benches. This guide inventories them and says honestly what each is for.
Figures measured as of 2026-09-24 on the published PolyOrderbooks archive.
Catalog
The four layers
Choose
Which layer for which job
Live dashboards read layer 3; filling-level backtests and audit work read layer 1; teaching and citable examples read layer 2; competitive scoping reads layer 4. The honest decision rule: match the layer to the claim — a paper citing "the dataset" must name which one.
The layers join cleanly because every timestamped one uses the UTC grid — a dashboard can overlay live and recorded panels without reconciliation.
Honest
The honest framing
This guide's job is legibility, not promotion: it names what exists, what each layer can prove, and where the recorded archive sits in that stack — and it points to the comparison pages for any claim about fitness for a specific use.
Trace
Tracing a number to its layer
The guide's real value is traceability: a claim like "final-second books are single-sided 90% of the time" traces to layer 1 (recorded ladders), its teaching copy traces to layer 2 (open Zenodo), and a miss about live pricing traces to layer 3 (official feeds). Before trusting any Polymarket statistic, ask which layer produced it.
That four-layer discipline is what keeps this catalog honest about what each dataset can and cannot prove. The comparison pages extend the same rule to third-party sources.
FAQ
What datasets exist for Polymarket?
The recorded 250ms archive, an open Zenodo snapshot (CC BY 4.0), the official live feeds, and third-party benches — each for a different job.
Which dataset is citable?
The Zenodo release with DOI 10.5281/zenodo.22084114 is the citable open artifact.
How do layers join?
Live and recorded layers share the UTC grid, so panels and studies overlay without time reconciliation.