Protodex vs PolyOrderbooks
Protodex vs PolyOrderbooks
Protodex is a one-time CSV dataset: 15-minute price and order book snapshots across 18,400+ Polymarket markets, sold on its site as a flat purchase. PolyOrderbooks is a queryable 250ms archive. If your workflow is a single offline backtest, the flat dataset can be the cheaper answer; the difference is what you trade for it.
Figures measured as of 2026-10-02 on the published PolyOrderbooks archive.
What Protodex sells
The flat-file dataset on its own site
Protodex lists a Polymarket Historical Dataset as clean CSV: 15-minute snapshots of the live Polymarket order book and mid price across 18,400+ markets over 76+ trading days, with rows carrying the market question, outcome side, mid, best bid and ask, bid/ask depth, spread, and UTC timestamp.
Its own page is careless with its headline numbers — the title claims 23.0M+ price snapshots while the body says 18.5M+ — which is exactly the kind of inconsistency to price into trust. The product is what it is: a reasonably clean, gap-checked, 15-minute grid you can buy once for $19 or refresh weekly for $19/month.
For a single research project or model-training run, that flat purchase can be the most direct path to Polymarket-shaped data available anywhere.
Granularity
What 15-minute sampling cannot see
A 15-minute grid captures six to twelve samples across a typical 90-to-180-minute market life — fine for trend-level asks, useless for the settlement dynamics that decide Up/Down markets, where the final seconds are the trade. Published measurement on this site shows books end one-sided in over 90% of resolved BTC five-minute contracts with median top-of-book size above 42,000 shares.
PolyOrderbooks serves that story directly: 250ms capture with prices and metrics in the same rows, paid windows out to 30–120 days, resolved markets queryable. You can, and should, also feed the result to your own Polars or DuckDB pipeline instead of receiving someone else's pre-gridded CSV.
Cost shape
One-time vs recurring
- Protodex full dataset — $19 one-time, as it stands at purchase; CSV, UTF-8, commercially usable with no redistribution.
- Protodex weekly auto-refresh — $19/month re-delivery for teams that keep models current.
- PolyOrderbooks free Starter — $0 for 250ms L2 with 3 days of history, no card, same resolution as paid plans.
- PolyOrderbooks Pro — $19/month at the time of writing for multi-day history at 250ms with metrics; enterprise S3 delivery (Parquet/CSV/JSON) for bulk.
Honest split
Pick by workflow, not by margin
If you need one bounded CSV for a bounded study and a 15-minute grid is adequately dense for it, Protodex is the honest recommendation — $19 is hard to argue with.
If your work is iterative — backtests that cross resolution, markets that settle inside your sampling window, or questions that need more than 76 days — a queryable 250ms archive costs more but keeps answering. The free Starter tier exists precisely so you can measure the difference before paying anything.
Worked example
The same job through both lenses
A concrete test: price the settlement-fundamental question with a 15-minute grid. A 90-minute Up/Down market yields six to twelve snapshots; the final-minute dynamics that actually decide the contract — the one-sided state, the bound traffic, the reprice seconds — are all inside the gap between two of your samples. The flat dataset answers "which trend won"; it cannot answer "what was tradable at 14:42:50", because no 15-minute sample landed on that second.
The same market in the 250ms archive yields thousands of frames across that hour, and one of them is the recorded second itself: floor, both sides, sizes, the reference print that moved it. That difference is not a performance question — it is a resolution question, and resolution is the deliverable when the settlement window is where contracts are won.
For a single bounded backtest where 15-minute trend questions genuinely suffice, Protodex at $19 once is the pragmatic buy; the price is hard to object to, and a flat CSV removes every integration loop. The trade is the window: 76 days, closed schema, no late markets beyond that span, and a headline dataset count that the provider's own page contradicts.
Decision
How to pick
Choose Protodex for one bounded, trend-level study with a fixed horizon and a closed schema. Choose the archive when the work iterates — backtests that change resolution, markets that settle inside the sampling grid, or questions that outlive 76 days.
Either way, verify the raw rows against one known event before trusting the whole stack: pull the same reprice hour from each source, diff the depth columns, and the honest choice becomes a data-quality statement instead of a price argument.
FAQ
Is Protodex's Polymarket dataset reliable?
It is a clean, sold-as-CSV dataset with a permissive commercial license, and its own page documents a 15-minute sampling interval. Its headline counts are inconsistent between its title (23.0M+) and body (18.5M+), so treat its figures with normal care.
Is 15-minute sampling enough for Polymarket backtests?
Only for questions above the move-to-move level. Up/Down settlement plays out in seconds; 15-minute grids cannot see them. Match the grid to the strategy's shortest decision horizon.
Why would I pay more for an archive instead of buying the CSV?
Because the CSV is one cut of the past and the archive is the ongoing record. If you will iterate on the question, or revisit markets that resolve inside your sampling window, a 250ms queryable archive keeps working where a flat file does not.