Polymarket Data in R

Polymarket Data in R

R handles the API directly: httr GETs the JSON, jsonlite parses it, and data.table reshapes. The wide timestamp schema maps to long-form studies without ceremony.

Figures measured as of 2026-10-02 on the published PolyOrderbooks archive.

Pull

The request

  • httr::GET(endpoint, query = list(start_ts, end_ts, resolution = "250ms"), add_headers("x-api-key" = key))
  • jsonlite::fromJSON on the response gives a data.frame with one row per timestamp.
  • Combine the three datasets (books, prices, metrics) by timestamp; the archive timestamps them at the same 250ms grid.
  • For sunset-style analysis, filter with data.table between two POSIX boundaries — no join gymnastics required.

Study

Typical study shapes

  • Settlement tables: per contract keep the final 60 seconds, mark one-sided frames, tabulate p(empty ask) by time-to-resolve.
  • Spread series: (ask - bid) per row, then rolling quantiles by market type.
  • Depth-imbalance regressions: log(bidDepth/askDepth) ahead of the mid move — time the lead on your own data.
  • Cross-venue: merge Polymarket rows with exchange data on floor(utc, second); the same-second studies illustrate the join.

Limits

Scale notes

A full crypto month at 250ms is a large in-memory object; filter to markets and windows before pulling, and write final frames to Parquet via the arrow package.

For bulk research, enterprise S3 delivery lands Parquet/CSV/JSON that arrow reads directly — R becomes analysis-only, which is where it is strongest.

Free Starter reads 3 days of history at 250ms; Pro extends windows to 30–120 days for crypto and 30 for sports.

Worked example

A settlement table in R

Load a resolved contract slice with httr and jsonlite, then filter to the final 60 seconds with data.table to build the settlement view: per row, tag askDdepth handling.

Summarize with one group-by: n frames, pct with bidDepth>0&askDepth>0, and the last executable mid before resolution — the columns the archive ships map directly to rows.

Sanity-checks

The numbers to sanity-check

On a known resolved market, R should reproduce the archived settlement outcome; if the last frame in R differs from the archive, the ts window boundary is off by one interval.

Notes

Going further

For cross-venue studies, join on the exact UTC second — Polymarket rows at 17:88:xx and Binance ticks share the second bucket, which is what makes same-second comparisons valid.

Honest fit

Where this tool wins

R is the tool where the archive's claims should be reproduced: the settlement flag, the one-sided ratios, the same-second cross-venue match to exchange ticks — none of it should be taken on faith if you can rebuild it from 250ms rows.

Keeping the workload analytical (pull → frame → test) rather than operational keeps the borrow-checker out of the way; R's verbs map cleanly to the field names.

Notes

First repro

Cross-venue joins collapse to floor(utc) seconds once timestamps are UTC: the ETH-vs-Binance same-second pattern is the canonical example in the published studies.

Get started

Your first solid pull

First repro targets one resolved market: pull the full 250ms window, filter to the final 60 seconds, and verify your one-sided flags match the archive CSV flags for the same contract.

A quick win that builds trust: the same-second cross-check with exchange ticks. Take one UTC second, join the Polymarket frame to the exchange tick on floor(utc), and confirm both moved — the pattern repeats for every market.

Conclusion

How to take it further

R closes the loop between the archive and the conclusions: the settlement flags, the one-sided ratios, and the cross-venue matches are all repoducible from 250ms rows, which is exactly how a data claim should be treated — reproduce it before you build on it.

Keep the toolbox small: httr for the fetch, jsonlite for the parse, data.table for the reshape, arrow for Parquet out. The field names are documented and stable, so once a script runs against one market it runs against every market in the archive.

As the workload grows past laptop memory, the same verbs move from R to DuckDB for the scan and R for the analysis; that hand-off is the least painful scale path in this stack, and the published studies in this guide are built the same way.

FAQ

Can R call the Polymarket order book API?

Yes — httr against /v1/markets/{slug}/books with resolution and time parameters, jsonlite to parse, no SDK required.

What parsing pitfalls exist?

Keep timestamps as UTC character or POSIX, do not let Excel-style parsing happen, and store API keys outside scripts.

How large can an R workspace go?

Comfortably tens of millions of rows in data.table on a modern machine; beyond that, filter windows server-side or use S3 Parquet with arrow.