Polymarket Data in Parquet

Polymarket Data in Parquet

Parquet compresses wide timestamp tables dramatically and scans only the columns a query touches. For 250ms Polymarket rows it is the format that makes laptop research possible.

Figures measured as of 2026-10-02 on the published PolyOrderbooks archive.

Why

What Parquet changes

  • Compression: price/depth columns are low-cardinality and highly duplicate; Parquet typically 10–20x smaller than CSV.
  • Columnar projection: a spread query reads two columns, not whole rows.
  • Direct reads: DuckDB, Polars, arrow, ClickHouse, and BigQuery all read Parquet natively — no ETL.
  • Type safety: timestamps and numerics typed at write time instead of re-inferred at read time.

Shape

The suggested layout

  • One file (or daily partition) per market/series: books.parquet containing ts, seq, bid, ask, mid, bidDepth, askDepth, vol.
  • Add a _metadata footer for row-group stats (min/max per file) so readers prune whole files.
  • Snowflake-like partitioning: /market_type=updown/date=2026-09-01/books.parquet
  • Keep resolution in the filename (books_250ms.parquet) — re-deriving it later is guessing.

Access

Getting Parquet out of the archive

  • Self-serve downloads expose CSV and JSON per market; convert via the SDK's pandas/to_parquet step for scripts.
  • Bulk Parquet/CSV/JSON is available on enterprise S3 delivery for full-archive work.
  • Parquet then feeds DuckDB/ClickHouse for the heavy SQL — the lifecycle most published studies use.

Worked example

A file you can query on sight

Slice a market into a single books.parquet with ts+seq+quotes+depths+vol; row-group stats make DuckDB skip whole groups for a windowed query.

Name files books_250ms_PART.parquet and keep a COLUMNS.md; six months later the file is self-describing, which is the point of the format.

Sanity-checks

The numbers to sanity-check

A 250ms month for one market is ~10.4M rows; Parquet rarely needs more than a few hundred MB for that, versus GBs of CSV — a size sanity check on your files is a proxy for schema health.

Notes

Going further

The archive ships bulk Parquet on enterprise delivery, so lake teams get typed files without writing a transform.

Honest fit

Where this format wins

Parquet is the material the rest of the pipeline assumes: typed timestamps, column pruning, and 10x-plus compression over CSV turn an unwieldy 250ms record into a structure any SQL engine reads without ETL.

The honest discipline is naming and row-group hygiene: resolution and partitioning encoded in the path, stats in the footer — because a Parquet file that lies about what it holds is worse than a CSV you can eyeball.

Notes

First repro

A full market-month at 250ms (~10.4M rows) as Parquet is typically a few hundred MB versus GBs of CSV — the size delta is itself a schema sanity check.

Get started

Your first solid pull

First file: one market, one day at 250ms, written as books_250ms.parquet, and immediately queried by DuckDB with a window filter to demonstrate column pruning reduces the scan.

Check size and stats next: PARQUET file_size and the min/max per row group should match the nominal ts range; that double check is the receipt every later pipeline trusts.

Conclusion

How to take it further

Parquet is the layer that makes every other choice work: typed timestamps, columnar scans, and compression that keeps a full 250ms market-month at a few hundred megabytes on a laptop.

The habits that protect the lake are cheap: store resolution in the name, sort by (market, ts), write row groups with healthy min/max, and keep a checksummed manifest. Each is one line of code and each prevents a silent-corruption debugging session.

From DuckDB on a laptop to Athena across a bucket, every reader speaks the same format — so the file layout you choose today is the one every future tool inherits.

FAQ

Why is Parquet better than CSV for Polymarket data?

10–20x smaller, typed timestamps, and columnar scans mean a 250ms series for a month is still laptop-queryable.

What tools read Polymarket Parquet?

DuckDB, Polars, arrow, pandas, ClickHouse, BigQuery, and standard SQL engines — all direct, no conversion.

Can I convert the CSV download to Parquet?

Yes — load the CSV with DuckDB or pandas and write to_parquet; timestamps and numerics re-type cleanly from the documented schema.