Polymarket Data in Node.js

Polymarket Data in Node.js

Node is the natural host for both scrapers and dashboards, and the REST API is plain JSON with a stable schema — a short typed client is all the glue you need.

Figures measured as of 2026-10-02 on the published PolyOrderbooks archive.

Client

The typed fetch

  • fetch(endpoint + queryParams, { headers: { "x-api-key": key } }) with resolution=250ms and start_ts/end_ts.
  • Parse with a zod schema (or plain interface) covering ts, bid, ask, mid, bidDepth, askDepth, vol — one type across all three datasets.
  • Return rows sorted by sequence; the API sequences within a market, so page-boundary joins are safe.
  • Handle 429 with Retry-After and resume at the last ts — redelivery is idempotent due to monotonic timestamps.

Pipeline

From rows to own storage

  • Node worker (node-cron or a queue) pulls windows and appends Parquet via arrow-js.
  • Serve via Fastify: a thin JSON gateway in front of your DuckDB/BigQuery store.
  • Lambda-friendly: stateless fetch with a persisted cursor in S3 keeps deploys cold-start-light.
  • For the MCP route, the repository's MCP server wraps the same endpoints for Cursor/Claude.

Notes

Keeping it truthful

Do not recompute depths from mids client-side — the archive already ships depth columns; your model should consume quotes and fills as measured.

Books, prices, and metrics come back over the same endpoints at the same 250ms resolution, so a dashboard needs one schema, not three.

Worked example

The cursor-worker skeleton

A worker fetches a window at 250ms, appends to Parquet via arrow-js, and persists a cursor in S3; on resume it re-fetches from start_ts and de-dupes by sequence, making retries safe.

A Fastify gateway serves your stored copy to the app with a NavBar-era typed client in ~40 lines because the schema is small and stable.

Sanity-checks

The numbers to sanity-check

Cursor math is the failure point: last-stored ts + 1 interval on resume must not skip frames; assert contiguous sequences per market in tests.

Notes

Going further

The MCP server in this repository wraps the same endpoints, so a Node process and an AI agent consume identical schemas.

Honest fit

Where this tool wins

Node is where the operational work lives: scheduled pulls, cursor persistence, Parquet append, and a typed gateway over your own copy — the pieces that productionize a data flow without a framework tax.

The honest contract is the schema: three datasets (books, prices, metrics) sharing one row shape of ts and fields; a typed client approaches the SDK's ergonomics in four dozen lines.

Notes

First repro

The MCP server in this repo invokes the kindred endpoints, so an agent and a cron worker type against the same response — one less integration to keep honest.

Get started

Your first solid pull

First worker pulls one market since last cursor to a Parquet file and prints health fields (last ts, rows). Then run it twice and confirm the second run appends zero duplicates — that is the idempotency test that matters.

Expose the /health route early; monitoring a data flow you cannot inspect is how drift gets in, and last-ts plus rowCount are the two numbers operators read first.

Conclusion

How to take it further

Node/TypeScript is the operational backbone of a serious deployment: scheduled pulls, idempotent appends, Parquet writes, and a typed gateway over your own copy. Once the worker resumes cleanly after a gap, the flow has passed the test that matters.

The schema is the contract: one row shape across books, prices, and metrics, typed once and reused everywhere including the MCP server. That single decision removes a whole class of drift bugs before they reach a dashboard.

Ship the /health route with last-ts and rowCount early; a data pipeline you can inspect is a pipeline you can trust, and the numbers operators read are the same numbers the sanity checks assert.

FAQ

Can Node.js query Polymarket order books?

Yes — fetch with the API key header and query params; responses are JSON arrays with a documented, stable schema.

Any TypeScript SDK?

Python is the primary SDK language; TypeScript teams typically write a zod-typed fetch in ~40 lines — the schema is small.

How do I backfill page gaps?

There are no pagination gaps when you filter by ts and resume from the last processed timestamp; sequence numbers make appends idempotent.