Polymarket Data in Node.js
Polymarket Data in Node.js
Node is the natural host for both scrapers and dashboards, and the REST API is plain JSON with a stable schema — a short typed client is all the glue you need.
Figures measured as of 2026-10-02 on the published PolyOrderbooks archive.
Client
The typed fetch
- fetch(endpoint + queryParams, { headers: { "x-api-key": key } }) with resolution=250ms and start_ts/end_ts.
- Parse with a zod schema (or plain interface) covering ts, bid, ask, mid, bidDepth, askDepth, vol — one type across all three datasets.
- Return rows sorted by sequence; the API sequences within a market, so page-boundary joins are safe.
- Handle 429 with Retry-After and resume at the last ts — redelivery is idempotent due to monotonic timestamps.
Pipeline
From rows to own storage
- Node worker (node-cron or a queue) pulls windows and appends Parquet via arrow-js.
- Serve via Fastify: a thin JSON gateway in front of your DuckDB/BigQuery store.
- Lambda-friendly: stateless fetch with a persisted cursor in S3 keeps deploys cold-start-light.
- For the MCP route, the repository's MCP server wraps the same endpoints for Cursor/Claude.
Notes
Keeping it truthful
Do not recompute depths from mids client-side — the archive already ships depth columns; your model should consume quotes and fills as measured.
Books, prices, and metrics come back over the same endpoints at the same 250ms resolution, so a dashboard needs one schema, not three.
Worked example
The cursor-worker skeleton
A worker fetches a window at 250ms, appends to Parquet via arrow-js, and persists a cursor in S3; on resume it re-fetches from start_ts and de-dupes by sequence, making retries safe.
A Fastify gateway serves your stored copy to the app with a NavBar-era typed client in ~40 lines because the schema is small and stable.
Sanity-checks
The numbers to sanity-check
Cursor math is the failure point: last-stored ts + 1 interval on resume must not skip frames; assert contiguous sequences per market in tests.
Notes
Going further
The MCP server in this repository wraps the same endpoints, so a Node process and an AI agent consume identical schemas.
Honest fit
Where this tool wins
Node is where the operational work lives: scheduled pulls, cursor persistence, Parquet append, and a typed gateway over your own copy — the pieces that productionize a data flow without a framework tax.
The honest contract is the schema: three datasets (books, prices, metrics) sharing one row shape of ts and fields; a typed client approaches the SDK's ergonomics in four dozen lines.
Notes
First repro
The MCP server in this repo invokes the kindred endpoints, so an agent and a cron worker type against the same response — one less integration to keep honest.
Get started
Your first solid pull
First worker pulls one market since last cursor to a Parquet file and prints health fields (last ts, rows). Then run it twice and confirm the second run appends zero duplicates — that is the idempotency test that matters.
Expose the /health route early; monitoring a data flow you cannot inspect is how drift gets in, and last-ts plus rowCount are the two numbers operators read first.
Conclusion
How to take it further
Node/TypeScript is the operational backbone of a serious deployment: scheduled pulls, idempotent appends, Parquet writes, and a typed gateway over your own copy. Once the worker resumes cleanly after a gap, the flow has passed the test that matters.
The schema is the contract: one row shape across books, prices, and metrics, typed once and reused everywhere including the MCP server. That single decision removes a whole class of drift bugs before they reach a dashboard.
Ship the /health route with last-ts and rowCount early; a data pipeline you can inspect is a pipeline you can trust, and the numbers operators read are the same numbers the sanity checks assert.
FAQ
Can Node.js query Polymarket order books?
Yes — fetch with the API key header and query params; responses are JSON arrays with a documented, stable schema.
Any TypeScript SDK?
Python is the primary SDK language; TypeScript teams typically write a zod-typed fetch in ~40 lines — the schema is small.
How do I backfill page gaps?
There are no pagination gaps when you filter by ts and resume from the last processed timestamp; sequence numbers make appends idempotent.