Sources
What's in the archive
Whole public archives queryable with read-only SQL. Reddit, Hacker News, TikTok, mailing lists, forums, Stack Exchange, OpenAlex, the scholarly paper catalog, Bluesky, prediction markets, and GitHub are documented in full; the wider corpus — newsletters, reference works, public records, and more — is described by family.
Documented in full
reddit.comments 30.7B+ posts + comments Live · see note The public Reddit archive — posts and comments queryable with read-only SQL by subreddit, author, timestamp, text, and score.Hacker News hackernews.items 45.3M+ stories + comments Live Every Hacker News story and comment, queryable with read-only SQL — source-native records, timestamps, authors, scores, and a live tail.TikTok tiktok.videos TikTok public-catalogue videos — 23.1 billion crawled videos from 139.6 million creators plus a 4.5-billion-video research snapshot (measured 2026-09-09), with descriptions, hashtags, engagement counts, and 3.39 billion caption-transcript rows over 1.38 billion videos — queryable with read-only SQL by creator and video.Forums forums.posts 44.4M+ posts + comments Live About 36 million posts and comments across roughly 4,300 forum sites — LessWrong, the EA Forum, DEV, DataSecretsLox, crypto-governance forums, and a long crawl-discovered tail — queryable with read-only SQL by site, thread, author, and text.Mailing lists mailing_lists.messages 795M+ messages Live Public mailing-list and Usenet archives — the Linux kernel list, GNOME, Fedora, extropians, SL4, and about 51,000 more lists — threaded and queryable with read-only SQL by list, author, date, and text.Stack Exchange stackexchange.posts 83.4M+ questions + answers Live Questions and answers across the Stack Exchange network — Stack Overflow, Mathematics, and every other landed site — queryable with read-only SQL by site, tag, author, score, and text.OpenAlex openalex.works 511.8M+ works Frozen The OpenAlex scholarly graph — 508 million works, 119 million author profiles, and 3 billion citation edges — queryable with read-only SQL by DOI, title, author, venue, year, and citation count.Scholarly papers academic.catalog 91.2M+ papers Live One merged bibliographic row per paper across arXiv, PubMed, PMC, Europe PMC (with bioRxiv and medRxiv), INSPIRE-HEP, and HuggingFace papers, enriched from OpenAlex — with full text on hand where a source supplied it.Bluesky bluesky.posts 2.3B+ posts Live Source-native Bluesky posts from the public firehose, archived and continuously observed — queryable with read-only SQL by author, time, and text.Manifold manifold.markets The full Manifold prediction-market corpus — every market, every bet, every comment — queryable with read-only SQL and kept minutes-fresh from the live API.Prediction markets markets.catalog Frozen One folded row per prediction market across Kalshi, Polymarket, and Manifold — title, category, status, resolution, open/close/settle times, volume, liquidity, and probability — queryable with read-only SQL.GitHub repositories github.repos Frozen Every public GitHub origin the Software Heritage archive has observed — 408 million repositories keyed by owner, including ones since deleted, renamed, or taken private — queryable with read-only SQL.How coverage is stated
Each source page states what is covered: the public source, the SQL tables it lands in, the record shape, how fresh it is, and what is known to be missing — so an agent can tell the difference between "no results" and "not ingested". Known incompleteness is also served per relation as extent and known_holes in /v1/scry/schema.
The rest of the corpus
The full inventory
Each relation states its own extent and known holes in the schema. The complete per-source inventory — counts, freshness, embedding coverage — is shared with evaluating customers.