Sources

What's in the archive

Whole public archives queryable with read-only SQL. Reddit, Hacker News, TikTok, mailing lists, forums, Stack Exchange, OpenAlex, the scholarly paper catalog, Bluesky, prediction markets, and GitHub are documented in full; the wider corpus — newsletters, reference works, public records, and more — is described by family.

Documented in full

Reddit reddit.comments 30.7B+ posts + comments Live · see note The public Reddit archive — posts and comments queryable with read-only SQL by subreddit, author, timestamp, text, and score.Hacker News hackernews.items 45.3M+ stories + comments Live Every Hacker News story and comment, queryable with read-only SQL — source-native records, timestamps, authors, scores, and a live tail.TikTok tiktok.videos TikTok public-catalogue videos — 23.1 billion crawled videos from 139.6 million creators plus a 4.5-billion-video research snapshot (measured 2026-09-09), with descriptions, hashtags, engagement counts, and 3.39 billion caption-transcript rows over 1.38 billion videos — queryable with read-only SQL by creator and video.Forums forums.posts 44.4M+ posts + comments Live About 36 million posts and comments across roughly 4,300 forum sites — LessWrong, the EA Forum, DEV, DataSecretsLox, crypto-governance forums, and a long crawl-discovered tail — queryable with read-only SQL by site, thread, author, and text.Mailing lists mailing_lists.messages 795M+ messages Live Public mailing-list and Usenet archives — the Linux kernel list, GNOME, Fedora, extropians, SL4, and about 51,000 more lists — threaded and queryable with read-only SQL by list, author, date, and text.Stack Exchange stackexchange.posts 83.4M+ questions + answers Live Questions and answers across the Stack Exchange network — Stack Overflow, Mathematics, and every other landed site — queryable with read-only SQL by site, tag, author, score, and text.OpenAlex openalex.works 511.8M+ works Frozen The OpenAlex scholarly graph — 508 million works, 119 million author profiles, and 3 billion citation edges — queryable with read-only SQL by DOI, title, author, venue, year, and citation count.Scholarly papers academic.catalog 91.2M+ papers Live One merged bibliographic row per paper across arXiv, PubMed, PMC, Europe PMC (with bioRxiv and medRxiv), INSPIRE-HEP, and HuggingFace papers, enriched from OpenAlex — with full text on hand where a source supplied it.Bluesky bluesky.posts 2.3B+ posts Live Source-native Bluesky posts from the public firehose, archived and continuously observed — queryable with read-only SQL by author, time, and text.Manifold manifold.markets The full Manifold prediction-market corpus — every market, every bet, every comment — queryable with read-only SQL and kept minutes-fresh from the live API.Prediction markets markets.catalog Frozen One folded row per prediction market across Kalshi, Polymarket, and Manifold — title, category, status, resolution, open/close/settle times, volume, liquidity, and probability — queryable with read-only SQL.GitHub repositories github.repos Frozen Every public GitHub origin the Software Heritage archive has observed — 408 million repositories keyed by owner, including ones since deleted, renamed, or taken private — queryable with read-only SQL.

How coverage is stated

Each source page states what is covered: the public source, the SQL tables it lands in, the record shape, how fresh it is, and what is known to be missing — so an agent can tell the difference between "no results" and "not ingested". Known incompleteness is also served per relation as extent and known_holes in /v1/scry/schema.

The rest of the corpus

  • Community & discussionReddit, 2005 to now: 26.9 billion comments — 94% of every comment ever written, audited against Reddit's own ID counter — and every post. Every Hacker News item since 2006, minutes behind the site. Stack Exchange across the network, LessWrong, the EA Forum, and some 4,300 forum sites.
  • Video & chatYouTube: metadata for 4.52 billion public videos, 1.12 billion caption transcripts over 824 million videos in 301 languages, and 1.22 billion comments. TikTok's public catalogue with captions where creators enabled them. Twitch chat across half a million channels, captured live.
  • Social streamsBluesky's firehose and the Mastodon fediverse, landed as they are posted, and 1.28 billion 4chan posts. Vastly larger social archives sit behind special access, granted person by person.
  • Scholarly literature508 million OpenAlex works with the citation graph, 316 million carrying a DOI. arXiv, PubMed, PMC, Europe PMC with bioRxiv and medRxiv, INSPIRE-HEP, and a DOI-keyed journal full-text corpus, folded into one catalog that states which corpora hold each paper and whether its full text is on hand. Every ClinicalTrials.gov study.
  • Public records59 million SEC EDGAR filing documents in full text — annual and quarterly reports, current reports, insider filings. The DOJ Epstein releases as a queryable artifact index. Grant and funding records from public releases.
  • Code & packages408 million GitHub repositories as Software Heritage observed them — including the ones since deleted, renamed, or taken private; one owner lookup returns a person's entire public footprint. The package registries as one catalog: npm, Go, PyPI, NuGet, Packagist, crates.io, and about thirty more.
  • Mailing lists & Usenet811 million messages across 132,000 public lists and newsgroups — the Linux kernel list, netdev, Fedora, extropians, SL4, and the long tail — threaded by reply and root, queryable by list, author, date, and text.
  • Prediction marketsKalshi (188 million markets), Polymarket, Manifold, and Metaculus: markets, comments, and resolution context, plus all 23 million Manifold bets, each carrying the market probability before and after it.
  • Books, reference & the webA bibliographic catalog across OCLC/WorldCat, OpenLibrary, ISBNdb, and Google Books; public-domain full text in passages behind special access, granted person by person; FanFiction.net 1998 to 2015; Wikipedia, embedded for vector ranking. A live crawl of the web, and the cleaned reading layer of Common Crawl — articles, forums, and lists chosen by link-graph centrality, boilerplate stripped, every row keeping its WARC triple.

The full inventory

Each relation states its own extent and known holes in the schema. The complete per-source inventory — counts, freshness, embedding coverage — is shared with evaluating customers.