Data
The viewer uses public, immutable lookup releases. The root manifest.json points to the latest release and lists all releases. Each immutable v/{release}/manifest.json records coverage, the tokenizer, shard counts and the release base. Every page pins one release for its session.
Published files
stats.jsoncontains monthly token and article denominators.c/{shard}.jsoncontains term counts and distinct article counts for body, headline and their union.d/{shard}.jsoncontains base64-encoded unsigned LEB128 pairs of article ordinal deltas and occurrence counts. Body and headline each have disjoint front-page and non-front-page lists.docs/months.jsonmaps ordinals to months. Ordinals are specific to a release.meta/{YYYY-MM}.jsoncontains article IDs, dates, page numbers, headlines and other metadata. It never contains article bodies.meta/sizes.jsonsupplies compressed sizes for export estimates.phrases.jsonlists tracked phrases.inventory.jsonrecords object sizes and SHA1 hashes.
Terms route to decimal-numbered shards with FNV-1a over their exact UTF-8 bytes, modulo the manifest's shard count. JSON objects are gzip-compressed and served with Content-Encoding: gzip; browsers decompress them automatically.
Reuse
Use the viewer's Chart values downloads for time series and Articles downloads for matching metadata. Both CSV and JSON retain the release and term kind. JSON also includes the citation and query. Complete article exports are not silently capped. The citation is shown beside the download controls.
Bulk downloads and API examples are not provided on this page yet. The old browser SQL and Parquet interface is no longer used.