Data

The viewer uses public, immutable lookup releases. The root manifest.json points to the latest release and lists all releases. Each immutable v/{release}/manifest.json records coverage, the tokenizer, shard counts and the release base. Every page pins one release for its session.

Published files

  1. stats.json contains monthly token and article denominators.
  2. c/{shard}.json contains term counts and distinct article counts for body, headline and their union.
  3. d/{shard}.json contains base64-encoded unsigned LEB128 pairs of article ordinal deltas and occurrence counts. Body and headline each have disjoint front-page and non-front-page lists.
  4. docs/months.json maps ordinals to months. Ordinals are specific to a release.
  5. meta/{YYYY-MM}.json contains article IDs, dates, page numbers, headlines and other metadata. It never contains article bodies.
  6. meta/sizes.json supplies compressed sizes for export estimates.
  7. phrases.json lists tracked phrases. inventory.json records object sizes and SHA1 hashes.

Terms route to decimal-numbered shards with FNV-1a over their exact UTF-8 bytes, modulo the manifest's shard count. JSON objects are gzip-compressed and served with Content-Encoding: gzip; browsers decompress them automatically.

Reuse

Use the viewer's Chart values downloads for time series and Articles downloads for matching metadata. Both CSV and JSON retain the release and term kind. JSON also includes the citation and query. Complete article exports are not silently capped. The citation is shown beside the download controls.

Bulk downloads and API examples are not provided on this page yet. The old browser SQL and Parquet interface is no longer used.