About

The People's Daily N-gram Viewer shows term frequencies in 人民日报. It publishes counts and article metadata, never full article text.

How it works

Each release contains precomputed monthly counts in small gzip-compressed JSON shards. The browser fetches the shards for the requested terms. It needs no query engine or WebAssembly. Field, date range, resolution and display changes reuse those counts.

Tokens follow the tokenizer and dictionary recorded in the release manifest. Curated phrases use non-overlapping exact substring matches. A term absent from the index is not evidence that it never occurs in the newspaper. Missing terms must be added in a later release.

Measures

Occurrence frequency divides term occurrences by segmented token totals. Article frequency divides distinct matching articles by the number of articles. Both adds body and headline occurrences and token totals, but counts each article once. Front page means page 1. In Both co-occurrence queries, terms may occur in different fields of the same article.

Co-occurrence means that all requested terms appear in the same article in the selected field. It shows each individual term and their intersection. Its measure is always Articles. Same-page co-occurrence is not provided.

Months within coverage without matches have zero counts. Months outside coverage are absent. Year values sum counts and denominators before computing frequencies. Incomplete first and last years are marked as partial. Coverage describes this release, not a claim that the historical newspaper corpus is complete.

Smoothing

Percentage mode uses a centered moving average of the displayed period frequencies. A 3-year window includes the preceding year, the current year and the following year. Monthly mode uses 12 × the selected number of years, with floor(window/2) periods on each side, so its even-sized window contains one extra month. Edges use smaller windows. Zero-match periods within coverage participate. Absolute counts are not smoothed.

Downloads and links

Chart and article downloads are CSV or JSON. JSON includes citation, release, query and generation time. CSV has release and kind columns and a plain header row. The kind identifies a token, tracked phrase or co-occurrence series. Article previews are paginated newest first. Complete article exports show an estimated metadata transfer size, report progress and can be cancelled. PNG and R exports retain the displayed chart values and visible series. Monthly R exports use calendar dates.

Permalinks include the pinned release and hidden series. Old links without a release use the latest release. Every published release is retained; links resolve its immutable manifest without downloading inventory. An unavailable release offers a link to the latest.

Contact

matthew.peter.robertson小老鼠uni-mannheim.de