News Ngram

Source on GitHub

Phrase frequency in the news, via GDELT. Each value is a share of the coverage GDELT monitored in that interval, so outlets that simply publish more do not dominate. Online news and TV are plotted separately — here is why.

Reading these charts

Two charts, never one

The same phrase runs about 6.5× hotter on the online-news chart than on the TV chart, and over the 94 months where both have data the two move together only loosely (r = 0.69). In January 2023, for instance, “climate change” sat near its high in online news and near its low on television.

They are not the same measurement. One counts articles, the other counts airtime. One watches the world, the other watches nine US networks. Putting them on a shared axis would flatten TV into a line along the floor; giving them two axes on one chart would let any pair of scales manufacture whatever agreement you wanted to see. So each gets its own chart, its own axis, and its own date range.

What the percentage on the y-axis is

Different things on each chart, which is the whole reason there are two.

On the online news chart it is the percent of the articles GDELT monitored worldwide in that interval that contained your phrase. Read 1.1% as: about one in ninety articles mentioned it.

On the TV chart it is the percent of monitored airtime. GDELT cuts each broadcast into 15-second clips and counts the ones whose captions contain your phrase; the line is that percentage averaged across the stations listed below that were on air at the time. For scale, CNN alone contributes roughly 130,000 clips a month.

Both are shares rather than counts, so an outlet or a network that simply produces more does not dominate. But a percent of articles and a percent of airtime are not the same quantity, and the ratio between them is not meaningful.

Because TV is a mean across stations, it hides disagreement between them — a phrase saturating one network and absent from the rest reads as a modest middling value.

Which stations the TV chart covers

Eleven — six cable networks, and the five major broadcast networks as their San Francisco affiliates. The year is when GDELT started indexing each one.

Those start years matter, because the panel grows over time: four stations in 2009, nine from mid-2010, all eleven from the end of 2013. A station enters the average only once it exists. GDELT reports a flat zero for a station in the months before it was indexed, which is indistinguishable from one that was on air and simply never said your phrase — averaging those zeros in would push the early years down for no reason, so they are excluded instead.

The broadcast networks are one city's affiliates, not national feeds. GDELT has no national ABC, CBS, NBC or PBS; it indexes them station by station, and San Francisco is the only market whose affiliates run through the end of the archive. Washington DC's stop between 2013 and 2019, Philadelphia's in 2018, Chicago's in 2015, and New York and Los Angeles are not in the archive at all. KGO carries ABC's national programming plus Bay Area local news, so treat it as a usable stand-in for ABC rather than as ABC itself — local stories will carry a California accent.

C-SPAN is deliberately excluded. Its three channels are in the archive, but they carry unedited floor and committee proceedings, which gives legislative phrasing far more weight than any newscast would. BBC News, Al Jazeera, Deutsche Welle and Russia Today are indexed as international channels and left out of a US average for the mirror-image reason. NPR is absent because GDELT publishes no radio data at all — the archive is built from television captions, and there is no radio equivalent to query.

Where the data stops

Online news covers 2017-01-01 onward. That is GDELT's own floor, not a setting here; earlier dates are clamped. TV covers 2009-07-02 through 2024-10-31 — it reaches back seven years further, but the archive stops in late 2024, so it cannot tell you about the present. Ask for a range outside either window and that chart quietly narrows to what exists; the date line under each title always states what you actually got.

Searching

Multi-word phrases are matched exactly. Query syntax passes straight through to GDELT, so source filters work: "no evidence" (domain:reuters.com OR domain:apnews.com). A line already containing ", ( or : is sent verbatim; anything else is quoted for you.

GDELT throttles aggressively and will refuse several requests in a row before answering. This page waits and retries, which is why a plot of several phrases takes a while.