W14 - Michel et al. Quantitative Analysis of Culture Using Millions of Digitized Books

Quantitative Analysis of Culture

Introduction to Culturomics

  • Authors: Jean-Baptiste Michel et al.

  • Corpus: A digital collection of texts, containing about 4% of all printed books.

  • Timeframe: Cultural trends are analyzed from 1800 to 2000, highlighting linguistic and cultural phenomena reflected in the English language.

  • Purpose: Provides insights into lexicography, grammar evolution, technology adoption, censorship, and historical epidemiology.

  • Culturomics: Extends scientific inquiry boundaries, allowing for quantitative analysis of cultural phenomena.

Corpus Creation

  • Books Studied: 5,195,769 digitized works from over 40 global university libraries.

  • Digitization Process: Involves optical character recognition (OCR) technology.

  • Data: The corpus includes English (~361 billion words), French (45 billion), Spanish (45 billion), German (37 billion), Chinese (13 billion), Russian (35 billion), and Hebrew (2 billion).

  • Growth Timeline: The corpus grew from several hundred thousand words in the early 1800s to 11 billion words by 2000.

Data Analysis Techniques

  • 1-grams and n-grams: A 1-gram is defined as a single uninterrupted string of characters. An n-gram is a sequence of 1-grams. The study considers n-grams that occur at least 40 times.

  • Usage Frequency Calculation: Involves tracking the number of n-gram occurrences relative to total words in the corpus for that year.

  • Example: The frequency of the term "slavery" peaked during the Civil War.

Cultural and Linguistic Trends

  • Two highlights of culturomic trends:

    • Cultural Change: Identifies concepts that evolve (e.g., "slavery").

    • Linguistic Change: Tracks how language adapts over time (e.g., shifts from "the Great War" to "World War I").

  • Over 2 billion trajectories can be explored through the dataset at culturomics.org.

The Size of the English Lexicon

  • Common 1-grams: Criteria based on frequency (greater than one per billion).

  • Word Estimates:

    • 1900: 544,000 words

    • 1950: 597,000 words

    • 2000: 1,022,000 words

  • Growth Rate: Average addition of ~8,500 words/year.

  • Comparison with Dictionaries: Gaps between lexicon and dictionaries highlight many undocumented words, termed "dark matter."

Dictionary Coverage of the Lexicon

  • Coverage Insights: High-frequency words are well represented in dictionaries, but numerous low-frequency terms are not.

  • Lexical Dark Matter: Approximately 52% of the English lexicon consists of words absent in standard references, especially in lower-frequency ranges.

  • Recent Additions: Examined changes in dictionaries, observing historical words often remain omitted until their frequency increases significantly.

Evolution of Grammar

  • Irregular Verbs: Analysis of regularity changes among English irregular verbs over the past 200 years.

  • Stability vs. Change: 16% of irregular verbs showed significant changes in usage regularity, indicating a slow transition.

  • Regularization Trends: Some verbs, previously irregular, are conforming to regular patterns, and vice versa.

Collective Memory and Forgetting

  • Analyzed frequency plots of 1-grams to assess societal memory decay over time.

  • Observations: Interest in older events diminishes rapidly compared to modern interests, showing changes in how collective memory functions.

Cultural Adoption of Technology

  • Invention Cohorts: Study of the adoption rates of inventions shows a trend toward quicker assimilation into society over time.

Fame and Cultural Memory

  • Fame Trajectories: Rise and fall of fame can be tracked through name frequency in historical databases.

  • Celebrity Dynamics: The age of initial fame and the rate of rise to prominence have decreased, indicating faster cycles of fame and memory.

Detecting Censorship and Suppression

  • Censorship Indicators: Analysis of name frequency to identify censorship patterns in different historical contexts (e.g., Nazi Germany).

  • Censorship Impact: Identified suppressed individuals and organizations through their diminutive representation in literature during oppressive regimes, indicating how suppression alters cultural influence.

Conclusion on Culturomics

  • Definition: Culturomics refers to the high-throughput collection and analysis of cultural data, providing quantitative evidence for various fields in the humanities.

  • Future Directions: The need for incorporating diverse formats beyond books (e.g., newspapers, artwork) into cultural analysis for comprehensiveness and accuracy.