Enlightenment Ideals and Belief in Progress: Almelhem et al. (2026)
Distilled by claude-sonnet-4-6 · extracted Jun 28, 2026, verified Jun 28, 2026
JEL (IAR-assigned): C81, C88, N33, N63, O14, Z11 · assigned from the abstract, not the journal
What this is. The paper’s core results, the LDA-based classification method it applies, and the two estimating regressions with real equations: enough to know what it found and how, without reading the full 52 pages. To replicate or extend, read the original at https://doi.org/10.1093/qje/qjaf054.
The paper applies Latent Dirichlet Allocation to 264,443 English volumes from the HathiTrust Digital Library (printed in England, 1500-1900) to trace how the languages of science, religion, and political economy evolved in the centuries leading to the British Industrial Revolution. Three findings emerge. First, the languages of science and religion diverged in the mid-eighteenth century: science volumes that had used roughly 30% religious language in the early eighteenth century used only about 10% by 1850. Second, regression analysis shows that volumes using language at the nexus of science and political economy became the most progress-oriented beginning in the late seventeenth century, while volumes using purely scientific language were largely neutral. Third, within this nexus, those that also used the language of industrialization were the most progress-oriented from the mid-eighteenth century onward. The findings support Mokyr (2016)‘s Industrial Enlightenment thesis: it was pragmatic, industrially oriented scientific writing aimed at a broad literate audience, not elite scientific discourse, that carried progress-oriented culture into Britain’s economic take-off.
Core results
Section titled “Core results”Magnitudes and descriptions are as reported; locators point into the source PDF.
| # | Result | Locator | Magnitude |
|---|---|---|---|
| R1 | Languages of science and religion became distinct in the mid-eighteenth century; science volumes ceased to use religious language | Figure II, p. 284; Figure III, p. 286 | Science volumes used ~30% religious language in the early 18th century, declining to ~10% by 1850; by 1750 essentially no volumes sit at the science-religion vertex of the language simplex |
| R2 | Average progress sentiment rose from the mid-seventeenth century and persisted through the period | Figure V, p. 291 | Progress score (percentile) rose from ~10th-15th percentile pre-1650 to ~50th-60th percentile by 1800-1850 |
| R3 | Science-political economy nexus volumes were the most progress-oriented from ~1700 onward | Figure VII, p. 295; Online Appendix Table B.1 | Predicted progress score for 50%/50% science-political economy mix is highest among all language combinations beginning late 17th century; volumes using purely scientific or purely religious language score lower |
| R4 | Industrial language at the science-political economy nexus amplified progress orientation from the mid-eighteenth century | Figure XI, p. 305; Online Appendix Table B.4 | Within the 50%/50% science-political economy nexus, volumes at the 75th percentile of industrial language had approximately 2x the predicted progress score of zero-industry volumes by 1800 |
| R5 | Pattern is robust to alternative category definitions; replacing political economy with law or economics yields the same 18th-century rise | Figure VIII, p. 298 | Science-law and science-economics nexus volumes both show a sharp 18th-century rise in predicted progress sentiment, mirroring the science-political economy finding; the pattern does not emerge for arts and literature |
Overall (paper’s conclusion). The results are consistent with Mokyr (2016)‘s claim that the Industrial Enlightenment diffused progress-oriented views of science into industry and political economy. It was the literate artisan and applied-science audience, not the elite scientific community, whose language became most progress-oriented in the run-up to Britain’s industrialization.
Theory / model
Section titled “Theory / model”The paper has no formal economic model. The empirical analysis tests three subsidiary hypotheses derived from Mokyr (2016) and Mokyr (2009)‘s Industrial Enlightenment and Culture of Growth theses:
- The language of science and religion became increasingly distinct during the Enlightenment (secularization of science).
- The language of science became more progress-oriented during the Enlightenment, with the effect concentrated at the nexus of science and political economy rather than in pure scientific discourse.
- Volumes using the language of industrialization at the science-political economy nexus were particularly progress-oriented in the period before and during Britain’s Industrial Revolution.
Identification strategy. The analysis is descriptive: the regressions are accounting exercises documenting how progress-oriented language correlates with category weights and industrial language over time. The authors explicitly state that the regressions “are not meant to imply a causal relationship, as omitted variable biases and reverse causation may be present” (p. 293, p. 305). The evidence is structural in the sense of testing whether the pattern predicted by Mokyr (2016) is present in the data, but causality is not claimed.
Method
Section titled “Method”The method builds on the lda-topic-model technique applied to historical text corpora. Related work using this approach includes Erikson (2021), who applies LDA and sentiment analysis to political and economic tracts from England, 1550-1720, and Grajzl and Murrell (2024), who study English print culture across 1530-1700; both cover shorter time windows than the 400-year corpus here.
LDA topic model
Section titled “LDA topic model”The corpus is a document-term matrix , where is the number of volumes and is the vocabulary size. LDA (Blei, Ng, and Jordan (2003)) models each volume as a mixture over topics and each topic as a multinomial distribution over words. The optimal is chosen by 4-fold cross-validation on perplexity (Section II.D, p. 276-277). The output is, for each volume and topic , a weight representing how strongly the topic appears in that volume, with for each volume.
Topic categorization and volume classification
Section titled “Topic categorization and volume classification”Topics are grouped into three categories (science, religion, political economy) based on topic-pair co-occurrence. For each topic pair and each volume , let be the product of the two topics’ weights. The corpus-wide share of topic-pair is (equation 1, p. 279):
Categories are identified as the triplets of topics with the highest total share (Incidence) that are sufficiently distinct from each other. The three resulting categories are science (topics 3, 41, 43), religion (topics 10, 34, 38), and political economy (topics 13, 35, 36); see Table I, p. 282.
Each volume is assigned weights for all three categories by weighting the LDA topic weights by each topic’s time-varying category coefficient (equations 2-4, p. 285):
where is the category coefficient of topic for category , computed from the time-varying topic-pair shares over 20-year moving bins. By construction .
Sentiment (progress-oriented score)
Section titled “Sentiment (progress-oriented score)”A progress dictionary (Table II, p. 289) lists 7 modern English synonyms of “progress” (progress, improvement, stride, betterment, advance, rise, amelioration), all in use before 1643 per the Oxford English Dictionary. The progress score for volume is the share of dictionary words in the volume (equation 5, p. 289):
where is the count of word from the progress dictionary in volume , and is the total word count of volume . Scores are converted to percentile ranks over the full corpus for comparability.
Industrial score
Section titled “Industrial score”An industrial score is constructed from the weighted index of machine-related root words transcribed from the five volumes of Appleby’s Illustrated Handbook of Machinery (Appleby 1877-1903). The top 10 industrial words (by index frequency) include: crane (51), electr (42), weight (37), rope (27), cost (27); see Table IV, p. 301. Each volume’s industrial score is the normalized sum of industrial word counts weighted by each word’s Appleby index frequency.
Empirical specifications
Section titled “Empirical specifications”Regression 1: Progress sentiment and language category weights (eq. 6, p. 294)
Section titled “Regression 1: Progress sentiment and language category weights (eq. 6, p. 294)”Volumes are placed into 20-year bins by publication date. The baseline estimating equation is:
where is the progress score (percentile) of volume in bin ; , , are the volume’s category weights from equations (2)-(4); is excluded as the reference category; are 20-year bin fixed effects; and is the vector of all variables and interactions in equation (6), with time-varying slope allowing all coefficients to change across bins. Standard errors are clustered by year of publication. Full results are in Online Appendix Table B.1; predicted values for key language mixes are plotted in Figure VII (p. 295).
Regression 2: Progress sentiment, category weights, and industrial language (eq. 7, p. 303)
Section titled “Regression 2: Progress sentiment, category weights, and industrial language (eq. 7, p. 303)”The industrial score is added as an additional regressor with all two-way and three-way interactions:
where is the normalized industrial language score of volume and all other notation follows equation (6). As before, all slope coefficients are interacted with bin fixed effects to allow time variation. Full results are in Online Appendix Table B.4; predicted values for the 50%/50% science-political economy location at varying industry percentiles are plotted in Figure XI (p. 305).
Both regressions are run on pre-1650 data excluded in robustness checks (Online Appendix Figure B.19, B.27); results are similar. Alternative dictionaries using 1708 Dictionarium Anglo-Britannicum progress words and a ChatGPT-generated Enlightenment-era synonym list also yield similar patterns (Online Appendix Figures B.10-B.13).
Datasets used
Section titled “Datasets used”| Dataset | Role in paper | Wiki page |
|---|---|---|
| HathiTrust Digital Library (HDL) extracted features, 264,443 volumes (Almelhem et al. 2025, Harvard Dataverse) | Main corpus: bag-of-words representation of all English volumes printed in England, 1500-1900; used for LDA and sentiment analysis | no page yet |
| Appleby’s Illustrated Handbook of Machinery, vols. 1-5 (Appleby 1877-1903) | Source of industrial root-word index used to construct per-volume industrial scores | no page yet |
Sample: 264,443 unique English volumes printed in England, 1500-1900, after removing duplicates and non-English volumes from an initial set of 420,081. Data available at Harvard Dataverse: https://doi.org/10.7910/DVN/DQRO8L (Almelhem et al. 2025).
When to read the full paper
Section titled “When to read the full paper”Use the original if you are: replicating the LDA estimation or the sentiment regressions (Online Appendix Sections A-H contain the full data-cleaning protocol, all 60 topic definitions, robustness checks with alternative dictionaries and unbinned data, and author fixed-effect specifications); extending the corpus to other European languages to test the McCloskey (2006) thesis; or using the qualitative volume examples in Section VI (Clare 1735, Saul 1735, Stephenson 1831) to understand what progress-oriented industrial language looked like in practice.
Attribution and rights
Section titled “Attribution and rights”Source: peer-reviewed, The Quarterly Journal of Economics 141(1), 2026. This distillation was extracted by an LLM on 2026-06-28 and is not human-verified or independently reproduced. The CC BY 4.0 licence permits mirroring; the verbatim PDF is not hosted in this batch. Replication data are available at Harvard Dataverse (Almelhem et al. 2025), DOI: 10.7910/DVN/DQRO8L.
Attribution (CC BY 4.0). Almelhem, Ali, Murat Iyigun, Austin Kennedy, and Jared Rubin. “Enlightenment Ideals and Belief in Progress in the Run-up to the Industrial Revolution: A Textual Analysis.” The Quarterly Journal of Economics 141, no. 1 (2026): 263-314. DOI: 10.1093/qje/qjaf054. (c) The Author(s) 2025. Published by Oxford University Press on behalf of President and Fellows of Harvard College. Licensed under CC BY 4.0. This page is an adaptation by the Institute for Automated Research: core results extracted and re-expressed; changes were made.