Data Literacy

Presentation

Motivation Data are everywhere. Using data for good, however, requires more than knowing how to calculate the arithmetic mean of a bunch of numbers or plot a graph. This course develops the conceptual tools needed to understand, evaluate, and communicate quantitative evidence by examining how data are produced, how concepts become measurements, how samples allow us to learn about populations, and how numerical and visual summaries can reveal important patterns or obscure them.

Approach The course is organized around three fundamental ways of learning from data: description, association, and causation. Students begin by describing distributions and comparing groups before asking how variables relate to one another through cross-tabulations, correlations, and regression. This understanding of association provides the foundation for causal reasoning, as students learn why association alone cannot establish causation and use counterfactual thinking and experiments to understand the challenges of drawing causal conclusions from observational data.

Principle Throughout the course, the emphasis is less on performing statistical procedures than on reasoning critically with evidence. Students learn to recognize when comparisons or visualizations distort what data appear to show and when selection problems, confounding, measurement error, or other threats undermine the inferences drawn from them. By the end of the course, students should be better equipped to evaluate empirical claims wherever they encounter them, whether in academic research, public policy, journalism, business, or everyday life.

Objectives

The course builds toward the following learning objectives:

  1. Explain how concepts, observations, variables, populations, samples, and measurement decisions shape the data we analyze.
  2. Select and interpret appropriate numerical and visual summaries for different types of data.
  3. Quantify and interpret the uncertainty involved in using samples to learn about populations.
  4. Compare groups and evaluate associations while distinguishing meaningful patterns from potentially misleading comparisons.
  5. Interpret basic regression results as descriptions of conditional relationships without automatically treating them as causal.
  6. Distinguish descriptive, associational, and causal claims.
  7. Identify common threats to valid inference, including measurement error, selection bias, confounding, reverse causality, missing data, and inappropriate aggregation.
  8. Explain the logic of counterfactual causal inference and why randomization is useful for identifying causal effects.
  9. Critically evaluate quantitative claims in research, policy, journalism, business, and public discourse.

Modules

01—What Are Data?

This module introduces observations, variables, cases, datasets, populations, samples, units of analysis, and data-generating processes. Students examine how data are representations produced through choices about what to observe and record.

Readings

Diez, David M., Christopher D. Barr, and Mine Çetinkaya-Rundel. 2019. OpenIntro Statistics. 4th ed. OpenIntro, Inc. https://www.openintro.org/book/os/. 📑 Chapter 1 (Introduction to Data).

Imai, Kosuke, and Nora Webb Williams. 2022. Quantitative Social Science: An Introduction in Tidyverse. Princeton University Press. 📑 Chapter 1 (Introduction).

Borgman, Christine L. 2015. Big Data, Little Data, No Data: Scholarship in the Networked World. MIT Press. 📑 Chapter 2 (What Are Data?).

Kitchin, Rob. 2014. The Data Revolution: Big Data, Open Data, Data Infrastructures and Their Consequences. SAGE Publications. 📑 Chapter 1 (Conceptualising Data).

D’Ignazio, Catherine, and Lauren F. Klein. 2020. Data Feminism. MIT Press. 📑 Chapter 6 (The Numbers Don’t Speak for Themselves).

02—From Concepts to Measurements

This module examines constructs, operationalization, indicators, scales of measurement, reliability, validity, and measurement error. Students learn to ask how abstract social, political, and economic concepts become empirical variables.

Readings

Adcock, Robert, and David Collier. 2001. “Measurement Validity: A Shared Standard for Qualitative and Quantitative Research.” American Political Science Review 95 (3): 529–46. https://doi.org/10.1017/S0003055401003100.

Imai, Kosuke, and Nora Webb Williams. 2022. Quantitative Social Science: An Introduction in Tidyverse. Princeton University Press. 📑 Chapter 3 (Measurement).

Gelman, Andrew, Jennifer Hill, and Aki Vehtari. 2020. Regression and Other Stories. Cambridge University Press. 📑 Chapter 2 (Data and Measurement).

Stevens, Stanley Smith. 1946. “On the Theory of Scales of Measurement.” Science 103 (2684): 677–80. https://doi.org/10.1126/science.103.2684.677.

Collier, David, Jody LaPorte, and Jason Seawright. 2012. “Putting Typologies to Work: Concept Formation, Measurement, and Analytic Rigor.” Political Research Quarterly 65 (1): 217–32. https://doi.org/10.1177/1065912912437162.

Hand, D. J. 1996. “Statistics and the Theory of Measurement.” Journal of the Royal Statistical Society: Series A (Statistics in Society) 159 (3): 445–73. https://doi.org/10.2307/2983326. ⚠️ Advanced material — Read for the main argument and key concepts.

03—Description: Data Types and Distributions

This module introduces categorical and numerical variables and the logic of describing their distributions. Students examine shape, center, spread, outliers, transformations, and the relationship between variable type and appropriate summary.

Readings

Gerring, John. 2012. “Mere Description.” British Journal of Political Science 42 (4): 721–46. https://doi.org/10.1017/S0007123412000130.

Cairo, Alberto. 2016. The Truthful Art: Data, Charts, and Maps for Communication. New Riders. 📑 Chapter 7 (Visualizing Distributions).

Wilke, Claus O. 2019. Fundamentals of Data Visualization. O’Reilly Media, Incorporated. 📑 Chapter 7 (Visualizing Distributions), Chapter 8 (Empirical Cumulative Distribution Functions and Q-Q Plots), and Chapter 9 (Visualizing Many Distributions at Once).

04—Summarizing Data

This module develops fluency with counts, proportions, percentages, rates, means, medians, quantiles, variance, standard deviation, and standardized quantities. Students learn to choose summaries that preserve the substantive meaning of the data.

Readings

Huntington-Klein, Nick. 2022. The Effect: An Introduction to Research Design and Causality. CRC Press. 📑 Chapter 3 (Describing Variables).

Llaudet, Elena, and Kosuke Imai. 2023. Data Analysis for Social Science: A Friendly and Practical Introduction. Princeton University Press. 📑 Chapter 3 (Inferring Population Characteristics via Survey Research).

Cairo, Alberto. 2016. The Truthful Art: Data, Charts, and Maps for Communication. New Riders. 📑 Chapter 6 (Exploring Data with Simple Charts).

05—Seeing Data

This module focuses on tables, charts, visual encodings, scales, denominators, graphical integrity, accessibility, and common forms of visual distortion. Students learn to read visualizations critically and to recognize design choices that support accurate interpretation.

Readings

Cleveland, William S., and Robert McGill. 1984. “Graphical Perception: Theory, Experimentation, and Application to the Development of Graphical Methods.” Journal of the American Statistical Association 79 (387): 531–54. https://doi.org/10.1080/01621459.1984.10478080.

Cairo, Alberto. 2016. The Truthful Art: Data, Charts, and Maps for Communication. New Riders. 📑 Chapter 5 (Basic Principles of Visualization).

Bergstrom, Carl T., and Jevin D. West. 2020. Calling Bullshit: The Art of Skepticism in a Data-Driven World. Random House. 📑 Chapter 7 (Data Visualization).

Wilke, Claus O. 2019. Fundamentals of Data Visualization. O’Reilly Media, Incorporated. 📑 Chapter 2 (Visualizing Data: Mapping Data onto Aesthetics) and Chapter 17 (The Principle of Proportional Ink).

06—Populations, Samples, and Uncertainty

This module introduces sampling, representativeness, sampling variability, standard errors, confidence intervals, and margins of error. Students ask what a sample can legitimately tell us about a larger population and what uncertainty remains.

Readings

Llaudet, Elena, and Kosuke Imai. 2023. Data Analysis for Social Science: A Friendly and Practical Introduction. Princeton University Press. 📑 Chapter 7 (Quantifying Uncertainty).

Spiegelhalter, David. 2019. The Art of Statistics: How to Learn from Data. Basic Books. 📑 Chapter 3 (Why Are We Looking at Data Anyway? Populations and Measurement) and Chapter 7 (How Sure Can We Be About What Is Going On? Estimates and Intervals).

Wilke, Claus O. 2019. Fundamentals of Data Visualization. O’Reilly Media, Incorporated. 📑 Chapter 16 (Visualizing Uncertainty).

Meng, Xiao-Li. 2018. “Statistical Paradises and Paradoxes in Big Data (i): Law of Large Populations, Big Data Paradox, and the 2016 US Presidential Election.” The Annals of Applied Statistics 12 (2): 685–726. https://doi.org/10.1214/18-AOAS1161SF. ⚠️ Advanced material — Read for the main argument and key concepts.

07—Comparing Groups

This module examines absolute and relative differences, ratios, percentage changes, baselines, standardization, composition, and meaningful group comparisons. Students pay particular attention to the denominator and comparison category underlying an empirical claim.

Readings

Wooldridge, Jeffrey M. 2019. Introductory Econometrics: A Modern Approach. 7th ed. Cengage Learning. 📑 Math Refresher A (Basic Mathematical Tools; A-3 Proportions and Percentages).

Diez, David M., Christopher D. Barr, and Mine Çetinkaya-Rundel. 2019. OpenIntro Statistics. 4th ed. OpenIntro, Inc. https://www.openintro.org/book/os/. 📑 Chapter 2 (Summarizing Data).

08—Association Between Variables

This module introduces cross-tabulations, conditional proportions, covariance, correlation, direction, strength, nonlinearity, and dependence. Students learn what can, and cannot, be concluded when variables move together.

Readings

Huntington-Klein, Nick. 2022. The Effect: An Introduction to Research Design and Causality. CRC Press. 📑 Chapter 4 (Describing Relationships).

Cairo, Alberto. 2016. The Truthful Art: Data, Charts, and Maps for Communication. New Riders. 📑 Chapter 9 (Seeing Relationships).

De Mesquita, Ethan Bueno, and Anthony Fowler. 2021. Thinking Clearly with Data: A Guide to Quantitative Reasoning and Analysis. Princeton University Press. 📑 Chapter 2 (Correlation: What Is It and What Is It Good For?).

Wilke, Claus O. 2019. Fundamentals of Data Visualization. O’Reilly Media, Incorporated. 📑 Chapter 12 (Visualizing Associations Among Two or More Quantitative Variables).

09—Regression

This module introduces bivariate and multivariable regression as tools for describing conditional relationships. Students examine slopes, predictions, controls, residuals, uncertainty, and interpretation without automatically assigning causal meaning.

Readings

Llaudet, Elena, and Kosuke Imai. 2023. Data Analysis for Social Science: A Friendly and Practical Introduction. Princeton University Press. 📑 Chapter 4 (Predicting Outcomes Using Linear Regression).

Imai, Kosuke, and Nora Webb Williams. 2022. Quantitative Social Science: An Introduction in Tidyverse. Princeton University Press. 📑 Chapter 4 (Prediction).

Huntington-Klein, Nick. 2022. The Effect: An Introduction to Research Design and Causality. CRC Press. 📑 Chapter 13 (Regression).

De Mesquita, Ethan Bueno, and Anthony Fowler. 2021. Thinking Clearly with Data: A Guide to Quantitative Reasoning and Analysis. Princeton University Press. 📑 Chapter 6 (Regression for Describing and Forecasting).

Gelman, Andrew, Jennifer Hill, and Aki Vehtari. 2020. Regression and Other Stories. Cambridge University Press. 📑 Chapter 6 (Background on Regression Modeling), Chapter 7 (Linear Regression with a Single Predictor), Chapter 8 (Fitting Regression Models), Chapter 9 (Prediction and Bayesian Inference), Chapter 10 (Linear Regression with Multiple Predictors), and Chapter 11 (Assumptions, Diagnostics, and Model Evaluation). ⚠️ Advanced material — Read for the main argument and key concepts.

Schoon, Eric W, David Melamed, and Ronald L Breiger. 2024. Regression Inside Out. Cambridge University Press. 📑 Chapter 2 (OLS Inside Out), Chapter 3 (Generalizing Regression Inside Out), and Chapter 4 (Turning Variance Inside Out). ⚠️ Advanced material — Read for the main argument and key concepts.

10—Causation: Why Association Is Not Enough

This module examines confounding, omitted variables, reverse causality, selection, collider bias, and spurious relationships. Students learn why even strong and precisely estimated associations may fail to support causal conclusions.

Readings

Bergstrom, Carl T., and Jevin D. West. 2020. Calling Bullshit: The Art of Skepticism in a Data-Driven World. Random House. 📑 Chapter 4 (Causality).

De Mesquita, Ethan Bueno, and Anthony Fowler. 2021. Thinking Clearly with Data: A Guide to Quantitative Reasoning and Analysis. Princeton University Press. 📑 Chapter 10 (Why Correlation Doesn’t Imply Causation).

Huntington-Klein, Nick. 2022. The Effect: An Introduction to Research Design and Causality. CRC Press. 📑 Chapter 5 (Identification).

Gelman, Andrew, and Guido Imbens. 2013. Why Ask Why? Forward Causal Inference and Reverse Causal Questions. National Bureau of Economic Research. ⚠️ Advanced material — Read for the main argument and key concepts.

11—Thinking in Counterfactuals

This module introduces potential outcomes, treatment and control states, treatment effects, the fundamental problem of causal inference, and identification. Students focus on understanding what it would mean for a causal claim to be empirically defensible.

Readings

Cunningham, Scott. 2021. Causal Inference: The Mixtape. Yale University Press. 📑 Chapter 4 (Potential Outcomes Causal Model).

Pearl, Judea, and Dana Mackenzie. 2018. The Book of Why: The New Science of Cause and Effect. Basic Books. 📑 Chapter 8 (Counterfactuals: Mining Worlds That Could Have Been).

Lundberg, Ian, Rebecca Johnson, and Brandon M Stewart. 2021. “What Is Your Estimand? Defining the Target Quantity Connects Statistical Evidence to Theory.” American Sociological Review 86 (3): 532–65. https://doi.org/10.1177/00031224211004187. 🚨 Very advanced material — Read for exposure, not mastery.

12—Experiments and Randomization

This module covers treatment and control groups, random assignment, balance, compliance, attrition, A/B tests, internal validity, and external validity. Students examine why randomization helps solve the problem of causal comparison and where experimental evidence can still fail.

Readings

De Mesquita, Ethan Bueno, and Anthony Fowler. 2021. Thinking Clearly with Data: A Guide to Quantitative Reasoning and Analysis. Princeton University Press. 📑 Chapter 12 (Randomized Experiments).

Llaudet, Elena, and Kosuke Imai. 2023. Data Analysis for Social Science: A Friendly and Practical Introduction. Princeton University Press. 📑 Chapter 2 (Estimating Causal Effects with Randomized Experiments).

Imai, Kosuke, and Nora Webb Williams. 2022. Quantitative Social Science: An Introduction in Tidyverse. Princeton University Press. 📑 Chapter 2 (Causality).

Angrist, Joshua D, and Jörn-Steffen Pischke. 2014. Mastering’metrics: The Path from Cause to Effect. Princeton University Press. 📑 Chapter 1 (Randomized Trials).

Angrist, Joshua D, and Jörn-Steffen Pischke. 2009. Mostly Harmless Econometrics: An Empiricist’s Companion. Princeton university press. 📑 Chapter 2 (The Experimental Ideal). ⚠️ Advanced material — Read for the main argument and key concepts.

Duflo, Esther, Rachel Glennerster, and Michael Kremer. 2007. “Using Randomization in Development Economics Research: A Toolkit.” Handbook of Development Economics 4: 3895–962. https://doi.org/10.1016/S1573-4471(07)04061-2. ⚠️ Advanced material — Read for the main argument and key concepts.

13—Causation from Observational Data

This module introduces the logic behind natural experiments, matching, regression discontinuity, difference-in-differences, and instrumental variables. Students focus not on technical mastery of each estimator but on understanding how observational designs attempt to approximate credible counterfactual comparisons.

Readings

Llaudet, Elena, and Kosuke Imai. 2023. Data Analysis for Social Science: A Friendly and Practical Introduction. Princeton University Press. 📑 Chapter 5 (Estimating Causal Effects with Observational Data).

Huntington-Klein, Nick. 2022. The Effect: An Introduction to Research Design and Causality. CRC Press. 📑 Chapter 6 (Causal Diagrams), Chapter 7 (Drawing Causal Diagrams), and Chapter 8 (Causal Paths and Closing Back Doors).

De Mesquita, Ethan Bueno, and Anthony Fowler. 2021. Thinking Clearly with Data: A Guide to Quantitative Reasoning and Analysis. Princeton University Press. 📑 Chapter 11 (Controlling for Confounders), Chapter 13 (Regression Discontinuity Designs), and Chapter 14 (Difference-in-Differences Designs).

Samii, Cyrus. 2016. “Causal Empiricism in Quantitative Research.” The Journal of Politics 78 (3): 941–55. https://doi.org/10.1086/686690. ⚠️ Advanced material — Read for the main argument and key concepts.

Angrist, Joshua D, and Jörn-Steffen Pischke. 2014. Mastering’metrics: The Path from Cause to Effect. Princeton University Press. 📑 Chapter 3 (Instrumental Variables), Chapter 4 (Regression Discontinuity Designs), and Chapter 5 (Differences-in-Differences). ⚠️ Advanced material — Read for the main argument and key concepts.

Freedman, David A. 1991. “Statistical Models and Shoe Leather.” Sociological Methodology 21: 291–313. https://doi.org/10.2307/270939. ⚠️ Advanced material — Read for the main argument and key concepts.

14—How Data Can Mislead

This module examines missing data, selection bias, survivorship bias, p-hacking, multiple comparisons, bad denominators, aggregation, ecological inference, motivated interpretation, and other pathways from technically correct calculations to misleading conclusions. Students learn to identify these problems and assess how they can distort the interpretation of otherwise technically correct calculations.

Readings

Bergstrom, Carl T., and Jevin D. West. 2020. Calling Bullshit: The Art of Skepticism in a Data-Driven World. Random House. 📑 Chapter 5 (Numbers and Nonsense) and Chapter 6 (Selection Bias).

Reinhart, Alex. 2015. Statistics Done Wrong: The Woefully Complete Guide. No Starch Press. 📑 Chapter 1 (An Introduction to Statistical Significance), Chapter 2 (Statistical Power and Underpowered Statistics), and Chapter 3 (Pseudoreplication: Choose Your Data Wisely).

Cairo, Alberto. 2019. How Charts Lie: Getting Smarter about Visual Information. W. W. Norton & Company. 📑 Chapter 3 (Charts That Lie by Displaying Dubious Data), Chapter 4 (Charts That Lie by Displaying Insufficient Data), Chapter 5 (Charts That Lie by Concealing or Confusing Uncertainty), and Chapter 6 (Charts That Lie by Suggesting Misleading Patterns).

15—Synthesis: Reading Empirical Claims Critically

This module integrates measurement, description, association, causation, uncertainty, and research design. Students practice conducting an evidence audit that ends with a calibrated statement of what an empirical claim is and is not supported by the available evidence.

Readings

Spiegelhalter, David. 2019. The Art of Statistics: How to Learn from Data. Basic Books. 📑 Chapter 10 (Answering Questions and Claiming Discoveries), Chapter 12 (How Things Go Wrong), and Chapter 13 (How We Can Do Statistics Better).

Wasserstein, Ronald L., and Nicole A. Lazar. 2016. “The ASA Statement on p-Values: Context, Process, and Purpose.” The American Statistician 70 (2): 129–33. https://doi.org/10.1080/00031305.2016.1154108.

Ioannidis, John P. A. 2005. “Why Most Published Research Findings Are False.” PLOS Medicine 2 (8): e124. https://doi.org/10.1371/journal.pmed.0020124.

Gelman, Andrew, Jennifer Hill, and Aki Vehtari. 2020. Regression and Other Stories. Cambridge University Press.

Greenland, Sander, Stephen J. Senn, Kenneth J. Rothman, et al. 2016. “Statistical Tests, p Values, Confidence Intervals, and Power: A Guide to Misinterpretations.” European Journal of Epidemiology 31: 337–50. https://doi.org/10.1007/s10654-016-0149-3. ⚠️ Advanced material — Read for the main argument and key concepts.

Evaluation

TBD.