Computational Social Science

Presentation

Motivation The digital transformation of social life has created both new sources of evidence and new ways of studying social phenomena. Text, networks, geographic information, digital traces, administrative records, and other large or complex datasets allow researchers to investigate questions that were previously difficult or impossible to study systematically. This course introduces computational social science as an approach that combines these new forms of data with computational methods and social-scientific reasoning.

Approach The course introduces students to the major research designs and methodological families of computational social science. Students learn how digital data can be collected through web sources and APIs, then examine how text, networks, spatial and temporal data, and simulation can be studied through computational text analysis and natural language processing, machine learning, social network analysis, geographic information systems, and causal inference. Throughout, methods are connected to substantive social-science questions rather than presented simply as computational techniques.

Principle Throughout the course, the emphasis is on the complete journey from social-scientific question to computational evidence. Students learn to judge computational performance alongside measurement, research design, validation, and uncertainty, while recognizing how ethics, privacy, representativeness, and reproducibility shape what can responsibly be learned from computational evidence. By the end of the course, students should have the foundational knowledge needed to engage with the literature, pursue more advanced study, and continue learning independently.

Objectives

The course builds toward the following learning objectives:

  1. Formulate a substantive social-science question as a computational research problem.
  2. Evaluate what digital, administrative, textual, relational, spatial, and behavioral data make possible computationally while recognizing the biases they may introduce.
  3. Acquire and organize research data using APIs, web sources, or other computational collection strategies when appropriate.
  4. Select and apply computational methods appropriate to textual, unsupervised, supervised, relational, spatial, and simulation-based research problems.
  5. Distinguish descriptive, exploratory, predictive, explanatory, and causal research goals and evaluate methods accordingly.
  6. Validate computational measurements and models against appropriate empirical benchmarks and diagnostics.
  7. Interpret social network analysis both as a family of methods and as a relational perspective on social explanation.
  8. Identify and evaluate ethical, privacy, representativeness, and reproducibility concerns in computational research.
  9. Design and communicate an end-to-end computational social science study whose methodological choices are justified by the substantive question.

Modules

01—Computational Social Science

This module introduces the origins, scope, and major debates of computational social science. Students examine digital social science, big data, computational thinking, methodological pluralism, and the relationship between computation, substantive theory, and traditional social-science methods.

Readings

Lazer, David, Alex Pentland, Lada Adamic, et al. 2009. “Computational Social Science.” Science 323 (5915): 721–23. https://doi.org/10.1126/science.1167742.

Edelmann, Achim, Tom Wolff, Danielle Montagne, and Christopher A. Bail. 2020. “Computational Social Science and Sociology.” Annual Review of Sociology 46: 61–81. https://doi.org/10.1146/annurev-soc-121919-054621.

Salganik, Matthew J. 2019. Bit by Bit: Social Research in the Digital Age. Princeton University Press. 📑 Chapter 1 (Introduction).

Lazer, David MJ, Alex Pentland, Duncan J Watts, et al. 2020. “Computational Social Science: Obstacles and Opportunities.” Science 369 (6507): 1060–62. https://doi.org/10.1126/science.aaz8170.

Donoho, David. 2017. “50 Years of Data Science.” Journal of Computational and Graphical Statistics 26 (4): 745–66. https://doi.org/10.1080/10618600.2017.1384734.

02—From Social Questions to Computational Research Designs

This module moves from theory and substantive questions to concepts, units of analysis, measurement, hypotheses, and method choice. Students explicitly distinguish prediction from explanation and ask how computational techniques can be aligned with different scientific goals.

Readings

King, Gary, Robert O Keohane, and Sidney Verba. 2021. Designing Social Inquiry: Scientific Inference in Qualitative Research. Princeton University Press. 📑 Chapter 1 (The Science in Social Science).

Grimmer, Justin. 2015. “We Are All Social Scientists Now: How Big Data, Machine Learning, and Causal Inference Work Together.” PS: Political Science & Politics 48 (1): 80–83. https://doi.org/10.1017/S1049096514001784.

Hofman, Jake M, Duncan J Watts, Susan Athey, et al. 2021. “Integrating Explanation and Prediction in Computational Social Science.” Nature 595 (7866): 181–88. https://doi.org/10.1038/s41586-021-03659-0.

Watts, Duncan J. 2017. “Should Social Science Be More Solution-Oriented?” Nature Human Behaviour 1: 0015. https://doi.org/10.1038/s41562-016-0015.

03—Digital Data and the Computational Research Lifecycle

This module examines digital traces, platforms, administrative records, sensors, transactional data, found data, provenance, representativeness, and data-generating processes. Students learn to ask what population and behavior a digital dataset actually represents.

Readings

Salganik, Matthew J. 2019. Bit by Bit: Social Research in the Digital Age. Princeton University Press. 📑 Chapter 2 (Observing Behavior).

Lazer, David, Ryan Kennedy, Gary King, and Alessandro Vespignani. 2014. “The Parable of Google Flu: Traps in Big Data Analysis.” Science 343 (6176): 1203–5. https://doi.org/10.1126/science.1248506.

Olteanu, Alexandra, Carlos Castillo, Fernando Diaz, and Emre Kıcıman. 2019. “Social Data: Biases, Methodological Pitfalls, and Ethical Boundaries.” Frontiers in Big Data 2: 13. https://doi.org/10.3389/fdata.2019.00013.

King, Gary. 2013. Big Data Is Not about the Data! Presentation at the Golden Seeds Innovation Summit, New York, NY. https://gking.harvard.edu/files/gking/files/evbase-gs.pdf.

04—Collecting Data from the Web

This module introduces APIs, HTTP, scraping, crawling, rate limits, authentication, dynamic content, and responsible acquisition. Students connect technical procedures to terms of service, research ethics, provenance, and the problem of transforming online information into defensible research data.

Readings

Landers, Richard N., Robert C. Brusso, Katelyn J. Cavanaugh, and Andrew B. Collmus. 2016. “A Primer on Theory-Driven Web Scraping: Automatic Extraction of Big Data from the Internet for Use in Psychological Research.” Psychological Methods 21 (4): 475–92. https://doi.org/10.1037/met0000081.

Munzert, Simon, Christian Rubba, Peter Meißner, and Dominic Nyhuis. 2015. Automated Data Collection with r: A Practical Guide to Web Scraping and Text Mining. Wiley. 📑 Chapter 9 (Scraping the Web).

Luscombe, Alex, Kevin Dick, and Kevin Walby. 2022. “Algorithmic Thinking in the Public Interest: Navigating Technical, Legal, and Ethical Hurdles to Web Scraping in the Social Sciences.” Quality & Quantity 56 (3): 1023–44. https://doi.org/10.1007/s11135-021-01164-0.

05—Preparing Computational Social Data

This module covers cleaning, parsing, reshaping, entity resolution, record linkage, missingness, metadata, reproducible preprocessing, and computational measurement. Students examine the transformations that stand between raw digital traces and analyzable social-science data.

Readings

Atteveldt, Wouter van, Damian Trilling, and Carlos Arcíla Calderón. 2022. Computational Analysis of Communication. Wiley Blackwell. 📑 Chapter 5 (Files and Data Frames) and Chapter 6 (Data Wrangling).

Cirone, Alexandra, and Arthur Spirling. 2021. “Turning History into Data: Data Collection, Measurement, and Inference in HPE.” Journal of Historical Political Economy 1 (1): 127–54.

Enamorado, Ted, Benjamin Fifield, and Kosuke Imai. 2019. “Using a Probabilistic Model to Assist Merging of Large-Scale Administrative Records.” American Political Science Review 113 (2): 353–71. https://doi.org/10.1017/S0003055418000783.

Lazer, David, Eszter Hargittai, Deen Freelon, et al. 2021. “Meaningful Measures of Human Society in the Twenty-First Century.” Nature 595 (7866): 189–96. https://doi.org/10.1038/s41586-021-03660-7.

Bail, Christopher A. 2012. “The Fringe Effect: Civil Society Organizations and the Evolution of Media Discourse about Islam Since the September 11th Attacks.” American Sociological Review 77 (6): 855–79.

06—Text as Data

This module treats documents as observations and introduces corpora, tokenization, dictionaries, features, document-term representations, annotation, measurement, and validation. Students examine how language can be converted into evidence without losing sight of the construct being measured.

Readings

Grimmer, Justin, Margaret E Roberts, and Brandon M Stewart. 2022. Text as Data: A New Framework for Machine Learning and the Social Sciences. Princeton University Press. 📑 Chapter 1 (Introduction) and Chapter 2 (Social Science Research and Text Analysis).

Grimmer, Justin, and Brandon M. Stewart. 2013. “Text as Data: The Promise and Pitfalls of Automatic Content Analysis Methods for Political Texts.” Political Analysis 21 (3): 267–97. https://doi.org/10.1093/pan/mps028.

Gentzkow, Matthew, Bryan Kelly, and Matt Taddy. 2019. “Text as Data.” Journal of Economic Literature 57 (3): 535–74. https://doi.org/10.1257/jel.20181020.

07—Natural Language Processing

This module introduces classification, information extraction, named entities, embeddings, semantic similarity, and language models. Students learn the distinction between general NLP capabilities and social-science applications in which task validity and substantive interpretation are central.

Readings

Jurafsky, Daniel, and James H. Martin. 2026. Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models. 3rd ed. https://web.stanford.edu/~jurafsky/slp3/. 📑 Chapter 2 (Words and Tokens), Chapter 3 (N-gram Language Models), and Chapter 4 (Logistic Regression and Text Classification). ⚠️ Advanced material — Read for the main argument and key concepts.

Eisenstein, Jacob. 2019. Introduction to Natural Language Processing. MIT Press. 📑 Chapter 1 (Introduction), Chapter 2 (Linear Text Classification), Chapter 8 (Sequence Labeling), Chapter 14 (Distributional and Distributed Semantics), and Chapter 17 (Information Extraction). ⚠️ Advanced material — Read for the main argument and key concepts.

Manning, Christopher D., Prabhakar Raghavan, and Hinrich Schütze. 2008. Introduction to Information Retrieval. Cambridge University Press. 📑 Chapter 6 (Scoring, Term Weighting and the Vector Space Model). ⚠️ Advanced material — Read for the main argument and key concepts.

08—Machine Learning for Social Research

This module introduces supervised learning, train/test separation, cross-validation, regularization, feature engineering, performance metrics, interpretation, and prediction. Students examine the relationship and tension between predictive performance and social-scientific explanation.

Readings

Molina, Mario, and Filiz Garip. 2019. “Machine Learning for Sociology.” Annual Review of Sociology 45 (1): 27–45. https://doi.org/10.1146/annurev-soc-073117-041106.

James, Gareth, Daniela Witten, Trevor Hastie, Robert Tibshirani, and Jonathan Taylor. 2023. An Introduction to Statistical Learning: With Applications in Python. Springer. 📑 Chapter 2 (Statistical Learning).

Mullainathan, Sendhil, and Jann Spiess. 2017. “Machine Learning: An Applied Econometric Approach.” Journal of Economic Perspectives 31 (2): 87–106. https://doi.org/10.1257/jep.31.2.87.

Ward, Michael D, Brian D Greenhill, and Kristin M Bakke. 2010. “The Perils of Policy by p-Value: Predicting Civil Conflicts.” Journal of Peace Research 47 (4): 363–75. https://doi.org/10.1177/0022343309356491.

Yarkoni, Tal, and Jacob Westfall. 2017. “Choosing Prediction over Explanation in Psychology: Lessons from Machine Learning.” Perspectives on Psychological Science 12 (6): 1100–1122. https://doi.org/10.1177/1745691617693393.

Rogers, Simon, and Mark Girolami. 2016. A First Course in Machine Learning. CRC press. 📑 Chapter 1 (Linear Modelling: A Least Squares Approach). ⚠️ Advanced material — Read for the main argument and key concepts.

09—Unsupervised Learning and Discovery

This module examines clustering, topic models, dimensionality reduction, latent structure, exploratory analysis, stability, and interpretability. Students learn how algorithms can reveal patterns while recognizing that discovered structure does not explain itself.

Readings

James, Gareth, Daniela Witten, Trevor Hastie, Robert Tibshirani, and Jonathan Taylor. 2023. An Introduction to Statistical Learning: With Applications in Python. Springer. 📑 Chapter 12 (Unsupervised Learning).

Grimmer, Justin, Margaret E Roberts, and Brandon M Stewart. 2022. Text as Data: A New Framework for Machine Learning and the Social Sciences. Princeton University Press. 📑 Chapter 10 (Principles of Discovery) and Chapter 13 (Topic Models).

Denny, Matthew J, and Arthur Spirling. 2018. “Text Preprocessing for Unsupervised Learning: Why It Matters, When It Misleads, and What to Do about It.” Political Analysis 26 (2): 168–89. https://doi.org/10.1017/pan.2017.44.

Roberts, Margaret E., Brandon M. Stewart, Dustin Tingley, et al. 2014. “Structural Topic Models for Open-Ended Survey Responses.” American Journal of Political Science 58 (4): 1064–82. https://doi.org/10.1111/ajps.12103.

Nelson, Laura K. 2020. “Computational Grounded Theory: A Methodological Framework.” Sociological Methods & Research 49 (1): 3–42. https://doi.org/10.1177/0049124117729703.

10—Social Network Analysis

This module introduces nodes, ties, centrality, homophily, communities, brokerage, diffusion, and network dependence. Students also examine a central conceptual debate: whether social network analysis is best understood as a collection of methods or as a theoretical and relational framework for explaining social phenomena.

Readings

Borgatti, Stephen P., Ajay Mehra, Daniel J. Brass, and Giuseppe Labianca. 2009. “Network Analysis in the Social Sciences.” Science 323 (5916): 892–95. https://doi.org/10.1126/science.1165821.

Borgatti, Stephen P., Martin G. Everett, and Jeffrey C. Johnson. 2018. Analyzing Social Networks. 2nd ed. SAGE. 📑 Chapter 1 (Introduction), Chapter 9 (Characterizing Whole Networks), and Chapter 10 (Centrality).

Zufall, Elise, and Tyler A Scott. 2024. “Syntactic Measurement of Governance Networks from Textual Data, with Application to Water Management Plans.” Policy Studies Journal 52 (4): 941–54. https://doi.org/10.1111/psj.12556.

Henry, Adam Douglas. 2011. “Ideology, Power, and the Structure of Policy Networks.” Policy Studies Journal 39 (3): 361–83.

Granovetter, Mark. 1985. “Economic Action and Social Structure: The Problem of Embeddedness.” American Journal of Sociology 91 (3): 481–510. https://doi.org/10.1086/228311.

Jackson, Matthew O. 2008. Social and Economic Networks. Princeton University Press. 📑 Chapter 2 (Representing and Measuring Networks), Chapter 3 (Empirical Background on Social and Economic Networks), and Chapter 7 (Diffusion through Networks). ⚠️ Advanced material — Read for the main argument and key concepts.

11—Spatial Data and Geographic Analysis

This module introduces geographic information systems alongside broader spatial and temporal analysis. Students examine spatial data models, layers, coordinate systems, spatial joins, geoprocessing, geographic units, spatial dependence, trajectories, sequences, panels, event data, temporal aggregation, and dynamics.

Readings

Lovelace, Robin, Jakub Nowosad, and Jannes Muenchow. 2019. Geocomputation with r. Chapman; Hall/CRC. 📑 Chapter 2 (Geographic Data in R) and Chapter 4 (Spatial Data Operations).

O’Sullivan, David, and David J. Unwin. 2010. Geographic Information Analysis. 2nd ed. Wiley. 📑 Chapter 1 (Geographic Information Analysis and Spatial Data) and Chapter 7 (Area Objects and Spatial Autocorrelation).

Wang, Fahui, and Lingbo Liu. 2023. Computational Methods and GIS Applications in Social Science. CRC Press. 📑 Chapter 1 (Getting Started with ArcGIS: Data Management and Basic Spatial Analysis Tools), Chapter 8 (Spatial Statistics and Applications), and Chapter 14 (Spatiotemporal Big Data Analytics and Application in Urban Studies).

Glaeser, Edward L, Scott Duke Kominers, Michael Luca, and Nikhil Naik. 2018. “Big Data and Big Cities: The Promises and Limitations of Improved Measures of Urban Life.” Economic Inquiry 56 (1): 114–37.

Keele, Luke J, and Rocio Titiunik. 2015. “Geographic Boundaries as Regression Discontinuities.” Political Analysis 23 (1): 127–55. https://doi.org/10.1093/pan/mpu014.

Fotheringham, A. Stewart, and David W. S. Wong. 1991. “The Modifiable Areal Unit Problem in Multivariate Statistical Analysis.” Environment and Planning A 23 (7): 1025–44. https://doi.org/10.1068/a231025. 🚨 Very advanced material — Read for exposure, not mastery.

12—Simulation and Agent-Based Modeling

This module introduces simulation as a way to study processes that are difficult to analyze directly and, for pedagogical convenience, considers Monte Carlo methods alongside the distinct approach of agent-based modeling. Students examine interaction and emergent social patterns in agent-based models alongside calibration, sensitivity analysis, model interpretation, and the broader logic of learning from computational worlds constructed by researchers.

Readings

Smaldino, Paul. 2023. Modeling Social Behavior: Mathematical and Agent-Based Models of Social Dynamics and Cultural Evolution. Princeton University Press. 📑 Chapter 1 (Doing Violence to Reality), Chapter 3 (The Schelling Chapter), and Chapter 10 (Models and Reality).

Gelman, Andrew, Jennifer Hill, and Aki Vehtari. 2020. Regression and Other Stories. Cambridge University Press. 📑 Chapter 5 (Simulation).

Railsback, Steven F., and Volker Grimm. 2019. Agent-Based and Individual-Based Modeling: A Practical Introduction. 2nd ed. Princeton University Press. 📑 Chapter 1 (Models, Agent-Based Models, and the Modeling Cycle), Chapter 2 (Getting Started with NetLogo), Chapter 3 (Describing and Formulating ABMs: The ODD Protocol), Chapter 4 (Implementing a First Agent-Based Model), Chapter 5 (From Animations to Science), and Chapter 6 (Testing Your Program).

Epstein, Joshua M. 2008. “Why Model?” Journal of Artificial Societies and Social Simulation 11 (4): 12. https://www.jasss.org/11/4/12.html.

Macy, Michael W., and Robert Willer. 2002. “From Factors to Actors: Computational Sociology and Agent-Based Modeling.” Annual Review of Sociology 28: 143–66. https://doi.org/10.1146/annurev.soc.28.110601.141117.

Huang, Yue, Zhengqing Yuan, Yujun Zhou, et al. 2024. “Social Science Meets Llms: How Reliable Are Large Language Models in Social Simulations?” arXiv Preprint arXiv:2410.23426. ⚠️ Advanced material — Read for the main argument and key concepts.

13—Causal Inference with Computational Data

This module connects potential outcomes, experiments, observational identification, high-dimensional controls, computational measurement, and treatment-effect heterogeneity. Students examine why predictive sophistication cannot substitute for a credible identification strategy.

Readings

Samii, Cyrus. 2016. “Causal Empiricism in Quantitative Research.” The Journal of Politics 78 (3): 941–55. https://doi.org/10.1086/686690.

Ma, Jing. 2025. “Causal Inference with Large Language Model: A Survey.” Findings of the Association for Computational Linguistics: NAACL 2025, 5901–13.

Egami, Naoki, Christian J. Fong, Justin Grimmer, Margaret E. Roberts, and Brandon M. Stewart. 2022. “How to Make Causal Inferences Using Texts.” Science Advances 8 (42): eabg2652. https://doi.org/10.1126/sciadv.abg2652. ⚠️ Advanced material — Read for the main argument and key concepts.

Athey, Susan, and Guido W Imbens. 2019. “Machine Learning Methods That Economists Should Know About.” Annual Review of Economics 11 (1): 685–725. https://doi.org/10.1146/annurev-economics-080217-053433. ⚠️ Advanced material — Read for the main argument and key concepts.

Imai, Kosuke, and Kentaro Nakamura. 2026. “Causal Inference with Generative Artificial Intelligence: Application to Texts as Treatments.” Journal of the American Statistical Association, nos. just-accepted: 1–27. 🚨 Very advanced material — Read for exposure, not mastery.

14—Ethics, Privacy, Bias, and Responsible Computational Research

This module addresses consent, public versus private data, re-identification, algorithmic bias, platform populations, vulnerable communities, documentation, and responsible disclosure. Students evaluate not only whether data can be collected and analyzed, but whether they should be.

Readings

Salganik, Matthew J. 2019. Bit by Bit: Social Research in the Digital Age. Princeton University Press. 📑 Chapter 6 (Ethics).

Zook, Matthew, Solon Barocas, danah boyd, et al. 2017. “Ten Simple Rules for Responsible Big Data Research.” PLOS Computational Biology 13 (3): e1005399. https://doi.org/10.1371/journal.pcbi.1005399.

boyd, danah, and Kate Crawford. 2012. “Critical Questions for Big Data.” Information, Communication & Society 15 (5): 662–79. https://doi.org/10.1080/1369118X.2012.678878.

Buolamwini, Joy, and Timnit Gebru. 2018. “Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification.” Proceedings of the 1st Conference on Fairness, Accountability and Transparency, Proceedings of machine learning research, vol. 81: 77–91. https://proceedings.mlr.press/v81/buolamwini18a.html.

15—Synthesis: Evaluating an End-to-End Computational Social Science Study

This module integrates substantive theory, data generation, computational methods, validation, inference, ethics, reproducibility, and communication. Students evaluate how these elements work together in an end-to-end computational social science study and whether the resulting pipeline supports credible social-scientific inference.

Readings

Bond, Robert M., Christopher J. Fariss, Jason J. Jones, et al. 2012. “A 61-Million-Person Experiment in Social Influence and Political Mobilization.” Nature 489: 295–98. https://doi.org/10.1038/nature11421.

Tucker, Joshua A., Andrew Guess, Pablo Barberá, et al. 2018. Social Media, Political Polarization, and Political Disinformation: A Review of the Scientific Literature. William; Flora Hewlett Foundation. https://doi.org/10.2139/ssrn.3144139.

Card, Dallas, Serina Chang, Chris Becker, et al. 2022. “Computational Analysis of 140 Years of US Political Speeches Reveals More Positive but Increasingly Polarized Framing of Immigration.” Proceedings of the National Academy of Sciences 119 (31): e2120510119. https://doi.org/10.1073/pnas.2120510119.

Goldberg, Amir, and Sarah K Stein. 2018. “Beyond Social Contagion: Associative Diffusion and the Emergence of Cultural Variation.” American Sociological Review 83 (5): 897–932. https://doi.org/10.1177/0003122418797576.

Evaluation

TBD.