Generative AI for Applied Research
Presentation
Motivation Generative AI is rapidly changing how researchers work with text, documents, data, and computational systems. Large language models can classify and extract information, assist with coding, generate structured data, retrieve and synthesize evidence, and participate in complex research workflows, but their apparent fluency can also obscure problems of reliability and validity. This course examines how researchers can use these systems while confronting hallucination, bias, provenance, and reproducibility.
Approach The course approaches generative AI not simply as a productivity tool, but as a research technology that must be designed, evaluated, and validated. Students first learn how generative models developed from earlier traditions in Artificial Intelligence, how contemporary language models work, and how open-weight and proprietary model ecosystems differ, then translate that understanding into research-grade prompts and formal representations of research tasks. Those representations support structured outputs, API workflows, annotation systems, embeddings, retrieval systems, and human-in-the-loop pipelines.
Principle Throughout the course, the central question is not whether an AI system can produce a plausible answer, but whether researchers can justify trusting the resulting evidence. Students learn to preserve provenance, construct validation datasets, compare AI outputs against appropriate benchmarks, analyze model failures, integrate human judgment, and communicate limitations transparently. By the end of the course, students should be able to design an auditable and reproducible generative-AI research pipeline whose outputs can be evaluated according to the standards of empirical social science.
Objectives
The course builds toward the following learning objectives:
- Explain the concepts and historical developments needed to situate contemporary generative AI within Artificial Intelligence and computational social science.
- Explain how contemporary generative language models work and distinguish between important features of open-weight and proprietary model ecosystems.
- Design prompts, knowledge representations, ontologies, schemas, and task protocols that translate social-science concepts into reproducible model operations.
- Use model APIs to construct parameterized research workflows that can be scaled and reproduced.
- Apply generative models to annotation, classification, extraction, coding, retrieval, and other social-science measurement tasks.
- Design human-in-the-loop procedures that assign appropriate responsibilities to researchers, annotators, and models.
- Evaluate generative-AI outputs against appropriate benchmarks using gold standards, reliability and validity criteria, error analysis, and sensitivity tests.
- Diagnose hallucination, instability, bias, contamination, and other model failures and assess their implications for substantive conclusions.
- Build and document an auditable AI-assisted research pipeline that preserves provenance and version history, incorporates appropriate ethical safeguards, and communicates its limitations transparently.
Modules
01—From Artificial Intelligence to Generative AI
This module clarifies what is meant by Artificial Intelligence and provides a concise historical genealogy from symbolic and statistical approaches to machine learning, deep learning, foundation models, and contemporary generative systems. Students use this history to situate generative AI as a research technology rather than as an isolated recent product category.
Readings
Russell, Stuart J., and Peter Norvig. 2020. Artificial Intelligence: A Modern Approach. 4th ed. Pearson. 📑 Chapter 1 (Introduction), Chapter 27 (Philosophy, Ethics, and Safety of AI), and Chapter 28 (The Future of AI).
Bommasani, Rishi et al. 2021. On the Opportunities and Risks of Foundation Models. https://arxiv.org/abs/2108.07258.
Ziems, Caleb, William Held, Omar Shaikh, Jiaao Chen, Zhehao Zhang, and Diyi Yang. 2024. “Can Large Language Models Transform Computational Social Science?” Computational Linguistics 50 (1): 237–91. https://doi.org/10.1162/coli_a_00502.
Bail, Christopher A. 2024. “Can Generative AI Improve Social Science?” Proceedings of the National Academy of Sciences 121 (21): e2314021121. https://doi.org/10.1073/pnas.2314021121.
02—How Generative Language Models Work
This module introduces tokens, embeddings, transformers, attention, pretraining, instruction tuning, inference, context windows, and stochastic generation. Students also compare open-weight and proprietary model ecosystems and examine how access, transparency, privacy, reproducibility, local deployment, and model control differ across them.
Readings
Alammar, Jay, and Maarten Grootendorst. 2024. Hands-on Large Language Models. O’Reilly Media. 📑 Chapter 1 (Introduction to Language Models) and Chapter 3 (Looking Inside Transformer LLMs).
Brown, Tom B., Benjamin Mann, Nick Ryder, et al. 2020. “Language Models Are Few-Shot Learners.” Advances in Neural Information Processing Systems 33: 1877–901. https://proceedings.neurips.cc/paper/2020/hash/1457c0d6bfcb4967418bfb8ac142f64a-Abstract.html.
Jurafsky, Daniel, and James H. Martin. 2026. Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models. 3rd ed. https://web.stanford.edu/~jurafsky/slp3/. 📑 Chapter 1 (Introduction), Chapter 7 (Transformers and Pretraining), and Chapter 8 (Post-training). ⚠️ Advanced material — Read for the main argument and key concepts.
Dong, Qingxiu, Lei Li, Damai Dai, et al. 2024. “A Survey on in-Context Learning.” Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 1107–28. ⚠️ Advanced material — Read for the main argument and key concepts.
Vaswani, Ashish, Noam Shazeer, Niki Parmar, et al. 2017. “Attention Is All You Need.” Advances in Neural Information Processing Systems 30: 5998–6008. https://proceedings.neurips.cc/paper/7181-attention-is-all-you-need. 🚨 Very advanced material — Read for exposure, not mastery.
03—Prompting as Research Design
This module treats prompts as methodological instruments. Students examine instructions, context, examples, decomposition, prompt templates, task specifications, robustness, and reproducible prompt protocols, while treating contemporary agentic features, reusable skills, and tool-use patterns as changing implementations of these more durable design principles.
Readings
Phoenix, James, and Mike Taylor. 2024. Prompt Engineering for Generative AI. " O’Reilly Media, Inc.". 📑 Chapter 1 (The Five Principles of Prompting).
Schulhoff, Sander, Michael Ilie, Nishant Balepur, et al. 2024. The Prompt Report: A Systematic Survey of Prompting Techniques. https://doi.org/10.48550/arXiv.2406.06608.
Abraham, Louis, Charles Arnal, and Antoine Marie. 2025. “Prompt Selection Matters: Enhancing Text Annotations for Social Sciences with Large Language Models.” Journal of Computational Social Science 8: 73. https://doi.org/10.1007/s42001-025-00388-6.
Atreja, Shubham, Joshua Ashkinaze, Lingyao Li, Julia Mendelsohn, and Libby Hemphill. 2025. “What’s in a Prompt?: A Large-Scale Experiment to Assess the Impact of Prompt Design on the Compliance and Accuracy of LLM-Generated Text Annotations.” Proceedings of the International AAAI Conference on Web and Social Media 19 (1). https://doi.org/10.1609/icwsm.v19i1.35807.
04—Knowledge Representation, Ontology Engineering, and Structured Extraction
This module introduces knowledge representation as the conceptual foundation for schemas, taxonomies, ontologies, knowledge graphs, entities, relations, and propositions. Students learn basic ontology-engineering principles of purpose, concept definition, classes, relations, constraints, iteration, and validation and then use those representations to constrain generative systems into structured, analyzable outputs with evidence and provenance.
Readings
Noy, Natalya F., and Deborah L. McGuinness. 2001. Ontology Development 101: A Guide to Creating Your First Ontology. Stanford Knowledge Systems Laboratory.
Hogan, Aidan, Eva Blomqvist, Michael Cochez, et al. 2021. “Knowledge Graphs.” ACM Computing Surveys (Csur) 54 (4): 1–37. https://doi.org/10.1145/3447772.
Tudorache, Tania. 2020. “Ontology Engineering: Current State, Challenges, and Future Directions.” Semantic Web 11 (1): 125–38. https://doi.org/10.3233/SW-190382.
Shimizu, Cogan, and Pascal Hitzler. 2025. “Accelerating Knowledge Graph and Ontology Engineering with Large Language Models.” Journal of Web Semantics 85: 100862. https://doi.org/10.1016/j.websem.2025.100862.
McVeety, Sam, and Amir Hormati. 2026. “Introducing the Open Knowledge Format.” Google Cloud, June 12. https://cloud.google.com/blog/products/data-analytics/how-the-open-knowledge-format-can-improve-data-sharing.
Narayanan, Pavan Kumar. 2024. “Getting Started with Data Validation Using Pydantic and Pandera.” In Data Engineering for Machine Learning Pipelines: From Python Libraries to ML Pipelines and Cloud Platforms. Springer.
Tsaneva, Stefani, Danilo Dessì, Francesco Osborne, and Marta Sabou. 2025. “Knowledge Graph Validation by Integrating LLMs and Human-in-the-Loop.” Information Processing & Management 62 (5): 104145. https://doi.org/10.1016/j.ipm.2025.104145.
Brach, William, Kristián Košťál, and Michal Ries. 2025. “The Effectiveness of Large Language Models in Transforming Unstructured Text to Standardized Formats.” IEEE Access 13: 91808–25. https://doi.org/10.1109/ACCESS.2025.3573030.
Polak, Maciej P., and Dane Morgan. 2024. “Extracting Accurate Materials Data from Research Papers with Conversational Language Models and Prompt Engineering.” Nature Communications 15: 1569. https://doi.org/10.1038/s41467-024-45914-8.
Xie, Zhengnan, Alice Saebom Kwak, Enfa Fane, et al. 2022. “Extracting Space Situational Awareness Events from News Text.” Proceedings of the Thirteenth Language Resources and Evaluation Conference, 6077–82.
05—Using LLMs through APIs
This module moves from conversational interfaces to research infrastructure. Students examine authentication, requests, parameters, model selection, open versus proprietary access patterns, batching, rate limits, retries, cost accounting, caching, logging, and reproducible execution, while treating tool calling and related agentic capabilities as optional, fast-changing extensions of API-based workflow design.
Readings
OpenAI. 2026. Developer Quickstart. Https://developers.openai.com/api/docs/quickstart.
OpenAI. n.d. “Structured Outputs.” Accessed September 2, 2026. https://developers.openai.com/api/docs/guides/structured-outputs.
OpenAI. n.d. “Function Calling and Tools.” Accessed September 2, 2026. https://developers.openai.com/api/docs/guides/function-calling.
06—LLMs for Annotation and Measurement
This module examines classification, coding, entity and relation extraction, construct measurement, codebooks, zero- and few-shot annotation, scale, and measurement validity. Students ask under what conditions model outputs can function as social-science measurements rather than merely plausible text.
Readings
Chae, Youngjin, and Thomas Davidson. 2026. “Large Language Models for Text Classification: From Zero-Shot Learning to Instruction-Tuning.” Sociological Methods & Research 55 (2): 501–67.
Gilardi, Fabrizio, Meysam Alizadeh, and Maël Kubli. 2023. “ChatGPT Outperforms Crowd Workers for Text-Annotation Tasks.” Proceedings of the National Academy of Sciences 120 (30): e2305016120. https://doi.org/10.1073/pnas.2305016120.
Törnberg, Petter. 2025. “Large Language Models Outperform Expert Coders and Supervised Classifiers at Annotating Political Social Media Messages.” Social Science Computer Review 43 (6). https://doi.org/10.1177/08944393241286471.
Pangakis, Nick, and Sam Wolken. 2025. “Keeping Humans in the Loop: Human-Centered Automated Annotation with Generative AI.” Proceedings of the International AAAI Conference on Web and Social Media 19: 1471–92. https://doi.org/10.1609/icwsm.v19i1.35883.
Licht, Hauke, Rupak Sarkar, Patrick Y. Wu, et al. 2025. “Measuring Scalar Constructs in Social Science with LLMs.” Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 32144–71.
Ashwin, Julian, Aditya Chhabra, and Vijayendra Rao. 2026. “Using Large Language Models for Qualitative Analysis Can Introduce Serious Bias.” Sociological Methods & Research 55 (3): 795–839. https://doi.org/10.1177/00491241251338246.
07—Human–AI Research Workflows
This module covers expert annotation, crowd annotation, adjudication, active review, escalation rules, disagreement, and the division of labor among researchers, annotators, and models. Students design human-in-the-loop systems in which consequential judgments remain explicit and auditable.
Readings
Schroeder, Hope, Deb Roy, and Jad Kabbara. 2025. “Just Put a Human in the Loop? Investigating LLM-Assisted Annotation for Subjective Tasks.” Findings of the Association for Computational Linguistics: ACL 2025, 25771–95. https://doi.org/10.18653/v1/2025.findings-acl.1323.
Zhao, Chengshuai, Zhen Tan, Chau-Wai Wong, Xinyan Zhao, Tianlong Chen, and Huan Liu. 2025. “SCALE: Towards Collaborative Content Analysis in Social Science with Large Language Model Agents and Human Intervention.” Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 8473–503. https://doi.org/10.18653/v1/2025.acl-long.416.
Dunivin, Zackary Okun, Mobina Noori, Seth Frey, and Curtis Atkinson. 2026. Self-Reflection in Automated Qualitative Coding: Improving Text Annotation Through Secondary LLM Critique. https://arxiv.org/abs/2601.09905.
Wang, Xinru, Hannah Kim, Sajjadur Rahman, Kushan Mitra, and Zhengjie Miao. 2024. “Human-LLM Collaborative Annotation Through Effective Verification of LLM Labels.” Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, 1–21.
Kim, Hannah, Kushan Mitra, Rafael Li Chen, Sajjadur Rahman, and Dan Zhang. 2024. “Meganno+: A Human-Llm Collaborative Annotation System.” Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, 168–76.
Schroeder, Hope, Marianne Aubin Le Quéré, Casey Randazzo, David Mimno, and Sarita Schoenebeck. 2025. “Large Language Models in Qualitative Research: Uses, Tensions, and Intentions.” Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 1–17. https://doi.org/10.1145/3706598.3713120.
Abramson, Corey M. et al. 2026. “Qualitative Research in an Era of Artificial Intelligence: A Pragmatic Approach to Data Analysis, Workflow, and Computation.” Annual Review of Sociology 52. https://doi.org/10.1146/annurev-soc-011824-104836.
Bisbee, James, and Arthur Spirling. 2025. “What to Do When Humans Are No Longer the Gold Standard: Large Language Models, State of the Art and Robustness.” Unpublished manuscript.
Heseltine, Michael, and Bernhard Clemm von Hohenberg. 2024. “Large Language Models as a Substitute for Human Experts in Annotating Political Text.” Research & Politics 11 (1): 20531680241236239.
Veselovsky, Veniamin, Manoel Horta Ribeiro, Akhil Arora, Martin Josifoski, Ashton Anderson, and Robert West. 2023. “Generating Faithful Synthetic Data with Large Language Models: A Case Study in Computational Social Science.” arXiv Preprint arXiv:2305.15041.
Zupic, Ivan. 2026. “A Human-Centered Workflow for Using Large Language Models in Content Analysis.” arXiv Preprint arXiv:2603.19271.
08—Embeddings and Semantic Representations
This module introduces vector representations, semantic similarity, nearest neighbors, clustering, classification, retrieval, construct representation, and validation. Students consider what it means to represent social meaning geometrically and how such representations can succeed or fail as measurements.
Readings
Alammar, Jay, and Maarten Grootendorst. 2024. Hands-on Large Language Models. O’Reilly Media. 📑 Chapter 2 (Tokens and Embeddings), Chapter 5 (Text Clustering and Topic Modeling), and Chapter 10 (Creating Text Embedding Models).
Rodriguez, Pedro L, and Arthur Spirling. 2022. “Word Embeddings: What Works, What Doesn’t, and How to Tell the Difference for Applied Research.” The Journal of Politics 84 (1): 101–15.
Jurafsky, Daniel, and James H. Martin. 2026. Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models. 3rd ed. https://web.stanford.edu/~jurafsky/slp3/. 📑 Chapter 5 (Embeddings).
Reimers, Nils, and Iryna Gurevych. 2019. “Sentence-BERT: Sentence Embeddings Using Siamese BERT-Networks.” Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, 3982–92. https://doi.org/10.18653/v1/D19-1410. ⚠️ Advanced material — Read for the main argument and key concepts.
Wulff, Dirk U., and Rui Mata. 2025. “Semantic Embeddings Reveal and Address Taxonomic Incommensurability in Psychological Measurement.” Nature Human Behaviour 9: 944–54. https://doi.org/10.1038/s41562-024-02089-y. ⚠️ Advanced material — Read for the main argument and key concepts.
09—Retrieval-Augmented Generation
This module introduces chunking, indexing, retrieval, grounding, context assembly, citations, document collections, and knowledge bases. Students learn to evaluate retrieval separately from generation and to distinguish answers grounded in a defined corpus from outputs based primarily on model training.
Readings
Alammar, Jay, and Maarten Grootendorst. 2024. Hands-on Large Language Models. O’Reilly Media. 📑 Chapter 8 (Semantic Search and Retrieval-Augmented Generation).
Jurafsky, Daniel, and James H. Martin. 2026. Speech and Language Processing: An Introduction to Natural Language Processing, Computational Linguistics, and Speech Recognition with Language Models. 3rd ed. https://web.stanford.edu/~jurafsky/slp3/. 📑 Chapter 11 (Information Retrieval and Retrieval-Augmented Generation).
Es, Shahul, Jithin James, Luis Espinosa-Anke, and Steven Schockaert. 2024. “RAGAs: Automated Evaluation of Retrieval Augmented Generation.” Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations, 150–58. https://doi.org/10.18653/v1/2024.eacl-demo.16.
Lewis, Patrick, Ethan Perez, Aleksandra Piktus, et al. 2020. “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks.” Advances in Neural Information Processing Systems 33: 9459–74. 🚨 Very advanced material — Read for exposure, not mastery.
10—Evaluating Generative AI
This module covers gold standards, inter-rater comparison, precision, recall, F1, calibration, construct validity, test sets, benchmark design, ablation, robustness, and task-specific evaluation. Students frame evaluation around the evidence required to justify using AI-assisted outputs in empirical research.
Readings
Abdurahman, Suhaib, Alireza Salkhordeh Ziabari, Alexander K. Moore, Daniel M. Bartels, and Morteza Dehghani. 2025. “A Primer for Evaluating Large Language Models in Social-Science Research.” Advances in Methods and Practices in Psychological Science 8 (2): 1–25. https://doi.org/10.1177/25152459251325174.
Chang, Yupeng, Xu Wang, Jindong Wang, et al. 2024. “A Survey on Evaluation of Large Language Models.” ACM Transactions on Intelligent Systems and Technology 15 (3): 1–45. https://doi.org/10.1145/3641289.
Peng, Ji-Lun, Sijia Cheng, Egil Diau, et al. 2024. A Survey of Useful LLM Evaluation. https://doi.org/10.48550/arXiv.2406.00936.
Brandt, Patrick T, Sultan Alsarra, Vito D’Orazio, et al. 2026. “Extractive Versus Generative Language Models for Political Conflict Text Classification.” Political Analysis 34 (3): 344–72.
Bisbee, James, Joshua D Clinton, Cassy Dorff, Brenton Kenkel, and Jennifer M Larson. 2024. “Synthetic Replacements for Human Survey Data? The Perils of Large Language Models.” Political Analysis 32 (4): 401–16. https://doi.org/10.1017/pan.2024.5.
Wallach, Hanna, Meera Desai, A. Feder Cooper, et al. 2025. “Position: Evaluating Generative AI Systems Is a Social Science Measurement Challenge.” Proceedings of the 42nd International Conference on Machine Learning, Proceedings of machine learning research, vol. 267: 82232–51. ⚠️ Advanced material — Read for the main argument and key concepts.
Wynter, Adrian de. 2025. “Awes, Laws, and Flaws from Today’s LLM Research.” Findings of the Association for Computational Linguistics: ACL 2025 (Vienna, Austria), 12834–54. https://doi.org/10.18653/v1/2025.findings-acl.664. ⚠️ Advanced material — Read for the main argument and key concepts.
11—Uncertainty, Hallucination, and Model Failure
This module develops a taxonomy of failure including fabrication, inconsistency, prompt sensitivity, positional effects, missing evidence, overconfidence, systematic bias, contamination, and error propagation. Students examine how local model failures can change substantive conclusions downstream.
Readings
Li, Junyi, Xiaoxue Cheng, Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2023. “HaluEval: A Large-Scale Hallucination Evaluation Benchmark for Large Language Models.” Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, 6449–64. https://doi.org/10.18653/v1/2023.emnlp-main.397. ⚠️ Advanced material — Read for the main argument and key concepts.
Niu, Cheng, Yuanhao Wu, Juno Zhu, et al. 2024. “RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models.” Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). ⚠️ Advanced material — Read for the main argument and key concepts.
Yu, Wenhao et al. 2024. “ReEval: Automatic Hallucination Evaluation for Retrieval-Augmented Large Language Models via Transferable Adversarial Attacks.” Findings of the Association for Computational Linguistics: NAACL 2024 (Mexico City, Mexico), 1333–51. https://doi.org/10.18653/v1/2024.findings-naacl.85. 🚨 Very advanced material — Read for exposure, not mastery.
12—Synthetic Data and Computational Simulation
This module examines synthetic text, simulated respondents, personas, agents, scenario generation, augmentation, privacy claims, distributional fidelity, and the limits of synthetic evidence. Students distinguish using generated material as a research instrument from treating generated observations as direct evidence about real populations.
Readings
Park, Joon Sung, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023. “Generative Agents: Interactive Simulacra of Human Behavior.” Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. https://doi.org/10.1145/3586183.3606763.
Wang, Angelina, Jamie Morgenstern, and John P. Dickerson. 2025. “Large Language Models That Replace Human Participants Can Harmfully Misportray and Flatten Identity Groups.” Nature Machine Intelligence 7: 400–411. https://doi.org/10.1038/s42256-025-00986-z.
Amershi, Saleema, Dan Weld, Mihaela Vorvoreanu, et al. 2019. “Guidelines for Human-AI Interaction.” Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems, 1–13. https://doi.org/10.1145/3290605.3300233.
Xie, Yueqi, Lemeng Liang, Shuzhen Li, et al. 2026. “Evaluating the Statistical Realism of LLM-Generated Social Science Data.” Proceedings of the National Academy of Sciences 123 (19): e2538145123. https://doi.org/10.1073/pnas.2538145123. ⚠️ Advanced material — Read for the main argument and key concepts.
13—Provenance and Reproducibility for AI-Assisted Research
This module covers model identifiers, versions, prompts, parameters, seeds, timestamps, evidence, logs, caching, data lineage, audit trails, and reproducibility under changing model services. Students consider the special challenges created by proprietary systems whose implementations may change without researcher control.
Readings
Kapoor, Sayash, and Arvind Narayanan. 2023. “Leakage and the Reproducibility Crisis in Machine-Learning-Based Science.” Patterns 4 (9): 100804. https://doi.org/10.1016/j.patter.2023.100804.
Feuerriegel, Stefan et al. 2026. “A Reporting Checklist for Large Language Models in Behavioural Science.” Nature Human Behaviour 10 (7): 1182–86. https://doi.org/10.1038/s41562-026-02492-7.
Gallifant, Jack, Majid Afshar, Saleem Ameen, et al. 2025. “The TRIPOD-LLM Reporting Guideline for Studies Using Large Language Models.” Nature Medicine 31 (1): 60–69. https://doi.org/10.1038/s41591-024-03425-5.
Zeng, Qiuhai, Claire Jin, Xinyue Wang, Yuhan Zheng, and Qunhua Li. 2025. “AIRepr: An Analyst-Inspector Framework for Evaluating Reproducibility of LLMs in Data Science.” Findings of the Association for Computational Linguistics: EMNLP 2025, 10170–201. https://doi.org/10.18653/v1/2025.findings-emnlp.539.
14—Ethics, Privacy, Bias, and Responsible Generative AI
This module addresses sensitive data, consent, intellectual property, bias, fairness, labor, disclosure, authorship, environmental costs, research integrity, and governance. Students examine the responsibilities created when generative systems mediate the production of scientific evidence.
Readings
Jiao, Junfeng, Saleh Afroogh, Yiming Xu, and Connor Phillips. 2025. “Navigating LLM Ethics: Advancements, Challenges, and Future Directions.” AI and Ethics 5 (6): 5795–819.
Bender, Emily M., Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?” Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency, 610–23. https://doi.org/10.1145/3442188.3445922.
Gallegos, Isabel O., Ryan A. Rossi, Joe Barrow, et al. 2024. “Bias and Fairness in Large Language Models: A Survey.” Computational Linguistics 50 (3): 1097–179. https://doi.org/10.1162/coli_a_00524.
Carlini, Nicholas, Florian Tramèr, Eric Wallace, et al. 2021. “Extracting Training Data from Large Language Models.” 30th USENIX Security Symposium (USENIX Security 21), 2633–50. https://www.usenix.org/conference/usenixsecurity21/presentation/carlini-extracting.
Palmer, Alexis, and Arthur Spirling. 2023. “Large Language Models Can Argue in Convincing Ways about Politics, but Humans Dislike AI Authors: Implications for Governance.” Political Science 75 (3): 281–91.
Porsdam Mann, Sebastian, Anuraag A. Vazirani, Mateo Aboy, et al. 2024. “Guidelines for Ethical Use and Acknowledgement of Large Language Models in Academic Writing.” Nature Machine Intelligence 6: 1272–74. https://doi.org/10.1038/s42256-024-00922-7.
15—Synthesis: Auditing a Generative-AI Research Pipeline
This module integrates research questions, constructs, knowledge representations, model tasks, schemas, APIs, human oversight, evaluation, error analysis, provenance, ethics, and reproducibility. Students audit how these elements work together across an AI-assisted research workflow, asking whether every consequential step between source material and empirical claim can be defended and audited.
Readings
Barrie, Christopher, Lisa P. Argyle, James Bisbee, et al. 2026. AI and Research Methods. https://doi.org/10.33774/apsa-2026-h59kk.
Paci, Simone. 2026. “Is Research Safe in the AI Revolution? LLMs Fall Short of Expert Benchmarks in Scientific Evidence Evaluation.” Chinese Political Science Review, 1–23.
Gebru, Timnit, Jamie Morgenstern, Briana Vecchione, et al. 2021. “Datasheets for Datasets.” Communications of the ACM 64 (12): 86–92. https://doi.org/10.1145/3458723.
Mitchell, Margaret, Simone Wu, Andrew Zaldivar, et al. 2019. “Model Cards for Model Reporting.” Proceedings of the Conference on Fairness, Accountability, and Transparency, 220–29. https://doi.org/10.1145/3287560.3287596.
Evaluation
TBD.