Working Efficiently with Data and Code
Presentation
Motivation Computational research can become complicated surprisingly quickly. Files multiply, code breaks, analyses become difficult to reproduce, and seemingly simple tasks consume hours of unnecessary work as projects spread across files, data, coding environments, and outputs. This course teaches the practical foundations for organizing computational work so that projects remain understandable, reproducible, efficient, and maintainable even as their scale and complexity increase.
Approach The course develops an integrated workflow for computational research. Students begin by thinking algorithmically about problems and organizing the files, data, and environments needed to solve them, then use the command line, programming languages, databases, Git and GitHub, reusable code, debugging and testing, dependency management, automation, and reproducible reporting as parts of the same system. The emphasis is on understanding how these elements work together rather than treating them as isolated technical skills.
Principle Throughout the course, the guiding principle is that good computational work is not merely code that runs. It should also be possible to understand what was done, determine where inputs came from, reproduce outputs, identify errors, modify the analysis, and hand the project to someone else. By the end of the course, students should be able to transform an improvised collection of scripts, datasets, and outputs into a well-structured and reproducible computational project.
Objectives
The course builds toward the following learning objectives:
- Decompose computational problems into explicit, ordered steps that can be reused and automated.
- Design coherent directory and naming structures for computational projects.
- Navigate and work within computational environments using appropriate graphical and command-line tools.
- Choose appropriate programming languages, data formats, and storage models for the requirements of a task.
- Use Git and GitHub to preserve project history, experiment safely, collaborate with others, and recover from mistakes.
- Write readable, modular, reusable code and diagnose failures systematically.
- Manage dependencies and computational environments so that projects remain reproducible across machines and over time.
- Automate repetitive and multi-stage computational tasks and identify performance bottlenecks.
- Produce a documented, reproducible computational research project in which inputs, transformations, code, and outputs form an auditable workflow.
Modules
01—Algorithmic Thinking and Computational Work as a System
This module introduces algorithmic thinking through problem decomposition, sequencing, abstraction, inputs, transformations, outputs, dependencies, and explicit workflow design. Students learn to treat computational work as a system whose steps can be inspected, tested, repeated, and improved.
Readings
Wing, Jeannette M. 2006. “Computational Thinking.” Communications of the ACM 49 (3): 33–35. https://doi.org/10.1145/1118178.1118215.
Futschek, Gerald. 2006. “Algorithmic Thinking: The Key for Understanding Computer Science.” International Conference on Informatics in Secondary Schools-Evolution and Perspectives, 159–68.
Irving, Damien, Kate Hertweck, Luke Johnston, Joel Ostblom, Charlotte Wickham, and Greg Wilson. 2021. Research Software Engineering with Python. CRC Press. 📑 Chapter 1 (Getting Started).
Janssens, Jeroen. 2021. Data Science at the Command Line. 2nd ed. O’Reilly Media. 📑 Chapter 1 (Introduction).
The Carpentries. n.d. “Instructor Training.” The Carpentries. Accessed September 4, 2026. https://preview.carpentries.org/instructor-training/. 📑 Episode 2 (Building Skill With Practice).
02—Files, Folders, Paths, and Naming
This module covers directory structures, absolute and relative paths, naming conventions, extensions, hidden files, project roots, and separation of inputs from generated outputs. Students learn how project organization reduces cognitive load and makes later automation and collaboration possible.
Readings
The Carpentries. n.d. “The Unix Shell.” Accessed September 2, 2026. https://swcarpentry.github.io/shell-novice/. 📑 Episode 2 (Navigating Files and Directories) and Episode 3 (Working With Files and Directories).
Wickham, Hadley, Mine Çetinkaya-Rundel, and Garrett Grolemund. 2023. R for Data Science. 2nd ed. O’Reilly Media. https://r4ds.hadley.nz/. 📑 Chapter 6 (Workflow: Scripts and Projects).
Alvarado-Mena, Edwin. 2026. “How to Organize Your Computer for Data Work.” July 28. https://alvaradocss.com/posts/organize-data-projects/.
03—Understanding Your Computer More Deeply
This module introduces the command line and operating system as computational interfaces. Students examine shell navigation, file operations, pipes, redirection, environment variables, processes, permissions, and the relationship among files, programs, memory, and running processes.
Readings
The Carpentries. n.d. “The Unix Shell.” Accessed September 2, 2026. https://swcarpentry.github.io/shell-novice/. 📑 Episode 1 (Introducing the Shell), Episode 4 (Pipes and Filters), and Episode 6 (Shell Scripts).
Irving, Damien, Kate Hertweck, Luke Johnston, Joel Ostblom, Charlotte Wickham, and Greg Wilson. 2021. Research Software Engineering with Python. CRC Press. 📑 Chapter 2 (The Basics of the Unix Shell), Chapter 3 (Building Tools with the Unix Shell), and Chapter 4 (Going Further with the Unix Shell).
Janssens, Jeroen. 2021. Data Science at the Command Line. 2nd ed. O’Reilly Media. 📑 Chapter 2 (Getting Started).
04—Programming Languages and Coding Environments
This module compares programming languages, interpreters, packages, ecosystems, notebooks, IDEs, and terminals. Students use R, Python, shell tools, and representative coding environments as examples to learn that language and interface choice should follow the task rather than habit.
Readings
Zelle, John M. 2017. Python Programming: An Introduction to Computer Science. 3rd ed. Franklin, Beedle & Associates. 📑 Chapter 1 (Computers and Programs) and Chapter 2 (Writing Simple Programs).
The Carpentries. n.d. “Programming Lessons.” Accessed September 5, 2026. https://software-carpentry.org/lessons/. 📑 Introductory Programming with Python or Programming with R series.
Janssens, Jeroen. 2021. Data Science at the Command Line. 2nd ed. O’Reilly Media. 📑 Chapter 10 (Polyglot Data Science).
Irving, Damien, Kate Hertweck, Luke Johnston, Joel Ostblom, Charlotte Wickham, and Greg Wilson. 2021. Research Software Engineering with Python. CRC Press. 📑 Chapter 5 (Building Command-Line Tools with Python).
05—Git and Version Control
This module introduces repositories, commits, diffs, history, staging, branches, merging, restoration, and meaningful commit practices. Students learn to use version control as infrastructure for making computational experimentation traceable and reversible.
Readings
The Carpentries. n.d. “Version Control with Git.” Accessed September 2, 2026. https://swcarpentry.github.io/git-novice/. 📑 Episode 1 (Automated Version Control), Episode 2 (Setting Up Git), Episode 3 (Creating a Repository), Episode 4 (Tracking Changes), Episode 5 (Exploring History), and Episode 6 (Ignoring Things).
Irving, Damien, Kate Hertweck, Luke Johnston, Joel Ostblom, Charlotte Wickham, and Greg Wilson. 2021. Research Software Engineering with Python. CRC Press. 📑 Chapter 6 (Using Git at the Command Line).
Chacon, Scott, and Ben Straub. 2014. Pro Git. 2nd ed. Apress. 📑 Chapter 1 (Getting Started) and Chapter 2 (Git Basics).
Bryan, Jennifer. 2018. “Excuse Me, Do You Have a Moment to Talk about Version Control?” The American Statistician 72 (1): 20–27. https://doi.org/10.1080/00031305.2017.1399928.
06—GitHub and Collaborative Workflows
This module develops version control into a collaborative workflow through remotes, cloning, pushing, pulling, issues, branches, pull requests, code review, and releases. Students learn how explicit project history supports coordination as well as individual work.
Readings
The Carpentries. n.d. “Version Control with Git.” Accessed September 2, 2026. https://swcarpentry.github.io/git-novice/. 📑 Episode 7 (Remotes in GitHub), Episode 8 (Collaborating), and Episode 9 (Conflicts).
Irving, Damien, Kate Hertweck, Luke Johnston, Joel Ostblom, Charlotte Wickham, and Greg Wilson. 2021. Research Software Engineering with Python. CRC Press. 📑 Chapter 7 (Going Further with Git) and Chapter 8 (Working in Teams).
Chacon, Scott, and Ben Straub. 2014. Pro Git. 2nd ed. Apress. 📑 Chapter 5 (Distributed Git).
07—Tidy Data, SQL, and Data Models
This module makes tidy data an explicit organizing principle and then broadens the discussion to data storage and interchange. Students examine CSV and JSON, relational tables, primary and foreign keys, joins, basic SQL queries, schemas, raw versus processed data, and the tradeoffs among relational, document-oriented, and graph-oriented/NoSQL models.
Readings
Wickham, Hadley. 2014. “Tidy Data.” Journal of Statistical Software 59: 1–23. https://doi.org/10.18637/jss.v059.i10.
Broman, Karl W, and Kara H Woo. 2018. “Data Organization in Spreadsheets.” The American Statistician 72 (1): 2–10. https://doi.org/10.1080/00031305.2017.1375989.
Wickham, Hadley, Mine Çetinkaya-Rundel, and Garrett Grolemund. 2023. R for Data Science. 2nd ed. O’Reilly Media. https://r4ds.hadley.nz/. 📑 Chapter 5 (Data Tidying), Chapter 19 (Joins), and Chapter 21 (Databases).
Beaulieu, Alan. 2020. Learning SQL: Generate, Manipulate, and Retrieve Data. O’Reilly Media. 📑 Chapter 1 (A Little Background), Chapter 3 (Query Primer), and Chapter 5 (Querying Multiple Tables).
IBM. n.d. “What Is Data Modeling?” Accessed August 21, 2026. https://www.ibm.com/think/topics/data-modeling.
08—Writing Reusable Code
This module covers functions, parameters, return values, modules, abstraction, interfaces, naming, the DRY principle, and separation of concerns. Students learn how to recognize repeated logic and convert it into components that can be tested and reused.
Readings
Wilson, Greg, D. A. Aruliah, C. Titus Brown, et al. 2014. “Best Practices for Scientific Computing.” PLOS Biology 12 (1): e1001745. https://doi.org/10.1371/journal.pbio.1001745.
Wickham, Hadley. 2019. Advanced r. 2nd ed. Chapman; Hall/CRC. 📑 Chapter 6 (Functions).
Janssens, Jeroen. 2021. Data Science at the Command Line. 2nd ed. O’Reilly Media. 📑 Chapter 4 (Creating Command-line Tools).
The tidyverse team. n.d. “The Tidyverse Style Guide.” Accessed September 2, 2026. https://style.tidyverse.org/.
Google. n.d. “Google’s r Style Guide.” Accessed September 2, 2026. https://google.github.io/styleguide/Rguide.html.
Google. n.d. “Google Python Style Guide.” Accessed September 2, 2026. https://google.github.io/styleguide/pyguide.html.
09—Debugging and Testing
This module treats debugging as a systematic method rather than trial and error. Students practice reading errors, isolating failures, constructing minimal examples, using assertions and logging, and designing unit and integration tests.
Readings
Irving, Damien, Kate Hertweck, Luke Johnston, Joel Ostblom, Charlotte Wickham, and Greg Wilson. 2021. Research Software Engineering with Python. CRC Press. 📑 Chapter 11 (Testing Software) and Chapter 12 (Handling Errors).
Zeller, Andreas. 2009. Why Programs Fail: A Guide to Systematic Debugging. 2nd ed. Morgan Kaufmann. 📑 Chapter 1 (How Failures Come to Be).
10—Dependencies and Computational Environments
This module examines package management, versions, lockfiles, virtual environments, configuration files, secrets, portability, and reproducible environments. Students learn to identify what another machine needs in order to reproduce a computational result.
Readings
Irving, Damien, Kate Hertweck, Luke Johnston, Joel Ostblom, Charlotte Wickham, and Greg Wilson. 2021. Research Software Engineering with Python. CRC Press. 📑 Chapter 14 (Creating Packages with Python; especially §14.2 Virtual Environments and §14.4 What Installation Does). ⚠️ Advanced material — Read for the main argument and key concepts.
Moreau, David, Kristina Wiebels, and Carl Boettiger. 2023. “Containers for Computational Reproducibility.” Nature Reviews Methods Primers 3: 50. https://doi.org/10.1038/s43586-023-00236-9. ⚠️ Advanced material — Read for the main argument and key concepts.
11—Automating Repetitive Work
This module introduces scripts, batch processing, loops, parameterization, pipelines, task runners, scheduled work, and the elimination of fragile manual steps. Students learn to use automation as a strategy for reliability as much as convenience.
Readings
Irving, Damien, Kate Hertweck, Luke Johnston, Joel Ostblom, Charlotte Wickham, and Greg Wilson. 2021. Research Software Engineering with Python. CRC Press. 📑 Chapter 9 (Automating Analyses with Make).
Janssens, Jeroen. 2021. Data Science at the Command Line. 2nd ed. O’Reilly Media. 📑 Chapter 6 (Project Management with Make).
12—Reproducible Documents and Reporting
This module introduces Quarto and other computational-document approaches to literate programming. Students examine citations, tables, figures, parameterized reports, and the separation of analysis from presentation so that reports can be regenerated from code and evidence.
Readings
Xie, Yihui, J. J. Allaire, and Garrett Grolemund. 2018. R Markdown: The Definitive Guide. Chapman; Hall/CRC. 📑 Chapter 1 (Installation) and Chapter 2 (Basics).
Posit. n.d. “Get Started with Quarto.” Accessed September 4, 2026. https://quarto.org/docs/get-started/.
Rule, Adam, Amanda Birmingham, Cristal Zuniga, et al. 2019. “Ten Simple Rules for Writing and Sharing Computational Analyses in Jupyter Notebooks.” PLOS Computational Biology 15 (7): e1007007. https://doi.org/10.1371/journal.pcbi.1007007.
Alvarado-Mena, Edwin. 2026. “Rendering Professional Documents with Quarto.” July 28. https://alvaradocss.com/posts/professional-documents-quarto/.
13—Scaling Computational Work
This module covers profiling, vectorization, memory use, caching, parallel processing, incremental computation, and the principle of optimizing only after identifying the actual bottleneck. Students learn to diagnose computational bottlenecks and choose optimization strategies that address the actual source of poor performance.
Readings
Janssens, Jeroen. 2021. Data Science at the Command Line. 2nd ed. O’Reilly Media. 📑 Chapter 8 (Parallel Pipelines). ⚠️ Advanced material — Read for the main argument and key concepts.
Irving, Damien, Kate Hertweck, Luke Johnston, Joel Ostblom, Charlotte Wickham, and Greg Wilson. 2021. Research Software Engineering with Python. CRC Press. 📑 Appendix E (Working Remotely). ⚠️ Advanced material — Read for the main argument and key concepts.
14—Building Reproducible Research Projects
This module integrates project architecture, entry points, configuration, READMEs, documentation, data provenance, workflow diagrams, licenses, dependency management, and project handoffs. Students evaluate reproducibility as a property of the whole project rather than a single script.
Readings
Wilson, Greg, Jennifer Bryan, Karen Cranston, Justin Kitzes, Lex Nederbragt, and Tracy K. Teal. 2017. “Good Enough Practices in Scientific Computing.” PLOS Computational Biology 13 (6): e1005510. https://doi.org/10.1371/journal.pcbi.1005510.
Marwick, Ben, Carl Boettiger, and Lincoln Mullen. 2018. “Packaging Data Analytical Work Reproducibly Using r (and Friends).” The American Statistician 72 (1): 80–88. https://doi.org/10.1080/00031305.2017.1375986.
Sandve, Geir Kjetil, Anton Nekrutenko, James Taylor, and Eivind Hovig. 2013. “Ten Simple Rules for Reproducible Computational Research.” PLOS Computational Biology 9 (10): e1003285. https://doi.org/10.1371/journal.pcbi.1003285.
National Academies of Sciences, Engineering, and Medicine. 2019. Reproducibility and Replicability in Science. The National Academies Press. https://doi.org/10.17226/25303. 📑 Chapter 3 (Understanding Reproducibility and Replicability), Chapter 4 (Reproducibility), and Chapter 6 (Improving Reproducibility and Replicability).
Gandrud, Christopher. 2020. Reproducible Research with r and RStudio. 3rd ed. Chapman; Hall/CRC. 📑 Chapter 2 (Getting Started with Reproducible Research).
15—Synthesis: Maintaining Reproducible Computational Work
This module integrates organization, version control, modularity, databases, automation, testing, reproducibility, documentation, and handoff. Students evaluate a realistic computational workflow as a maintainable research object, asking whether they and others can understand, run, audit, hand off, and extend it over time.
Readings
Irving, Damien, Kate Hertweck, Luke Johnston, Joel Ostblom, Charlotte Wickham, and Greg Wilson. 2021. Research Software Engineering with Python. CRC Press. 📑 Chapter 15 (Finale).
Kitzes, Justin, Daniel Turek, and Fatma Deniz, eds. 2018. The Practice of Reproducible Research: Case Studies and Lessons from the Data-Intensive Sciences. University of California Press. 📑 Case Study 4 (Estimating the Effect of Soldier Deaths on the Military Labor Supply).
Evaluation
TBD.