Special Topic Courses
Please contact the instructor if you have questions about prerequisites or other aspects of the course.
Fall 2026
- ST 495/590: Environmental Statistics
- Instructor: Erin Schliep
- Description: This course aims to provide an introduction to the types of statistical analyses used in environmental studies. Topics include sampling design, hierarchical modeling, spatial and temporal statistics, and extremes. The course focuses on applications in a variety of different areas including ecology, fish and wildlife monitoring, remote sensing, and climate.
- ST 790: Spectral methods for statistics and machine learning
- Instructor: Minh Tang
- Description: Spectral methods refer to a collection of algorithms based on the singular values and singular vectors (eigenvalues and eigenvectors) of large data matrices. This class of algorithms provide a simple yet surprisingly effective approach for inference on matrix-valued data. Examples include nonlinear dimension reduction, covariance matrix estimation, matrix denoising, community detection in networks, rankings from pairwise comparisons, tensor regression. This course will present a comprehensive introduction to spectral methods from a modern statistical perspective, with emphasis on both practical applications of spectral methods as well as their theoretical guarantees under general signal-plus-noise models.
- ST 790: Extreme Value Analysis
- Instructor: Sujit Ghosh
- Description: This course provides an introduction to the theory and practice of extreme value statistics, with applications to finance and climate. Topics include heavy-tailed distributions, univariate and multivariate extreme value theory, regression models for extremes, and dependence modeling using copulas. The course integrates theoretical foundations with hands-on data analysis using standard R packages.
Spring 2026
- ST 495/590: Data Visualization
- Instructor: Dan Harris
- Description: This course explores concepts and techniques in data visualization, equipping students with the skills to create sophisticated and impactful visual representations of data. Emphasis is placed on understanding the theoretical foundations, capturing and managing data, applying advanced visualization tools, and evaluating the effectiveness of different visualization techniques.
- ST 790: Nonparametric Bayesian Inference
- Instructor: Subhashis Ghoshal
- Description: This course is primarily intended for Statistics Ph.D. students who have research interest in theoretical, methodological, computational and applied aspects of contemporary Bayesian statistics, which primarily involves analysis of nonparametric, semi-parametric and high dimensional models.
- ST 790: Statistical Methods for Analysis with Missing Data
- Instructor: Marie Davidian
- Description: Missing data are common in practice and especially in health sciences research involving human subjects. The classical definition of missingness is that data that were intended to be collected according to a predetermined plan or experimental design were not. For example, in a clinical trial designed to collect data prospectively on every participant at a prespecified series of follow-up times, some subjects may fail to appear at the clinic for intended measurements at one or more times, refuse to provide some intended information, or drop out of the study and never return after a certain point. If the reasons for failure to appear or dropout are related to the issues under study, e.g., if subjects who are benefiting less from their assigned treatments are more likely to drop out, intuitively, failure to acknowledge this somehow could distort conclusions. Missingness also arises in analysis of “real world,” observational data collected retrospectively. Missing data have important implications for analysis. At the very least, there is a loss of information and thus a reduction in precision of inference on the population of interest relative to that intended. Of much greater concern is the potential for biased and misleading inferences that can result if the reasons for missingness are related to outcomes of interest and other features. Accordingly, principled methods to take this challenge into appropriate account are required. This course will provide an overview of statistical frameworks and methods for analysis in the presence of missing data. Because the literature is so vast, with missing data methods a topic of current research, the focus is on fundamental concepts and methodology that are the required background for reading the current specialized literature and for contending with missing data in practice. Both methodological developments and applications are emphasized, including implementation in software.
Fall 2025
- ST/BMA 590: Statistical Modeling in Ecology
- Instructor: Kevin Gross
- Description: This course explores advanced statistical methods commonly used in ecology, focusing on data that fall outside traditional linear regression and ANOVA methods. Topics covered include likelihood and numerical optimization, smooth regression techniques (e.g., loess and splines), generalized linear models (logistic and Poisson regression), mixed-effect models, and an introduction to Bayesian inference. The course will emphasize practical application and will use real-world examples from population and community ecology. The course requires a background in calculus, linear statistical models, matrix algebra, and basic probability. Computing instruction will be provided in R. The course is suited for graduate students or advanced undergraduates conducting research. Assessment is based on homework and class participation.
- ST 790: Nonparametric Regression
- Instructor: Luo Xiao
- Description: The course shall cover topics in nonparametric regression including density estimation, nonparametric regression for the univariate case, nonparametric regression in high dimension, deep learning and also a few selected topics, e.g., partially linear models, generalized additive models, semi-parametric regression for longitudinal data, functional data analysis, and deep generative models. The statistical methods that will be covered include kernel smoothing, splines and neural networks. Both the theoretic and practical implementation of these methods will be discussed. Real data examples will be used for illustration.
- ST 790: Navigating the PhD Program and Beyond: Perspectives, skills, and strategies
- Instructors: Srijan Sengupta and Jon Williams
- Description: This course will expose PhD students in statistics to a structured overview of the skills they are expected to learn, how they are expected to develop these skills, what type of career opportunities they are being trained for, and how to best align with a particular career path. The course will focus on the skill development necessary for research-focused academic, teaching-focused academic, and industry career paths. The course will mainly focus on the following aspects, in particular: beginning in research in statistics, statistics research landscape, professional development, scientific writing, research, teaching, and job searches. The course will involve guest speakers, individual presentations, and class participation.
Spring 2025
- ST 790: Advanced Design and Analysis of Experiments
- Instructor: Jon Stallrich
- Description: The course provides an overview of foundational approaches to ranking experimental designs under an optimality criterion. Approximate design theory and equivalence theorems are introduced and search algorithms are discussed for finding exact optimal designs. New techniques for screening experiments, such as definitive screening designs, OMARS designs, and supersaturated designs, will then be studied. An overview will be given for the sequential design and analysis of computer experiments, which is based around Gaussian processes and Bayesian optimization. Time permitting, the instructor may choose to introduce important design problems with online-controlled experiments, optimal designs for penalized regression, and optimal design for generalized linear models.
- ST 790: Dynamic Treatment Regimes
- Instructor: Marie Davidian
- Description: This course will provide a comprehensive introduction to methodology for data-based development and evaluation of dynamic treatment regimes. A dynamic treatment regime is a set of sequential decision rules, each corresponding to a key point in a disease or disorder process at which a decision on the next treatment action must be made. Each rule takes patient information to that point as input and returns the treatment s/he should receive from among the available options, thus tailoring treatment decisions to a patient’s individual characteristics. Dynamic treatment regimes formalize how clinicians make decisions in practice by synthesizing evolving information on a patient and are thus of considerable importance in precision medicine. Dynamic treatment regimes are also relevant in other contexts in which sequential decisions on interventions or policies must be made, as in education, engineering, economics and finance, and resource management. Of critical importance is the notion of an optimal treatment regime, one that, if used to select treatments for the patient population, would lead to the most beneficial outcome on average. Methods for estimation of dynamic treatment regimes and in particular optimal treatment regimes from data will be motivated and developed through a formal time-dependent causal inference framework. The gold standard study design for developing and evaluating regimes is the sequential multiple assignment randomized trial (SMART), considerations for which will be discussed. Inference for optimal treatment regimes is a nonstandard statistical problem and is thus notoriously difficult; an introduction to this challenge will be presented. Examples throughout the course will be drawn from cancer and other chronic disease research and research in the behavioral, educational, and other sciences. Use of the comprehensive R package DynTxRegime to implement many of the methods discussed in the lectures will be introduced. Students completing this course will have a foundation in causal inference and fundamental results and methods for dynamic treatment regimes that will provide the basis for study of the rapidly evolving literature on dynamic treatment regimes and data-driven sequential decision-making in precision medicine.
- ST 295: Introduction to the Foundations of Data Science with R
- Instructor: Elijah Meyer
- Description: In this course, we will discover patterns in data through the use of effective data visualization tools, exploratory data analysis, and modeling techniques. We will gain experience working with real-world messy data, and use data tidying and data wrangling techniques to make the data suitable for exploration. An emphasis will be placed on creating reproducible and shareable work. This course highlights the importance of and uses version control software for individual and collaborative work. The course will focus on the R statistical computing language.