↓Skip to main content

Optimization for Data Science - CAAM/DATA/STAT 31019 (Autumn 2026)

Optimization problems arising in modern data science often fall outside the traditional assumptions of classical optimization theory: objectives may be nonconvex, nonsmooth, highly overparameterized, and accessible only through noisy gradient information. This course develops algorithms and theory for this regime, including stochastic and variance-reduced gradient methods, adaptive and preconditioned algorithms, mirror descent and non-Euclidean geometry, nonconvex complexity theory and benign landscape analysis, weakly convex optimization, automatic and implicit differentiation, and min-max formulations. A central theme is the interplay between optimization and statistical estimation: algorithmic choices can influence regularization, generalization, and implicit bias, while statistical considerations determine the target accuracy. Throughout, the course will balance theoretical analysis with computational implementation.

Coordinates #

Time: MW, 3:00–4:20 p.m.
Location: Ryerson Physical Laboratory, Room 176
Assignments and announcements: Canvas.

Personnel #

Instructor:
Mateo Díaz (mateodd@uchicago.edu)
Office: Suzuki Center, 304
Office hours: M, 4:45–5:45 p.m.

Syllabus #

The syllabus can be found here.

Lecture notes #

The lecture notes will be posted here.

Homework #

Approximately three assignments will be posted on Canvas, combining proofs, algorithm analysis, and computational experiments. Most assignments will include implementation and testing of an optimization method. Submit your solutions as a PDF on Canvas. For problems requiring a computational implementation, include your code in the PDF and a link to a public GitHub repository containing the code.

Textbook #

There is no required textbook. Lecture notes and selected readings will be posted on this website and Canvas. The following are useful background references; additional papers will accompany the more specialized topics.

  • D. Drusvyatskiy, Convex Analysis and Nonsmooth Optimization (2020), lecture notes.
  • J. Nocedal and S. J. Wright, Numerical Optimization, second edition, Springer (2006), book information.
  • S. Bubeck, Convex Optimization: Algorithms and Complexity (2015), available online.
  • L. Bottou, F. E. Curtis, and J. Nocedal, Optimization Methods for Large-Scale Machine Learning (revised 2018), available online.

Grading system #

Course grades are based on homework, a midterm, a final project, and participation. We will use an optimization-based grading scheme, invented by Ben Grimmer, that chooses the weights for the different components to maximize each student’s grade subject to reasonable constraints.

For each student, let \(C_H,C_M,C_F,C_P\) be the homework, midterm, final-project, and participation scores, respectively, each on a scale from 0 to 100. Let \(H,M,F,P\) be the corresponding weights, measured in percentage points. The course grade, also on a scale from 0 to 100, is the optimal value of

\[\begin{aligned} \max_{H,M,F,P\in\mathbb{R}}\quad & \frac{C_H H+C_M M+C_F F+C_P P}{100} \\ \text{subject to}\quad & H+M+F+P=100, \\ & 10 \le H \le 25, \\ & 30 \le M, \\ & M \le F, \\ & M+F\le 80, \\ & 0\le P\le 10. \end{aligned}\]

Thus, homework receives between 10 and 25 percent of the course grade, the midterm receives at least 30 percent, and the final project receives at least as much weight as the midterm. The combined midterm and final-project weight is at most 80 percent, and participation receives between 0 and 10 percent.

Midterm and Final project #

There will be one in-class midterm; the date will be announced in class and by email. The midterm will be handwritten and closed-book, with no notes, books, or electronic resources permitted. There will not be a final exam.

The final project consists of a written report and an in-person presentation. Students will work in groups of three to six, choosing a topic broadly related to optimization for data science. Contributions may be theoretical, methodological, or focused on an application.

A central goal is to develop research judgment: recognizing interesting questions worth pursuing. Your group should read relevant papers, identify an open question, and work toward answering it over the quarter. A finished paper or complete solution is not expected; a strong project may formulate a compelling question, explain its significance, and present meaningful partial results.

The project should demonstrate substantive work beyond what an LLM can produce in a single response. One evaluation criterion will be whether an LLM, prompted by the instructor, can solve the proposed question and produce results on par with or superior to the report using one prompt. You are encouraged to meet with the instructor periodically for feedback. Detailed project instructions will be provided early in the course.

Participation #

The grading scheme permits a participation weight between 0 and 10 percent; full course marks are possible without participation credit. Full participation credit can be earned through engagement during or after lectures, office-hour discussions, or thoughtful questions.

See the syllabus for the full policies on collaboration, AI tools, deadlines, extensions, and accommodations.