Enrollment is closed for Fall 2026
We’ll also run the course in future semesters. Watch out for updates for future semesters. All course materials are openly accessible below.
Fall 2026
1 Credit, 7 Weeks First, S/U Grading, Fall enrollment closed
We’ll also run the course in future semesters. Watch out for updates for future semesters. All course materials are openly accessible below.

Organizer and Lead Instructor

Head TA

Teaching Assistant

Teaching Assistant

Advisor

Advisor

Advisor

Faculty Advisor
The course is led by CAIA members with faculty advising from Cornell.

Organizer and Lead Instructor

Head TA

Teaching Assistant

Teaching Assistant

Advisor

Advisor

Advisor

Faculty Advisor
How can we make AI systems do what we actually want? This student-led course introduces language model training, alignment, interpretability, evaluations, and scalable oversight. Through lectures, hands-on work, and discussion, we connect technical methods with broader questions about AI policy and governance.
The course includes Friday lectures, guided notebooks, Monday paper discussions, and a final project. We use point-based S/U grading: 145 points are available, and 100 points earns a pass.
Lecture. Each of seven Friday lectures earns 5 points, up to 35 points. Lectures introduce the week's core ideas, technical foundations, and research context.
Notebook. Weeks 2–6 each include a guided take-home notebook. Satisfactory completion earns 10 points, up to 50 points. Expect roughly 30 lines of student-written code within a provided framework, focused on a hands-on experiment and short analysis.
Discussion. Monday discussions examine a frontier or recent paper related to the preceding lecture. Each attendance earns 5 points, capped at 15 points.
Project proposal. Earn 5 points for defining an AI safety question, hypothesis, method, and expected result. You'll receive feedback before final-project work begins.
Final project. Earn 40 points for reproducing or extending a result, building an evaluation, comparing methods, or testing a safety hypothesis. Agentic coding tools are encouraged, but you remain responsible for understanding and validating the work. Compute credits will be provided.
Week 1
Introduction to AI Safety
Topics
Week 2
AI Alignment and RLHF
Topics
Slides
Discussion Readings
Week 3
Reward Hacking and Goal Misgeneralization
Topics
Slides
Discussion Readings
Week 4
Interpretability
Topics
Slides
Discussion Readings
Week 5
Evaluating Dangerous Capabilities
Topics
Slides
Slides (TBD)
Discussion Readings
Week 6
Control and Scalable Oversight
Topics
Slides
Slides (TBD)
Discussion Readings
Week 7
Policy, Governance, and Forecasting
Topics
Slides
Slides (TBD)
Discussion Readings