Teaching · Spring 2025

Foundations of Reinforcement Learning

Graduate course · Spring Semester 2025 · ETH Zürich

Course code263-5255-00L
TermSpring 2025
LevelGraduate course
Weeks9

Overview

This course focuses on the theoretical and algorithmic foundations of reinforcement learning, through the lens of optimization, modern approximation, and learning theory. It targets students with strong research interests in reinforcement learning, optimization under uncertainty, and data-driven control.

Course information

Language
English
Materials
Lectures, slides and announcements are distributed through ETH Moodle.
Contact
All course-related questions should be directed to forl2025eth@gmail.com.

Learning objectives

  • Define the key features of reinforcement learning that distinguish it from standard machine learning.
  • Identify the strengths and limitations of various reinforcement learning algorithms.
  • Formulate and solve sequential decision-making problems by applying relevant reinforcement learning tools.
  • Recognize the common boundary between optimization and reinforcement learning.
  • Generalize or discover new applications, algorithms, or theories of reinforcement learning towards independent research on the topic.

Course outline

Week 1Introduction

Learning objectives and course logistics · An overview of RL · A primer on optimization

Week 2Dynamic Programming and Linear Programming

Markov Decision Processes · Bellman equations and Bellman optimality · Value and policy iteration · Linear programming

Week 3Value-based RL

From planning to reinforcement learning · Model-free prediction · Model-free control · Function approximation · Convergence analysis

Week 4Policy-based RL I (Algorithms)

Overview of policy-based RL · Policy gradient estimation · Policy gradient methods · Natural policy gradient · Beyond PG: TRPO, PPO

Week 5Policy-based RL II (Theory)

Performance difference lemma · Global convergence of policy gradient methods · Global convergence of natural policy gradient methods · Remarks on sample efficiency

Week 6Multi-agent RL and Markov Games

From single agent to multiple agents · Normal form and repeated games · Markov games and algorithms · Zero-sum Markov games and algorithms

Week 7Imitation Learning

Offline imitation learning: behaviour cloning · Online interactive imitation learning: DAGGER, AggreVaTe · Inverse reinforcement learning: feature expectation matching, Max-Ent IRL · Generative adversarial imitation learning (GAIL)

Week 8Deep RL

Actor-critic methods · Overview of deep RL · Value-based deep RL · Policy-based and actor-critic deep RL · From deep learning theory to deep RL theory · Convergence analysis of neural TD-learning and neural actor-critic

Week 9Going Beyond: Model-based RL, Offline RL, Many-agent RL

Model-based RL · Offline RL · Many-agent RL · Summary and outlook

Recommended references

There is no required textbook. Lectures and class discussions are mostly based on classical and recent papers on the topic.

RL textbooks

  • [S09] Algorithms for Reinforcement Learning, Csaba Szepesvári, 2009.
  • [SB18] Reinforcement Learning: An Introduction, Richard S. Sutton and Andrew G. Barto, 2018.
  • [B19] Reinforcement Learning and Optimal Control, Dimitri P. Bertsekas, 2019.
  • [AJK20] Reinforcement Learning: Theory and Algorithms, Alekh Agarwal, Nan Jiang, and Sham M. Kakade, 2020.
  • [M21] Control Systems and Reinforcement Learning, S. Meyn, Cambridge University Press, 2021.
  • [KWM22] Algorithms for Decision Making, Mykel J. Kochenderfer, Tim A. Wheeler, and Kyle H. Wray, MIT Press, 2022.

Optimization foundations

ML and AI foundations

People

Portrait of Prof. Dr. Niao He

Prof. Dr. Niao He

Lecturer

Portrait of Jiawei Huang

Jiawei Huang

Head teaching assistant

Portrait of Adrian Müller

Adrian Müller

Teaching assistant

Portrait of Riccardo De Santi

Riccardo De Santi

Teaching assistant

Portrait of Ashwin Shenai

Ashwin Shenai

Teaching assistant

Portrait of Tianxu An

Tianxu An

Teaching assistant