
Prof. Dr. Niao He
Lecturer
Teaching · Spring 2025
Graduate course · Spring Semester 2025 · ETH Zürich
This course focuses on the theoretical and algorithmic foundations of reinforcement learning, through the lens of optimization, modern approximation, and learning theory. It targets students with strong research interests in reinforcement learning, optimization under uncertainty, and data-driven control.
Learning objectives and course logistics · An overview of RL · A primer on optimization
Markov Decision Processes · Bellman equations and Bellman optimality · Value and policy iteration · Linear programming
From planning to reinforcement learning · Model-free prediction · Model-free control · Function approximation · Convergence analysis
Overview of policy-based RL · Policy gradient estimation · Policy gradient methods · Natural policy gradient · Beyond PG: TRPO, PPO
Performance difference lemma · Global convergence of policy gradient methods · Global convergence of natural policy gradient methods · Remarks on sample efficiency
From single agent to multiple agents · Normal form and repeated games · Markov games and algorithms · Zero-sum Markov games and algorithms
Offline imitation learning: behaviour cloning · Online interactive imitation learning: DAGGER, AggreVaTe · Inverse reinforcement learning: feature expectation matching, Max-Ent IRL · Generative adversarial imitation learning (GAIL)
Actor-critic methods · Overview of deep RL · Value-based deep RL · Policy-based and actor-critic deep RL · From deep learning theory to deep RL theory · Convergence analysis of neural TD-learning and neural actor-critic
Model-based RL · Offline RL · Many-agent RL · Summary and outlook
There is no required textbook. Lectures and class discussions are mostly based on classical and recent papers on the topic.

Lecturer

Head teaching assistant

Teaching assistant

Teaching assistant

Teaching assistant

Teaching assistant