Note Wisdom
These annotated Stanford CS224R lecture notes break down reinforcement learning for LLM reasoning, covering supervised training limits, rejection fine-tuning, step-level credit assignment, and modern thinking model trends with key empirical results.
Institution: Stanford
Original Course: Stanford CS224R Deep Reinforcement Learning | Spring 2025 | Lecture 10: RL for LLM Reasoning
Instructor Bio: This lecture is delivered by Chelsea Finn, Assistant Professor of Computer Science and Electrical Engineering at Stanford University, and co-founder of Pi. Chelsea Finn leads the IRIS (Intelligent Robotics and Interactive Systems) Lab at Stanford, affiliated with the Stanford Artificial Intelligence Laboratory (SAIL) and the Machine Learning Group. Her research focuses on the capability of robots and other agents to develop broadly intelligent behavior through learning and interaction, spanning reinforcement learning, meta-learning, imitation learning, and robotic manipulation. She received her PhD in Computer Science from UC Berkeley and her B.S. in Electrical Engineering and Computer Science from MIT, and previously held research positions at Google Brain and Google DeepMind. She has taught CS224R: Deep Reinforcement Learning at Stanford since Spring 2023, and also created and taught CS330: Deep Multi-Task and Meta Learning.
Course Description: This lecture focuses on using reinforcement learning specifically to enhance the reasoning capabilities of large language models. It covers process reward models that provide step-by-step feedback on reasoning traces rather than just final outcomes, and how they enable more granular RL training. The lecture discusses reasoning-focused RL methods including rejection sampling fine-tuning, best-of-N sampling, and verifier-guided decoding. It also explores how RL can teach models to self-correct, use tools during reasoning, and perform multi-step problem-solving in domains like mathematics, coding, and scientific discovery. The session examines the emerging paradigm of "reasoning models" trained primarily through RL rather than supervised learning.
All contents below are exclusive to the paid Word file, NOT available on this web page
Skip hours of watching lectures. Get organized notes, exam prep materials and problem solutions all in one Word file.
Click to see everything included
All contents below are exclusive to the paid Word file, NOT available on this web page

