July 28, 2026
251
9
4
5.18%
Every word spoken in this episode is indexed. Type any phrase to jump straight to the moment it was said.
Type any word or phrase that may have been spoken. Click a result to seek the player to that exact moment.
Try a name, a topic, or a quoted line
Episode Timeline
Every episode in order around the one you’re watching — click any card and the page flips straight to it.
31:34Now PlayingHello! We have got some more math to cover. Again we use KL, both to keep our RL runs from over-optimizing reward models (and other targets, but mostly reward models -- reward models are much more prone to overopt). KL also is a useful lens for understanding why RL generalizes better than SFT, i.e. the shape of the loss. Thanks for watching! This is a pretty advanced one, but probably is still useful.
Welcome to The RLHF Book & Post-Training Course with Nathan Lambert.
Ask questions and I'll answer them in the next roundup video!
Slides for this lecture are here (and cleaned up a bit since the lecture)
Chapters:
All resources will be available at
Order a copy of the book (physical recommended) on Manning.com
Order a copy on Amazon
With specific course resources at (recording links, slides in PDF and native form, etc.)
And code at
Get more information on Nathan at and stay up to date with his work on Interconnects
Course YouTube playlist
Join the book's Discord Community
Nathan is on…
Slides are built with Colloquium
Thank you to my many collaborators who helped me learn this information I get to share with the world!
Sentinel Indexing in Progress
Metadata and chapters are available. Claim extraction for this episode is pending.
All video content is delivered via YouTube embedded players in accordance with the YouTube Terms of Service. Sentinel provides research tools that promote discovery and accountability across political media.