June 23, 2026
1,619
70
7
4.76%
Every word spoken in this episode is indexed. Type any phrase to jump straight to the moment it was said.
Type any word or phrase that may have been spoken. Click a result to seek the player to that exact moment.
Try a name, a topic, or a quoted line
49:41Now PlayingThis lecture starts slow, but covers key trends and training methods that came out of advancements in synthetic data. The core of this lecture is the math behind on-policy distillation, the most recent core post-training method to be added to my tool-belt.
Welcome to The RLHF Book & Post-Training Course with Nathan Lambert.
Ask questions and I'll answer them in the next roundup video!
Slides for this lecture are here
Chapters:
All resources will be available at
Order a copy of the book (physical recommended) on Manning.com
Order a copy on Amazon
With specific course resources at (recording links, slides in PDF and native form, etc.)
And code at
Get more information on Nathan at and stay up to date with his work on Interconnects
Course YouTube playlist
Join the book's Discord Community
Nathan is on…
Slides are built with Colloquium
Thank you to my many collaborators who helped me learn this information I get to share with the world!
Sentinel Indexing in Progress
Metadata and chapters are available. Claim extraction for this episode is pending.
All video content is delivered via YouTube embedded players in accordance with the YouTube Terms of Service. Sentinel provides research tools that promote discovery and accountability across political media.