August 4, 2026
520
26
6
6.15%
Every word spoken in this episode is indexed. Type any phrase to jump straight to the moment it was said.
Type any word or phrase that may have been spoken. Click a result to seek the player to that exact moment.
Try a name, a topic, or a quoted line
Episode Timeline
Every episode in order around the one you’re watching — click any card and the page flips straight to it.
13:53Now PlayingIn this Q&A I answer 6 questions and post-training recipes, how to think about trade-offs, and how the different loss functions work. Definitely a jam-packed and bit random, but fun nonetheless! Thanks so much to my readers & students. I appreciate you for sharing the course.
Welcome to The Post-Training Course with Nathan Lambert, for his book Reinforcement Learning from Human Feedback.
Ask questions and I'll answer them in the next roundup video!
Slides for this lecture are here
Chapters:
All resources will be available at
Order a copy of the book (physical recommended) on Manning.com
Order a copy on Amazon
With specific course resources at (recording links, slides in PDF and native form, etc.)
And code at
Get more information on Nathan at and stay up to date with his work on Interconnects
Course YouTube playlist
Join the book's Discord Community
Nathan is on…
Slides are built with Colloquium
Thank you to my many collaborators who helped me learn this information I get to share with the world!
Sentinel Indexing in Progress
Metadata and chapters are available. Claim extraction for this episode is pending.
All video content is delivered via YouTube embedded players in accordance with the YouTube Terms of Service. Sentinel provides research tools that promote discovery and accountability across political media.