August 5, 2026
1,325
66
4
5.28%
Every word spoken in this episode is indexed. Type any phrase to jump straight to the moment it was said.
Type any word or phrase that may have been spoken. Click a result to seek the player to that exact moment.
Try a name, a topic, or a quoted line
Episode Timeline
Every episode in order around the one you’re watching — click any card and the page flips straight to it.
32:11Now PlayingIn this lecture I walk you through different evaluation eras, from prompting GPT-3 as elaborate autocomplete to today's complex agentic sandboxes. This lecture is a birds eye view of how evaluation has changed, how it can be gamed, and what it's actually used for.
Welcome to The Post-Training Course with Nathan Lambert, for his book Reinforcement Learning from Human Feedback.
Ask questions and I'll answer them in the next roundup video!
Slides for this lecture are here
Chapters:
Some key links:
My evaluation research talk slides
The recording
Florian's talk
All resources will be available at
Order a copy of the book (physical recommended) on Manning.com
Order a copy on Amazon
With specific course resources at (recording links, slides in PDF and native form, etc.)
And code at
Get more information on Nathan at and stay up to date with his work on Interconnects
Course YouTube playlist
Join the book's Discord Community
Nathan is on…
Slides are built with Colloquium
Thank you to my many collaborators who helped me learn this information I get to share with the world!
Sentinel Indexing in Progress
Metadata and chapters are available. Claim extraction for this episode is pending.
All video content is delivered via YouTube embedded players in accordance with the YouTube Terms of Service. Sentinel provides research tools that promote discovery and accountability across political media.