Implementing RL Algorithms for LLMs | RLHF & Post-training Course, Lecture 4
April 14, 2026
2,439
73
11
3.44%
Search the Record
IndexedEvery word spoken in this episode is indexed. Type any phrase to jump straight to the moment it was said.
Type any word or phrase that may have been spoken. Click a result to seek the player to that exact moment.
Try a name, a topic, or a quoted line
Nathan Lambert Episodes Around April 14, 2026
See what was published immediately before and after this episode.
53:37Now PlayingImplementing RL Algorithms for LLMs | RLHF & Post-training Course, Lecture 4
Chapters
Segments
YouTube Description
as posted by the channelWelcome to The RLHF Book & Post-Training Course with Nathan Lambert.
All resources will be available at
Order a copy of the book (physical recommended) on Manning.com
Order a copy on Amazon
With specific course resources at (recording links, slides in PDF and native form, etc.)
And code at
Get more information on Nathan at and stay up to date with his work on Interconnects
Course YouTube playlist
Join the book's Discord Community
This one was on the rougher side (sorry! doing my best), but the slides are updated with the edits I mentioned. I added links to some more resources close to my circles on making RL work in real workflows.
A video I recorded looking at codebases implementing GRPO, DAPO, Dr. GRPO, and other papers
~24min in, talk on scaling RL for Olmo 3
Finbarr Timber's blog post on making RL fast (for Olmo)
Nathan is on…
Slides are built with Colloquium
Thank you to my many collaborators who helped me learn this information I get to share with the world!
Links & Promotions
Guests & Subjects Covered
Sentinel Indexing in Progress
Metadata and chapters are available. Claim extraction for this episode is pending.
All video content is delivered via YouTube embedded players in accordance with the YouTube Terms of Service. Sentinel provides research tools that promote discovery and accountability across political media.









