Early stages of the reinforcement learning era of language models
March 10, 2025
5,439
176
2
3.27%
Search the Record
IndexedEvery word spoken in this episode is indexed. Type any phrase to jump straight to the moment it was said.
Type any word or phrase that may have been spoken. Click a result to seek the player to that exact moment.
Try a name, a topic, or a quoted line
Nathan Lambert Episodes Around March 10, 2025
See what was published immediately before and after this episode.
59:31Now PlayingEarly stages of the reinforcement learning era of language models
Segments
YouTube Description
as posted by the channelHey friends! This is a recent talk I gave at the UC Santa Cruz Silicon Valley Extension to their Natural Language Processing (NLP) masters students, doctoral students, alumni, and friends.
In this talk I cover the recent trend of reinforcement finetuning of language models, how it came about, technically how it is done, early experiments using it at Ai2 and recent mainstream releases utilizing it (DeepSeek R1, Claude 3.7, Grok 3, etc.). I conclude with a future of extensive RL training rather than just finetuning.
You can find the slides here
Or, the full recording with talks from Alessio of Latent Space and Dylan of SemiAnalysis here
Very related to a recent talk I gave on my primary Interconnects channel
Thanks Sam & Jeff for hosting me! The next talk I post will include some more hot off the press RL research than this one :D
Links & Promotions
Guests & Subjects Covered
Sentinel Indexing in Progress
Metadata and chapters are available. Claim extraction for this episode is pending.
All video content is delivered via YouTube embedded players in accordance with the YouTube Terms of Service. Sentinel provides research tools that promote discovery and accountability across political media.









