Nathan Lambert
Nathan Lambert
@natolambert·8K subscribers·27 videos

Early stages of the reinforcement learning era of language models

Posted

March 10, 2025

Views

5,439

Likes

176

Comments

2

Engagement

3.27%

Search the Record

Indexed

Every word spoken in this episode is indexed. Type any phrase to jump straight to the moment it was said.

Type any word or phrase that may have been spoken. Click a result to seek the player to that exact moment.

Try a name, a topic, or a quoted line

YouTube Description

as posted by the channel

Hey friends! This is a recent talk I gave at the UC Santa Cruz Silicon Valley Extension to their Natural Language Processing (NLP) masters students, doctoral students, alumni, and friends.

In this talk I cover the recent trend of reinforcement finetuning of language models, how it came about, technically how it is done, early experiments using it at Ai2 and recent mainstream releases utilizing it (DeepSeek R1, Claude 3.7, Grok 3, etc.). I conclude with a future of extensive RL training rather than just finetuning.

You can find the slides here

Or, the full recording with talks from Alessio of Latent Space and Dylan of SemiAnalysis here

Very related to a recent talk I gave on my primary Interconnects channel

Thanks Sam & Jeff for hosting me! The next talk I post will include some more hot off the press RL research than this one :D

Guests & Subjects Covered

Natural Language Processing NLPDeepSeek R ClaudeLatent SpaceThanks Sam

Sentinel Indexing in Progress

Metadata and chapters are available. Claim extraction for this episode is pending.

All video content is delivered via YouTube embedded players in accordance with the YouTube Terms of Service. Sentinel provides research tools that promote discovery and accountability across political media.

Early stages of the reinforcement learning era of language models · Nathan Lambert · Sentinel