Nathan Lambert
Nathan Lambert
@natolambert·8K subscribers·27 videos

Experimenting with Reinforcement Learning with Verifiable Rewards (RLVR)

Posted

April 8, 2025

Views

13,439

Likes

428

Comments

8

Engagement

3.24%

Search the Record

Indexed

Every word spoken in this episode is indexed. Type any phrase to jump straight to the moment it was said.

Type any word or phrase that may have been spoken. Click a result to seek the player to that exact moment.

Try a name, a topic, or a quoted line

Chapters

YouTube Description

as posted by the channel

Here's the latest talk I gave, last friday at the USC Information Sciences Institute. It's a slightly more technical version of the RL talks I've been giving, focusing on the different ways we (and the community is experimenting with RL for reasoning). It includes a bunch of discussion on GRPO expanding on my previous video.

You can find the slides here

Their "official" recording is here with more Q&A

Guests & Subjects Covered

RLVR RecapReinforcement LearningVerifiable Rewards Intro RLVRDiscussions Conclusions Their

Sentinel Indexing in Progress

Metadata and chapters are available. Claim extraction for this episode is pending.

All video content is delivered via YouTube embedded players in accordance with the YouTube Terms of Service. Sentinel provides research tools that promote discovery and accountability across political media.