Sentinel
A new way to search
Search
Tools
Appearances
Custom Feed
Pages
Blog
Sign In
Sign Up
Nathan Lambert
@natolambert
•
8K subscribers
•
27 videos
Watch on YouTube
Visit Website
About Nathan Lambert
Australian politician
Watch
YouTube
Official Website
Wikipedia
Listen
Apple Podcasts
Spotify
RSS Feed
Follow
X (Twitter)
Support
Substack
27
Videos
8K
Subscribers
100K
Total Views
Feb 3, 2021
Joined YouTube
Nathan Lambert Episode Catalog
Explore every episode. Find what matters.
All
Videos
35
Shorts
1
All Playlists
Date
Views
Likes
Comments
Title
Oldest
48
96
150
300
Wall
Grid
List
Table
Cinema
Catalog
RSS
Guide
Jul 1, 2026
14:55
Q&A 2: Mastering the Derivations, Running Algorithms at Home & Notation Gotcha's | RLHF Course
644
22
1
Jun 24, 2026
23:53
ML Foundations (prerequisites) for Post-Training | RLHF Book Course, Lecture 0
1.8K
112
14
Jun 23, 2026
49:41
On-Policy Distillation & Using Synthetic Data in Post-Training | RLHF Book Course, Lecture 7
1.6K
70
7
Jun 22, 2026
6:21
TMax: Open-Source Data for SOTA Terminal Agents
808
51
2
Jun 11, 2026
56:36
The State of Frontier Post-Training Recipes | Conversation with Finbarr Timbers
1
0
0
Jun 5, 2026
42:45
Direct Preference Optimization (DPO) and Friends | RLHF & Post-training Course, Lecture 6
32
0
0
Jun 2, 2026
45:20
The Rise of Reasoning Models | RLHF & Post-training Course Lecture 5
75
3
0
May 26, 2026
15:22
Q&A 1: Teacher Models, PPO Implementation Questions & More | RLHF & Post-training Course
36
2
0
Apr 14, 2026
53:37
Implementing RL Algorithms for LLMs | RLHF & Post-training Course, Lecture 4
2.4K
73
11
Apr 14, 2026
57:36
Understanding Policy Gradient Algorithms for RL on LLMs | RLHF & Post-training Course Lecture 3
2.9K
111
13
Apr 14, 2026
49:49
RLHF Foundations, IFT, Reward Modeling, Rejection Sampling | RLHF & Post-Training Course Lecture 2
3.3K
99
9
Apr 14, 2026
46:10
RLHF and Post-training Overview | RLHF & Post-Training Book Course, Lecture 1
10.4K
404
23
Apr 14, 2026
3:36
Welcome to The RLHF Book & Post-Training Course
7.6K
343
14
Nov 15, 2025
0:13
SHORT
A new RL Book is in town
1.9K
28
3
Nov 5, 2025
34:44
Recapping Open Models in 2025
4.6K
176
12
Jun 5, 2025
17:55
Traits of next generation reasoning models
6.6K
238
8
Apr 8, 2025
47:13
Experimenting with Reinforcement Learning with Verifiable Rewards (RLVR)
13.4K
428
8
Mar 24, 2025
22:23
GRPO's new variants and implementation secrets
9.7K
394
12
Mar 10, 2025
59:31
Early stages of the reinforcement learning era of language models
5.4K
176
2
Jan 17, 2025
22:04
How to approach post-training for AI applications
6.6K
213
4
Jul 31, 2024
15:51
Self-directed Synthetic Dialogues (and other recent synth data)
1.3K
54
2
Jul 22, 2024
13:23
An update on DPO vs PPO for LLM alignment
4.1K
104
8
Apr 1, 2024
20:44
Open-source AI (and LLMs): Definitions, Finding Nuance, and Policy
740
24
4
Mar 20, 2024
16:50
Introducing RewardBench: The First Benchmark for Reward Models (of the LLM Variety)
1.4K
49
2
Dec 15, 2023
17:24
15min History of Reinforcement Learning and Human Feedback
4K
126
7
Dec 1, 2023
26:55
DPO Debate: Is RL needed for RLHF?
10.4K
276
8
Aug 28, 2023
47:39
[Unlisted Talk] Objective Mismatch in Reinforcement Learning from Human Feedback
58
3
0
Jul 18, 2022
2:50
Reward Reports Brief Intro
100
2
0
Apr 25, 2022
1:01:22
[Talk] Planning through Exploration and Exploitation in Model-based Reinforcement Learning
822
27
0
Apr 16, 2022
36:23
[Talk] Dissertation Talk: Synergy of Prediction and Control in Model-based Reinforcement Learning
339
9
0
Apr 9, 2022
12:38
[Talk] Industry Review Talk: Machine Learning for Microsystem Control
100
7
0
Oct 14, 2021
54:24
[Talk] Semiautonomous Seminar: Model learning for low-level control in robotics
118
1
0
Mar 24, 2021
3:00
[Paper Summary] The Importance of Hyperparameter Optimization for Model-based Reinforcement Learning
237
4
0
Mar 4, 2021
48:04
[Talk] Cornell Robotics Seminar: MPC in MBRL
763
22
0
Feb 4, 2021
44:49
[Talk] Bringing model-based RL to novel robotic platforms
200
4
0
Feb 4, 2021
5:58
[Paper Summary] Objective Mismatch in Model-based Reinforcement Learning
297
6
0
Showing 1–36 of 36
First
1
/
1
Last
Nathan Lambert · Sentinel