Introducing RewardBench: The First Benchmark for Reward Models (of the LLM Variety)
March 20, 2024
1,391
49
2
3.67%
Search the Record
IndexedEvery word spoken in this episode is indexed. Type any phrase to jump straight to the moment it was said.
Type any word or phrase that may have been spoken. Click a result to seek the player to that exact moment.
Try a name, a topic, or a quoted line
Nathan Lambert Episodes Around March 20, 2024
See what was published immediately before and after this episode.
16:50Now PlayingIntroducing RewardBench: The First Benchmark for Reward Models (of the LLM Variety)
Chapters
Segments
YouTube Description
as posted by the channelGet to know my latest major project -- we're building the science of LLM alignment one step at a time.
Sorry about the glitchy noise! I didn't think it was so bad that I needed to kill it.
Abstract
Reward models (RMs) are at the crux of successful RLHF to align pretrained models to human preferences, yet there has been relatively little study that focuses on evaluation of those reward models. Evaluating reward models presents an opportunity to understand the opaque technologies used for alignment of language models and which values are embedded in them. To date, very few descriptors of capabilities, training methods, or open-source reward models exist. In this paper, we present REWARDBENCH, a benchmark dataset and code-base for evaluation, to enhance scientific understanding of reward models. The REWARDBENCH dataset is a collection of prompt-win-lose trios spanning chat, reasoning, and safety, to benchmark how reward models perform on challenging, structured and out-of-distribution queries. We created specific comparison datasets for RMs that have subtle, but verifiable reasons (e.g. bugs, incorrect facts) why one answer should be preferred to another. On the REWARDBENCH leaderboard, we evaluate reward models trained with a variety of methods, such as the direct MLE training of classifiers and the implicit reward modeling of Direct Preference Optimization
(DPO), and on a spectrum of datasets. We present many findings on propensity for refusals, reasoning limitations, and instruction following shortcomings of various reward models towards a better understanding of the RLHF process.
Links!
* RewardBench paper (arxiv soon)
* ReardBench Code
* RewardBench Leaderboard
* Interconnects post on Costs vs. Rewards vs. Preferences
* Interconnects post on why we need reward models
* Interconnects post on why we need reward models (p2)
* Paper on history and risks of RLHF
* Talk on history of RLHF
* RewardBench dataset
* Other preference data test sets
* Reward bench results repo
Links & Promotions
Guests & Subjects Covered
Sentinel Indexing in Progress
Metadata and chapters are available. Claim extraction for this episode is pending.
All video content is delivered via YouTube embedded players in accordance with the YouTube Terms of Service. Sentinel provides research tools that promote discovery and accountability across political media.









