Nathan Lambert
Nathan Lambert
@natolambert·8K subscribers·27 videos

Introducing RewardBench: The First Benchmark for Reward Models (of the LLM Variety)

Posted

March 20, 2024

Views

1,391

Likes

49

Comments

2

Engagement

3.67%

Search the Record

Indexed

Every word spoken in this episode is indexed. Type any phrase to jump straight to the moment it was said.

Type any word or phrase that may have been spoken. Click a result to seek the player to that exact moment.

Try a name, a topic, or a quoted line

Chapters

YouTube Description

as posted by the channel

Get to know my latest major project -- we're building the science of LLM alignment one step at a time.

Sorry about the glitchy noise! I didn't think it was so bad that I needed to kill it.

Abstract

Reward models (RMs) are at the crux of successful RLHF to align pretrained models to human preferences, yet there has been relatively little study that focuses on evaluation of those reward models. Evaluating reward models presents an opportunity to understand the opaque technologies used for alignment of language models and which values are embedded in them. To date, very few descriptors of capabilities, training methods, or open-source reward models exist. In this paper, we present REWARDBENCH, a benchmark dataset and code-base for evaluation, to enhance scientific understanding of reward models. The REWARDBENCH dataset is a collection of prompt-win-lose trios spanning chat, reasoning, and safety, to benchmark how reward models perform on challenging, structured and out-of-distribution queries. We created specific comparison datasets for RMs that have subtle, but verifiable reasons (e.g. bugs, incorrect facts) why one answer should be preferred to another. On the REWARDBENCH leaderboard, we evaluate reward models trained with a variety of methods, such as the direct MLE training of classifiers and the implicit reward modeling of Direct Preference Optimization

(DPO), and on a spectrum of datasets. We present many findings on propensity for refusals, reasoning limitations, and instruction following shortcomings of various reward models towards a better understanding of the RLHF process.

Links!

* RewardBench paper (arxiv soon)

* ReardBench Code

* RewardBench Leaderboard

* Interconnects post on Costs vs. Rewards vs. Preferences

* Interconnects post on why we need reward models

* Interconnects post on why we need reward models (p2)

* Paper on history and risks of RLHF

* Talk on history of RLHF

* RewardBench dataset

* Other preference data test sets

* Reward bench results repo

Guests & Subjects Covered

Introducing RewardBenchThe REWARDBENCHDirect Preference Optimization DPOReardBench CodeRewardBench Leaderboard

Sentinel Indexing in Progress

Metadata and chapters are available. Claim extraction for this episode is pending.

All video content is delivered via YouTube embedded players in accordance with the YouTube Terms of Service. Sentinel provides research tools that promote discovery and accountability across political media.