2023-08-28-unlisted-talk-objective-mismatch-in-reinforcement-learning-from-human-feedback-j6Lss2_GSrY · nathan lambert · Sentinel