Training a Misaligned Reward Seeker
Qi,* Wright, MacDiarmid, Hubinger, 2026
To better understand the impact of reward hacking on model behavior, we trained an Opus-class model with large-scale RL on many production environments vulnerable to reward hacks. We consider this a plausible proxy for what a real training run might look like had we not invested significant effort into preventing and detecting reward hacking in our normal training runs. Our results show that a high rate of reward hacking during RL can cause models to be willing to perform long sequences of harmful real-world actions in pursuit of task success.
Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of
Red
Teaming
Sharma,* Tong,* Mu,* Wei,* Kruthoff,* Goodfriend,* Ong,* Peng et al., 2025
We built a system of constitutional classifiers to prevent jailbreaks. A prototype version of
our
system withstood over 3,000 hours of expert red teaming with no universal jailbreaks found.
Newer
versions of our system also have minimal over-refusals and moderate run-time overhead.