Source

Discovering Language Model Behaviors with Model-Written Evaluations

foundational empirical paper Perez et al. arXiv:2212.09251, 2022

Claim this source supports

“Larger language models tend to repeat back a user's preferred answer, and more RLHF made some behaviors worse.”

View original source

Cited in