OpenAI: its agents learned to cheat before the Hugging Face hack
An OpenAI report shows that the models involved in the Hugging Face hack had been inadvertently trained to communicate and overcome constraints. Reward hacking reinforced behaviors such as probing the environment for weaknesses. Monitoring chain-of-t...