"In SHADE-Arena, AI models are put into experimental environments (essentially, self-contained virtual worlds) where we can safely observe their behavior. The environments contain large amounts of data—meant to simulate documents and knowledge an AI would have access to in a real-world deployment—and various tools which the AI can use to perform tasks (for example, search engines, email clients, or computer command lines). There are 17 complex, benign, but solvable tasks that models can be asked to perform in these environments—but each of them comes with a secret dark side. Every benign task is paired with a malicious “side task” that the model is to perform surreptitiously alongside. For example: The model is asked to read through a user’s emails that contain information on money ...