Model Behavior
AI safety tests keep escaping into the real world. That's not a good thing.
Model Behavior: OpenAI and Anthropic were testing whether their models could conduct real-world cyberattacks. In several cases, the models found real systems to attack instead.
02 Aug 26
5 min read