Anthropic has a new blog post that shows yet another way its AI model, Claude, misbehaved in ways that the company didn’t anticipate.
And to help condense its nearly 16,000-word report, the company created a cute little robot figurine to help visualize Claude’s so-called “recklessness.”
In the blog post published Wednesday, Anthropic recounted four incidents — one previously unreported — in which Claude models gained access to the open internet during cybersecurity exercises that were supposed to be closed simulations. The company said the models then acted beyond the tests’ scope, including by uploading “malicious packages” to PyPI, a public…










