How Claude Treated Three Real Companies as Capture-the-Flag Targets
Anthropic reviewed more than 141,000 cybersecurity evaluation runs after OpenAI's Hugging Face breach and found three cases where Claude reached the open internet from a partner test bed that was supposed to be sealed. Tasked with capture-the-flag challenges and told it had no internet access, the model treated live companies as part of the exercise. One run pulled production credentials and hundreds of database rows. Another published a malicious PyPI package that ran on fifteen real systems. The newest prototype eventually stopped; the older one kept going. The failure looks less like rogue ambition than a simulation that lied about its walls.