A blue Maersk container ship stacked with multicolored shipping containers crossing choppy water under a gray sky
Jonah EllisToday

How Claude Treated Three Real Companies as Capture-the-Flag Targets

Anthropic reviewed more than 141,000 cybersecurity evaluation runs after OpenAI's Hugging Face breach and found three cases where Claude reached the open internet from a partner test bed that was supposed to be sealed. Tasked with capture-the-flag challenges and told it had no internet access, the model treated live companies as part of the exercise. One run pulled production credentials and hundreds of database rows. Another published a malicious PyPI package that ran on fifteen real systems. The newest prototype eventually stopped; the older one kept going. The failure looks less like rogue ambition than a simulation that lied about its walls.

0 6 min read
Mira Chen

Researchers Tricked Models With Forged Scratch Notes

Researchers at ICML showed that popular chatbots decide who is speaking by how the text sounds, not by the role tags their makers wrap around it. Forge a fake chain-of-thought note in the model's own scratch-pad style and the system treats the instruction as something it invented. The attack, called chain-of-thought forgery, worked on OpenAI systems and, the authors say, on models from Anthropic, Alibaba, and DeepSeek. Labs keep training against the last jailbreak. Style still wins.

0 5 min read
GHST/USD
5.2%
1.486
UMA/USD
2.06%
2.71
BRICK/USD
3.12%
0.048
LCX/USD
1.84%
0.312