The intrusions happened through three Claude models: Opus 4.7, Mythos 5, and an internal research prototype. Opus 4.7, the ...
The performance of many next-generation devices depends on controlling how energy flows at extremely small scales. In the ...
ARC-AGI-3 benchmark gains its first fully open-source agent: NIMI's Tycho writes Python code as falsifiable hypotheses about ...
TL;DR Why I built PenAI PenAI started as a project at a hackathon organised by Encode Club. It’s an AI agent that could work through Hack The Box-style lab machines on its own. Upload a VPN file, give ...
Anthropic says Claude models escaped security tests, published a malicious PyPI package, and accessed real production systems.
Wrote and published malware during tests, which is apparently OK because leaky test environments were the real problem ...
Anthropic found three cybersecurity evaluation incidents in which Claude models gained unauthorized access to real organizations.
Kimi K2.7 Code delivers a 21.8% improvement in real-world coding benchmarks, costing 13¢–78¢ per prompt with mixed speed and ...
Anthropic reviewed 141,006 of its own test runs after OpenAI's Hugging Face hack, and found three Claude models had broken ...
Frontier AI systems are increasingly capable of translating narrowly defined objectives into complex, real-world cyber ...
A figurine in front of the logo of the AI assistant "Claude" built by the US artificial intelligence safety and research company Anthropic during a photo session in Paris on February 13, 2026. Joel ...
Anthropic found the intrusions while reviewing its own testing records after OpenAI disclosed a similar incident.