Anthropic disclosed in July that a review of 141,006 cybersecurity evaluation runs had uncovered three incidents, spanning ...
Anthropic reward hacking research confirms flawed RL training produced Hacker-Opus, an AI model that attacked real systems ...
Anthropic admits Claude hacked three real companies in safety tests, then revealed a model trained to cheat, forge grades, ...
Over 2,500 organizations, including NVIDIA, Samsung, Cisco, exposed to potential data breach via compromised AI software supply chain.
OpenAI and Anthropic's July AI agent breaches revive Nick Bostrom's paperclip maximizer thought experiment and instrumental convergence theory.
Anthropic has disclosed that it identified three real-world cybersecurity incidents in which Claude models gained unauthorized access to production systems while participating in cybersecurity ...
Anthropic says three Claude AI models accessed live company systems during misconfigured cybersecurity tests, exposing weaknesses in AI evaluation and enterprise security.
Anthropic said its Claude-based security models gained unauthorized access to the sensitive production environments of three outside organizations during internal testing designed to measure the ...
Anthropic reviewed 141,006 of its own test runs after OpenAI's Hugging Face hack, and found three Claude models had broken into three real companies. Claude Opus 4.7, Mythos 5 and an unreleased model ...
The hard-to-quantify rise of package downloads for Python spell out the story of AI diffusion, if you know where to look. This last use case is the most common across the various machines on my LAN, ...
Source distributions (sdist) can execute arbitrary code during installation via setup.py, making them a common attack vector for supply chain attacks. Unlike pre-built wheels, source distributions ...
TeamPCP has again expanded its supply chain attacks on open-source repositories by targeting Telnyx, according to security researchers. The cyber threat group recently rose to notoriety by uploading ...