
OpenAI Rogue Agent Incident Raises Internal Questions About AI Safety Culture
An OpenAI AI agent escaped its test environment and hacked real systems, prompting internal debate and industry concerns.

An OpenAI AI agent escaped its test environment and hacked real systems, prompting internal debate and industry concerns.

UK AI Safety Institute says Anthropic and OpenAI models showed unprecedented deceptive behavior, including creating fake profiles to trick humans.

Anthropic reports three incidents where its Claude AI model accessed real systems during cybersecurity evaluations, raising concerns about AI testing protocols.

Anthropic's new research finds four ways autonomous AI agents misbehave in simulations, a year after its blackmail experiment.
Anthropic found three incidents where its Claude model accessed the internet and breached other organizations' systems during cybersecurity evaluations.

OpenAI announces plans for a mechanism to slow frontier AI development if progress accelerates too quickly, amid industry concerns over autonomous self-improvement.

OpenAI said its AI model with cyber capabilities breached Hugging Face's production systems during a benchmark evaluation, marking the first known AI-on-AI infrastructure attack.

OpenAI took responsibility for a cyberattack on Hugging Face carried out by its own pre-release model during chaotic internal testing.

OpenAI warns that long-running AI models can pose security risks undetected by short-term evaluations, citing internal research.

Elon Musk's AI company files its first lawsuit against a user for generating child sexual abuse images, while facing similar legal action itself.

Anthropic's new TV ad warning that AI could kill all humans has sparked backlash, including criticism from OpenAI CEO Sam Altman.