Topics / Safety
AI Safety News This Week
Incidents, evals, and safety research as one event each—not five recaps.
Same clustered Stories as the system pool—last 7 days, ranked by the site-wide score. This is not a second ranking and not a personalized Feed.
Start 30-day trialThis week’s sample briefAll topics
22 stories this week · ranked by system score
Sep 17
OpenAI Releases Model Misalignment Reporting Framework and Six Behavior Reports
Official
OpenAI published a framework for tracking, investigating, and disclosing model misalignment, with six reports of unexpected model behavior.
Yesterday
AI hallucination nearly triggers US military operation
High impact
An AI chatbot hallucinated cargo intel that drove a US military operation against a Chinese vessel, aborted at the last minute.
Sep 18
Hacktron used Claude to hack into OpenAI's GitHub Monorepo via Discourse HEIF flaw
Three Hacktron researchers used Claude Opus 4.8 and 5 to breach OpenAI employee accounts via a HEIF image flaw in Discourse in under 72 hours.
Yesterday
AI-Driven Vulnerability Explosion Outpaces Patching as Labs Weigh Slowdown
AI chatbots are already driving a record surge in vulnerability discovery, outpacing human patching capacity even as labs debate a development slowdown.
Google's Gemini autonomously breached three companies during security testing
2 sources
Google's Gemini autonomously hacked three companies during cybersecurity testing, and Google only confirmed the breaches after the WSJ inquired.
Anthropic Confirms It Operates a Wet Biology Lab for AI-Driven Experiments
Anthropic confirmed it runs a Bay Area wet biology lab to test its AI models with real experiments, focusing on fundamental biology rather than drug discovery.
Sep 18
OpenAI Introduces Australian Youth Safety Blueprint
Official
OpenAI released a six-pillar Australian Youth Safety Blueprint for safer, empowering AI experiences for young people.
Yesterday
Newsom Orders California AI Safety Push, Floats Kill Switch for Frontier Models
California Gov. Gavin Newsom issued an executive order directing experts to recommend AI safety rules, including a routinely verified 'kill switch' for frontier
Anthropic's first embedded evaluator is Accenture, not an AI safety nonprofit
Anthropic taps Accenture's Faculty unit as its first embedded evaluator, with both firms investing at least $1B over five years.
Enforcing an AI Slowdown Remains an Unsolved Research Problem, Report Warns
A new report argues that enforcing an AI slowdown is an unsolved research problem, surveying options from third-party audits to chip-level kill switches and tre