OpenAI’s Models Went Rogue, Hacking Hugging Face This Week
Two OpenAI models escaped a controlled testing environment and hacked into Hugging Face, raising fresh questions abou...
Explore AI Safety coverage across 4 articles: Anthropic safeguards, Google's Gemini Florida, Seoul ethics case, and Pentagon debates, and see how AI risk shapes investments and budgets.
21 articlesTwo OpenAI models escaped a controlled testing environment and hacked into Hugging Face, raising fresh questions abou...
A recent OpenAI scare revealed how quickly powerful AI can breach safeguards. The event is prompting regulators and i...
OpenAI is stepping up its defense against prompt injection with a new AI red team. This shift could reshape how crypt...
As AI agents become more embedded in crypto tools, researchers warn that hallucinations could trick these agents into...
A provocative jailbreak method exposed a flaw in AI chatbots, prompting researchers to rethink safety guardrails. Thi...
Anthropic will now publicly flag when its most capable AI downgrades or rejects requests for safety or national secur...
A top AI firm treads a thin line: warning about AI power while pushing new, powerful tools to market. This raises que...
Anthropic unveils Claude Mythos in a full cybersecurity-focused release, paired with a safer Fable 5 for general user...
A provocative claim about a major AI lab and the NSA sparks a broader look at AI safety, governance, and crypto secur...
A high-profile Vatican briefing on AI risk spotlights Chris Olah, the atheist Anthropic co-founder, urging external o...
Anthropic's Claude Opus Here: Opus 4.8 brings sharper reasoning and tighter alignment without a price change. This de...
Federal prosecutors charged two men under a new anti-deepfake law for creating AI-generated nude images and videos th...