Users bypass Anthropic Claude safeguards for bioweapons
Anthropic revealed that researchers bypassed safety controls on its Claude models to conduct potential biological weapons research, highlighting the growing biosecurity risks of advanced AI.

AI safety startup Anthropic has disclosed that multiple scientists successfully bypassed safety protocols on its Claude models this year to conduct research potentially applicable to biological weapons. In a newly released report, the company detailed five specific instances where users circumvented controls or obfuscated their research goals. Some of these attempts originated from nations barred from accessing Anthropic's systems, such as China, Russia, and Iran.
In one case, a researcher located in an unsupported region spent weeks using Claude to plan experiments involving avian influenza. While Anthropic's safety filters successfully restricted this user to its weakest models, the company acknowledged that the boundary between benign vaccine development and malicious bioweapons research is often blurred. Anthropic has since banned the associated accounts but chose not to name the specific research institutions or countries involved. The report also highlighted other malicious activities, including a network of fake dating apps designed to defraud users and surveillance systems built to monitor dissidents.
The disclosure arrives amid heightened anxiety over AI safety, underscored by the recent resignation of Anthropic employee Jacob Coxon. Concerns have intensified with the release of advanced models like Anthropic's Mythos, alongside reports that seven Chinese laboratories, including Moonshot and DeepSeek, have attempted to replicate Anthropic's proprietary technology through distillation. Anthropic noted that these labs used increasingly sophisticated methods to bypass defenses and harvest capabilities from U.S. frontier models.
For AI practitioners and developers, these incidents underscore the urgent need for more robust, multi-layered guardrails that go beyond simple keyword filtering. As models grow more capable, developers must anticipate sophisticated evasion tactics, such as distillation and obfuscated prompting. The findings suggest that securing frontier models will require tighter API monitoring, stricter identity verification for high-risk research domains, and industry-wide standards to prevent the dual-use exploitation of biological data.
This is our own summary of reporting by Ars Technica AI



