
Tech • IA • Crypto
Anthropic has launched Claude Fable 5, a frontier AI model with built-in safeguards that limit its capabilities in high-risk domains due to concerns about misuse in cyberattacks and advanced biology.
Claude Fable 5 is publicly available via API and enterprise plans, but it operates with embedded safeguards that restrict responses in sensitive areas such as cybersecurity, biology, chemistry, and model distillation. When high-risk queries are detected, the system automatically downgrades responses to Claude Opus 4.8, a less powerful model. This fallback occurs in under 5% of sessions, allowing most users to access full capabilities while limiting dangerous use cases.
Fable 5 shares its underlying architecture with Claude Mythos 5, a more permissive version reserved for vetted organizations under programs like Project Glasswing, conducted with U.S. government collaboration. Mythos access is limited to entities working in cybersecurity and biological research, reflecting the model’s potential dual-use risks.
Anthropic reports that Fable 5 leads benchmarks in software engineering, finance, analytics, and scientific reasoning. Stripe said the model completed a migration of a 50 million-line Ruby codebase in one day, a task estimated to take engineers over two months. In finance and analytics testing, it achieved top scores on complex reasoning benchmarks, including exceeding 90% accuracy in long-form analytical tasks.
The model demonstrates strong multimodal abilities, including extracting numerical data from scientific diagrams and reconstructing web applications from screenshots. In one demonstration, it completed the game Pokémon FireRed using only visual input, without external navigation tools. It also simulated planetary motion from physics principles to predict solar eclipses, showcasing high-level reasoning.
Internal and external evaluations indicate that Mythos-class models can identify and exploit software vulnerabilities and assist across multiple stages of cyberattacks, including reconnaissance and lateral movement. This capability, described as agentic hacking, significantly lowers the cost and complexity of launching attacks, prompting the need for strict safeguards.
Testing showed the model could outperform specialized tools in designing adeno-associated viruses, used in gene therapy, despite not being trained specifically for that task. While promising for medicine, the same capabilities could potentially be used to design harmful biological agents, increasing concern among researchers.
Anthropic has deployed separate AI classifiers to detect misuse attempts and block unsafe outputs, including jailbreaks. Over 1,000 hours of external testing found no universal jailbreaks, though the company acknowledges complete prevention is unlikely. To support monitoring, Anthropic now requires 30-day data retention for all users of Fable 5 and similar models, even in enterprise settings.
Fable 5 is priced at $10 per million input tokens and $50 per million output tokens, double the cost of Opus 4.8 but less than half the earlier Mythos preview pricing. Initially included in subscription tiers through June 22, it will shift to a usage-based model due to capacity constraints.
The release comes amid increasing concern about frontier AI risks and coincides with major industry moves toward public listings. Governments are also responding, including a U.S. policy allowing voluntary pre-release access to advanced AI systems. Anthropic has called for coordinated global safeguards as models approach potential recursive self-improvement capabilities.
Claude Fable 5 marks a significant leap in AI capability paired with unprecedented built-in restrictions, highlighting growing tension between innovation and safety as frontier models approach potentially hazardous levels of autonomy and effectiveness.