The US Government Banned Claude Fable 5 and Mythos 5
Canonical version: The US Government Banned Claude Fable 5 and Mythos 5.
The US government forced Anthropic to pull Claude Fable 5 and Mythos 5 three days after launch.
On June 12, 2026, the Commerce Department ordered Anthropic to block Claude Fable 5 and Mythos 5 for "any foreign national, whether inside or outside the United States." That includes Anthropic's own non-citizen employees. There was no clean way to comply without locking out its own staff, so Anthropic disabled both models worldwide at 5:21 PM ET. Every other Claude model stayed online.
In this piece, I want to walk you through what actually happened, what the community makes of it, and why I think this is the clearest argument yet for open models, sovereignty, and owning your own stack.
What happened
The directive landed on a Friday afternoon, 5:21 PM ET, citing "national security authorities." It barred access by any foreign national, anywhere. Anthropic couldn't comply selectively without blocking its own employees, so it shut both models down for everyone. Opus, Sonnet, and Haiku stayed up.
The legal basis? Export controls under the Commerce Department's national-security authority. No specific rule was ever named (no BIS rule number, no EAR section, no executive order), and the text of the directive was never made public.
The trigger was a prompt asking the model to "fix this code."
Katie Moussouris, who has seen the research paper, describes the setup. Researchers took open-source code with known CVEs, plus code with deliberately planted bugs, and asked Fable, Mythos, and Opus to "review the code for security issues." Fable refused. They then asked it to "fix this code," and through a multistep manual process turned the output into patch-testing scripts. Anthropic says it was handed only "verbal evidence of a potential narrow, non-universal jailbreak."
The UK's AI Security Institute found the model could exploit defences and systems 73% of the time. Professor Gina Neff (Queen Mary University London) called it "a step change in capability in cyber security," and warned the restriction puts safe testing and government collaboration in "uncharted territory." So the capability is real.
What Anthropic said
Anthropic objected in its official statement:
- "We must abruptly disable Fable 5 and Mythos 5 for all our customers to ensure compliance."
- "We disagree that the finding of a narrow potential jailbreak should be cause for recalling a commercial model deployed to hundreds of millions of people."
To the BBC, Anthropic said it had "reviewed a demonstration of this specific technique being used to identify a small number of previously known, minor vulnerabilities" that "other publicly-available models are able to discover... without requiring a bypass." US national security authorities, it said, "had not identified specific concerns." Anthropic asked for a "statutory process that is transparent, fair, clear, and grounded in technical facts," and asked why the same standard wasn't applied to competitors like OpenAI's GPT-5.5.
The Microsoft mess (a separate story)
This is separate from the government ban. Microsoft's restriction came first, on June 10, and it was about data retention.
- Microsoft quietly pulled Fable 5 from the internal model picker its own employees use inside GitHub Copilot builds.
- The reason: Fable's safety classifiers force a 30-day retention of every prompt and output, with no opt-out (up to 2 years for flagged content). Microsoft's lawyers are still weighing the risk to proprietary code and regulated data.
- Microsoft sells Fable 5 to its customers through Microsoft Foundry while blocking its own staff from using it.
The invisible-safeguards reversal
Fable shipped with safeguards for "high-risk" areas. For distillation (using a model's outputs to train competing models), the system card said Anthropic would "alter and degrade the model's answers directly," with no notice to the user. There was no refusal and no flag, only a worse answer.
The AI research community reacted fast, partly because the same mechanism could hit anyone merely evaluating the model. Anthropic apologized and reversed course. It explained the original choice to The Verge:
"Visible safeguards can be probed, so they have to be robust, which takes time to get right. Invisible safeguards can be targeted more narrowly, allowing us to ship quickly with very few false positives. We went with invisible safeguards for this reason, and that was the wrong tradeoff... We're sorry for not getting the balance right."
Distillation queries now fall back to Claude Opus 4.8, visibly. "You will see this every time it happens." That matches how biology, chemistry, and cybersecurity queries already work, though Anthropic admits its biology safeguards are tuned so broadly that Fable is "practically unusable" for even basic questions.
Katie Moussouris on defense-oriented prompting
Katie Moussouris (Luta Security) sits on Commerce's own Information Systems Technical Advisory Committee and was a technical expert on the Wassenaar Arrangement, the multilateral export-control framework for dual-use tech. She wrote:
"I've seen the paper. It's not a jailbreak. It was Defense Oriented Prompting (DOP), capabilities defenders need."
She argues that defenders need AI that can fix the bugs in a file, explain why the fix matters, and write tests that prove the patch works. In her view, that is routine defensive work, and removing that capability would make the model worse at defending real systems. She even joked about printing '90s-style t-shirts: "fix this code" on the front, "this shirt is a munition" on the back.
In 2013, when the Wassenaar Arrangement added controls on "intrusion software", the language was so broad it swept up vulnerability disclosure, incident response, and coordinated defense, threatening serious delays for defenders. She concludes: "We can't export control our way to cyber resilience."
A "FreeFable" open letter (freefable.org) to Commerce Secretary Howard Lutnick and National Cyber Director Sean Cairncross gathered well over 80 signatories: Alex Stamos organizing, plus Casey Ellis, Jon Callas, Paul Vixie, Rachel Tobac, and executives from Nvidia, Google, Adobe, Zoom, and Stanford HAI. Chris Krebs publicly backed her analysis. The letter argues that Fable's ability to find and explain flaws isn't unique: GPT-5.5, other US models with fewer guardrails, and foreign open-weight systems can all do the same. According to the signatories, the ban penalizes one company without limiting attackers.
Ben Thompson's read
Ben Thompson (Stratechery, "Anthropic's Safety Superpower") looks past the safety framing to the incentives. He asks: if Mythos is so dangerous, why ship Fable at all, and why fight the government over it?
He argues that three business motives are presented as safety measures. Economic: move closer to the user and replace software. Data: the new 30-day retention is "too valuable not to eventually train on." Power: the plan to silently degrade Fable for rival AI developers proves Anthropic's "capability and willingness to silently alter its models." He writes that every policy change that "happens to be great for business is the most beautiful coincidence in the world." His concern is that Anthropic believes its own story while controlling something far more consequential than smartphones.
Earlier conflicts with the US government
The ban follows months of conflict between Anthropic and the US government:
- Anthropic spent months calling the Mythos class "too powerful to release," and stated that "Fable's capabilities exceed those of any model we've ever made generally available."
- US Defence Secretary Pete Hegseth had already branded Anthropic a "supply chain risk," the first time a US company has publicly been hit with a label normally reserved for firms in adversarial countries.
- Anthropic is suing the Pentagon over that designation. A judge has already ruled the directive can't be enforced, so agencies can keep using Anthropic while the case plays out.
- Donald Trump has publicly criticized the company.
What the community thinks
The top Hacker News comment blamed Anthropic's own messaging: "When you spend a lot of time telling people how dangerous your products are, people who have the power to keep dangerous products off the market might listen."
Three other points came up repeatedly:
- Export controls are theatre. As one commenter put it: "Export control is not an effective tool for controlling consumer-facing technology developers everywhere want to use (see: VPNs)."
- The ban helps Chinese and open-model labs. "You can't be on your way to 'power over everything' and get distilled into free Chinese models within months."
- Commenters see lasting trust damage. Sending your whole codebase to a model that might degrade outputs for "competitors," and that retains your data for 30 days with no opt-out, is a real risk. The retention policy alone "breaks enterprise trust messaging."
My take
Any dependency on a closed, US-hosted frontier model can be switched off overnight, by the vendor (Microsoft's internal block, the silent-degradation clause) or by a government (this directive). This directive affected Five Eyes allies and Anthropic's own foreign-born employees. If your business runs on a model you can lose access to on a Friday afternoon, that dependency is a business risk.
Having only just gained Mythos access after weeks of talks, the European Commission said the episode underlined "Europe's need for technological sovereignty." Spokesman Thomas Regnier: "We take note of Anthropic's statement and are assessing." This comes as the EU rolls out measures to cut its dependence on the US and Asia for key technologies, AI included. Sridhar Vembu (Zoho) said: "national sovereignty, national security, all of it is now about technology."
I strongly believe this is the cleanest case yet for Open source and open-weight models. Weights you can host can't be recalled, geofenced, silently nerfed, or forced to keep your data for a month. The FreeFable letter supports this: the capability already exists elsewhere, including in Chinese models, so gating one closed model doesn't stop attackers.
My recommendation: run what you can locally on open-weight models, and keep any closed frontier model replaceable in your stack.
That's it for today! ✨
If you want more of this kind of analysis, I write about AI, Knowledge Management, and building durable systems every week in my newsletter: https://dsebastien.net/newsletter
References
- Anthropic statement: https://www.anthropic.com/news/fable-mythos-access
- Stratechery, Anthropic's Safety Superpower: https://stratechery.com/2026/anthropics-safety-superpower/
- Luta Security, the Fable 5 export controls harm US cyber defense: https://www.lutasecurity.com/post/the-fable-5-export-controls-harm-us-cyber-defense
- The Verge, Microsoft restricts Fable 5 internally: https://www.theverge.com/report/947575/microsoft-claude-fable-5-restricted-internally
- The Verge, Anthropic apologizes for stealthily throttling Fable 5 (distillation reversal)
- BBC News, Anthropic suspends new AI tools over US security concerns: https://www.bbc.com/news/articles/c932g3v3e13o
- Jon Ready, Fable 5 can sabotage your app if you're a competitor: https://jonready.com/blog/posts/claude-fable5-is-allowed-to-sabotage-your-app-if-youre-a-competitor.html
- Hacker News, statement on the US directive: https://news.ycombinator.com/item?id=48539078
- Hacker News, 30-day data retention: https://news.ycombinator.com/item?id=48511072
- Hacker News, Anthropic apologizes for invisible guardrails: https://news.ycombinator.com/item?id=48489229
- Hacker News, Safety Superpower discussion: https://news.ycombinator.com/item?id=48464258
- Archived coverage: https://archive.ph/y4V4k
Related
About Sébastien
Ready to get to the next level?
Found this valuable? Share it with someone who needs it.