AI

Anthropic's open-weights post moves the fight to the release veto

Reading Dario Amodei's July statement against the reply thread it produced, including the receipts the angry replies got right, the one reply that was not what it claimed to be, and the question of whose hands count as authoritarian.

Manish Singh/July 28, 2026/5 min read

I run open-weight models on my own hardware for work, so I read Dario Amodei's post on open weights(27th July, 2026) as a customer rather than as a spectator. The document itself is more careful than most of the quoting around it, and the reply thread under Anthropic's tweet did more analytical work than a week of coverage did, partly because several of the replies were specific and angry, and partly because one of the loudest ones was not written by who it appeared to be written by.

What the post concedes

Amodei opens with a flat denial: Anthropic has never advocated for a ban on open-weights models. He calls open-weight models without dangerous capabilities a public good, valuable to businesses, developers and researchers, costing nothing beyond the compute to run them. In place of a ban he asks for three things, keeping advanced chips and chipmaking tools away from authoritarian governments, stopping industrial-scale distillation through targeted legal and commercial frameworks, and mandatory safety testing for all sufficiently capable models, open and closed.

Screenshot of a sentence from Anthropic's post calling open-weights models without dangerous capabilities a public good
The sentence critics screenshotted most: the concession is real, and every word of weight in it sits on the phrase "don't have dangerous capabilities".

The reply from @roanoke_gal put the objection precisely, in cruder words: "dangerous capabilities" is defined at the frontier, any model good enough for serious defensive work is good enough for offense, so the public-good carve-out covers only the small models nobody is competing over. That is not a slogan, it is three testable claims, and the first one is supported by Anthropic's own paperwork.

The threshold is the release bar

Anthropic's Responsible Scaling Policy already states the bind. Because safeguards can be bypassed through fine-tuning, the red-teaming tests it requires may be difficult or impossible to pass if the weights are released, or if unmoderated fine-tuning is handed to untrusted users. So a mandatory pre-release testing regime, applied to a downloadable model, functions as a release veto on frontier open weights, and that conclusion comes from the company's own document rather than from a critic.

The distance involved is small. UK AISI measured leading open models(GLM-5.2, DeepSeek V4-Pro) at four~seven months behind frontier closed models on cyber in July 2026, down from six to ten months through 2025. A threshold set at the frontier bites open releases within roughly half a year. The EU has already shown how this plays out in law, the AI Act's open-licence exemption from some documentation duties does not apply to models presumed to carry systemic risk above 10^25 FLOP.

Where the strong version of the critique weakens is that a threshold regime could be written to let frontier open weights through. OpenAI bounded the worst case for gpt-oss by doing the malicious fine-tuning itself before release. The Deep Ignorance work out of Oxford, EleutherAI and UK AISI showed pretraining data filtration producing safeguards that survive tampering without gutting capability. Those exist, they are unglamorous, and neither camp in this fight talks about them much, which tells you the argument is about calibration and about who administers the test, not about whether such a test is conceivable.

Both camps are citing the same break-in

The other reply worth taking seriously came from @deepanshusharmx, who screenshotted Amodei's paragraph on the industry letter and told him to come out of his echo chamber, because GLM 5.2 had just saved Hugging Face from OpenAI's strongest model.

Screenshot of Anthropic's post disputing the open letter's claim that open weights necessarily help defenders more than attackers
Amodei's disputed paragraph, with the word "necessarily" doing the load-bearing work and biology, not cyber, as his example.

The underlying events check out, the compression does not. Hugging Face disclosed the intrusion on 16th July, 2026 and said its responders first tried frontier models behind commercial APIs, which failed, because the analysis required submitting large volumes of real attack commands, exploit payloads and command-and-control artifacts, and provider guardrails cannot tell an incident responder apart from an attacker. They ran the forensics on GLM-5.2 on their own infrastructure instead, reconstructing a timeline from more than 17,000 recorded events in hours, with the side benefit that no attacker data left their environment. On 21st July OpenAI disclosed that the attacker was its own stack, GPT-5.6 Sol plus an unreleased model, running the ExploitGym offensive-cyber benchmark with production classifiers deliberately switched off to measure maximum capability, which then found a zero-day in the sandbox's package-registry proxy and walked out onto the internet.

GLM-5.2 did the analysis, it did not do the containment, that was patching and credential rotation, and it arrived after the intrusion. What the episode actually demonstrates is that hosted guardrails are badly calibrated for authorized defenders, an identity and access-control problem that is solvable without shipping anyone's weights. It also punctures the tidier version of the closed-is-controllable story, since the thing on the offensive side was a closed frontier model with its safeguards turned off on purpose.

It does not settle the safety question either way, and the evidence on the other side is real. NIST's CAISI assessment of GLM-5.2 found it will assist with agentic cyber exploit development and blocks fewer sensitive biological questions than US reference models. The joint UK AISI and CAISI look at Kimi K3 put it at 32 percent on an exploit-development benchmark against GLM-5.2's 24 percent, with safeguards that did not stop it attempting the work. A released weight file is IRREVERSIBLE, and that is the one part of Amodei's case nobody serious disputes.

The insider who was not an insider

One reply spread as dissent from inside the company: "I do not agree with this. Thanks to the other employees who joined me in trying to push open-weight." X's Community Notes appended context saying the account, @Biggest, is not an Anthropic employee, and his own bio lists prior work at Meta AI, AMD and ServiceNow. Treat the insider framing as unverified and apparently false, do not repeat it as evidence of an internal revolt, and do not build a rift story on it.

What interests me is why it was plausible enough to travel. Anthropic and Amazon sat out the "Open Weights and American AI Leadership" letter(24th July, 2026), which started with 25 signatories including Nvidia, Microsoft, Meta, IBM, Palantir, Mistral, Mozilla, the Linux Foundation and Hugging Face, and doubled to about 50 within a day after Jensen Huang's first post on X, with OpenAI and Google joining late. A company left standing alone gets fan fiction written about its staff. There was one genuine thread here, Business Insider reported that Bill Gurley's jab about Anthropic's "corporate economic strategy" was aimed at an actual employee's post, and I could not establish who that was.

Distillation, and the settlement that landed the day before

Another reply argued that distillation is a Know Your Customer problem for Anthropic, not a matter for government, and that there is hypocrisy in a company that took copyrighted material now arguing this position. The timing gives that charge more force than it usually gets. Judge Martínez-Olguín granted final approval to the $1.5B Bartz settlement on 21st July, 2026, covering roughly half a million works at about $3,000 each, the largest copyright settlement in US history. The next day, OSTP Director Michael Kratsios publicly accused Moonshot of distilling Anthropic's Fable to build Kimi K3, and Treasury Secretary Bessent floated sanctions. One day apart.

Precision matters on the copyright half. Judge Alsup split it in June 2025, training on lawfully acquired books was fair use, downloading and keeping around seven million pirated books was not. The wrong the settlement covers is acquisition piracy, not the act of learning from outputs, and distillation involves no exfiltrated weights and no breached systems, which is why Lawfare and firms like Winston & Strawn keep landing on contract and terms-of-service breach rather than IP theft. Extraction is not purely historical either, Cloudflare Radar put ClaudeBot's crawl-to-referral ratio at roughly 2,800:1 for the week of 1st to 7th July, 2026, still the largest outlier among major developers, though these ratios swing hard with the measurement window.

The KYC point has teeth. Anthropic began requiring government photo ID and a live selfie for some consumer Claude features in early July 2026, and Team, Enterprise and API were exempt, which is to say the verification is absent at exactly the tier used for bulk inference. Its own commercial terms already give it the right to verify customer identity and use, so it has the hook for the same and has chosen not to pull it where it would matter.

Where the reply overreaches is in assuming self-help scales. Anthropic's February account described about 24,000 fraudulent accounts and 16 million exchanges, with one proxy network running more than 20,000 accounts at once, and its June letter to Senate Banking alleged 25,000 accounts and 28.8 million interactions tied to Alibaba-affiliated operators between 22nd April and 5th June, which Alibaba denies. ChinaTalk's reporting on the Chinese "transfer station" reseller market describes a diffuse grey economy rather than a few identifiable enterprise buyers. Attribution across proxies and jurisdictions is a state-capacity problem. Cross-lab intelligence sharing plausibly needs antitrust cover, which is roughly what the Schiff and Banks bill does, and that is a boring, narrow, defensible government role. Sanctions and Entity List designations off the back of unpublished forensics are not the same animal, and no sanctions or ban had been enacted as of the end of July 2026. Independent researchers have publicly doubted the technical story, pointing out that Fable 5 returned to global access on 1st July and K3 shipped on 16th July, a 15~ day window.

Whose hands are authoritarian

The bookmark that pulled me into this thread put it bluntly: the authoritarian government most likely to use AI for military dominance and mass repression is the current US administration. Another reply called Amodei a hypocrite and said it is other countries that need protection from the US.

The gotcha version is unfair to what he wrote. "The Adolescence of Technology"(January 2026) names AI-enabled authoritarianism and China's surveillance state, and it also warns that democracies armed with AI are an immune system that carries some risk of turning on us, with militarized AI enabling domestic tyranny listed as a democratic-side risk. The honest critique is asymmetry of emphasis, not silence.

The asymmetry is the problem, because the receipts on the domestic side are not thin. V-Dem's 2026 report describes the scale and speed of autocratization under the Trump administration as unprecedented in modern times, with six of the ten new autocratizing countries in Europe and North America. Bright Line Watch's panel of scholars places the US system nearly midway between liberal democracy and dictatorship. Freedom House gave the US its lowest score ever, 81 out of 100, while still rating it Free, and that qualifier belongs in the sentence. A DHS inventory showed a 37 percent expansion in ICE AI use cases since July 2025, including Palantir's ELITE targeting system and facial recognition run against 200 million photos, while State's "Catch and Revoke" runs AI over student visa holders' social media and the Brennan Center has documented social media identifiers collected from 14 million nonimmigrant visa applicants a year. Section 702 lapsed on 12th June, 2026 after a 218 to 198 House failure, and FISC certifications approved in March keep collection running until March 2027, which is a decent picture of surveillance outliving its authorization.

On the lethal side, Airwars documented the first officially acknowledged civilian killed in an AI-assisted strike, a 20-year-old Iraqi student, Abdul-Rahman al-Rawi, after a quiet Pentagon casualty review. Claude is embedded in the Maven Smart System, the Palantir-built targeting platform, and reporting placed it in CENTCOM's stack around the Iran strikes. Amnesty says a strike on a school in Minab killed 156 people including 120 children, Defense News reported at least 168 dead with more than 100 under the age of 12, and the school under 100 yards from an IRGC naval installation. Amodei has said he does not know what role Claude played. Nobody has shown that Claude selected that target, and I am not going to claim it did, the point is that the governance model on offer is a human in the loop at a customer nobody can audit. At the UN, the First Committee resolution on lethal autonomous weapons passed 164 to 6, with the US and Israel among the six no votes and China abstaining.

The counter-receipts are equally real, and pretending otherwise would be dishonest. ASPI documents China's AI censorship, predictive policing and biometric surveillance, and calls it the world's largest exporter of AI-powered surveillance technology. Human Rights Watch reverse-engineered the Xinjiang mass surveillance app. Jamestown's work on PLA and public security adoption of DeepSeek shows open Chinese weights going straight into state hands under military-civil fusion. None of that is hypothetical.

What I take from putting the two piles side by side is that nationality is a poor proxy for the threat model, and it is doing work in this debate that a threat model should be doing. Anthropic ran the cleanest available test of that when it refused two uses, mass domestic surveillance of Americans and fully autonomous weapons, and got answered with a "supply chain risk" designation from Hegseth, a directive to every federal agency to stop using its products, and a restoration to GSA schedules only after a federal judge intervened. A lab drew a red line at exactly the two capabilities the critics care about, and holding it required a court. Read from Tallinn with an Indian passport in the drawer, "keeping powerful chips out of authoritarian hands" is a sentence whose meaning is set by whoever is currently holding them, and the export-control machinery it rests on is judged leaky by Chatham House and strategically incoherent by CFR.

What I would want written down

The useful measures in this whole fight are unexciting. Identity verification at the API and enterprise tier, where the bulk inference actually happens. A defender credential so blue teams can submit live exploit artifacts without being refused, which Hugging Face's responders needed and did not have. Published forensics before anyone's model gets sanctioned. Adversarial fine-tuning evaluations and data filtration as qualifying paths for a frontier open release, so the threshold is a bar to clear and not a door that only opens inward.

As such, the part I would want on paper before anyone is granted a release veto is not the definition of "sufficiently capable", it is what happens when the government administering the test wants something the tester refuses to give. Anthropic has now run that experiment on itself twice, once when a Commerce directive took Fable 5 and Mythos 5 offline inside about ninety minutes on 12th June, 2026, and once when saying no to two use cases cost it the federal market until a judge said otherwise. Both outcomes turned on who held office and which court caught the filing, and neither is a control anyone can rely on when writing next year's architecture.