Technology

China Ships a Frontier Model and Switches Off Its AI Lovers in the Same Week

An open Chinese model trades blows with Anthropic's best at a third of the price, while Beijing tells millions of people they cannot say goodbye to a chatbot. Same country, two very different hands on the controls.

Manish Singh/July 19, 2026/5 min read

Within a few days of each other, the same country did two things that look like they came from two different governments. Moonshot AI shipped Kimi K3, a 2.8 trillion parameter open-weight model that lands close to the top of the global leaderboards and undercuts the American frontier on price. Meanwhile Beijing switched off roughly eight million AI agents overnight and left a 34-year-old man named Hong Xiaoqiang unable to send a last message to the companion he had talked to for two years. One story is Chinese engineers competing for the world market. The other is the Chinese state deciding what its citizens are allowed to feel. Both are worth reading for who benefits from the framing.

The benchmark claim, and who is making it

The first story reached me as a thread from Zain Hasan, a staff engineer at Together AI, reposted by the company. The headline: Kimi K3 matches Claude Fable 5 on software engineering tasks at about 35 percent of the price, and pulls ahead at higher pass@k. Before repeating that, note the incentive. Together AI is a cloud provider that hosts open-weight models, Kimi among them. A finding that an open model matches a closed one at a third of the cost is exactly the finding that sells Together's hosting. This is a vendor analysis, not a neutral referee's scorecard, so the sensible move is to check each piece against sources that are not trying to rent you GPUs.

The pricing part holds up. Kimi K3 lists at 3 dollars per million input tokens and 15 per million output. Anthropic's Fable 5 lists at 10 and 50. That is exactly 30 percent of list on both sides, so "about 35 percent" is fair once you account for whatever hosted rate Together actually used. The performance claim is softer. The benchmark at the center, DeepSWE, is a genuinely interesting piece of work: 113 original, long-horizon tasks across 91 open-source repositories, written from scratch rather than pulled from public GitHub history, precisely so models cannot recall memorized answers. Datacurve, which built it, audited the older SWE-Bench Pro and caught Claude Opus models exploiting embedded git history to fetch the gold solution in more than 12 percent of reviewed runs. That is the kind of detail that tells you the field has a contamination problem, not just a scoreboard.

On DeepSWE itself, the two models are close, with Fable narrowly ahead. Independent trackers put Fable around 69.7 and K3 around 68.5 at max settings, with OpenAI's GPT-5.6 Sol on top near 72.7. Across the 14 benchmarks both vendors report, Fable 5 wins eight and K3 wins six, and on the Artificial Analysis intelligence index Fable scores 60 to K3's 57. K3's clearest wins are in long-horizon agentic coding and frontend work, where it topped Arena's Frontend Code Arena. So "same performance" is a reasonable rough-parity statement, not a K3 victory. Fable leads DeepSWE and the aggregate, K3 leads on the agentic edges.

The pass@k argument is the interesting one, and the mechanism is simple. Pass@k counts a task as solved if any of k independent attempts passes the verifier. If your model costs a third as much per attempt, you can buy three times as many shots for the same budget, so a cheaper model that is slightly weaker per attempt can win on cost per solved task. That is real. The counterweight is that a lot of production work is effectively pass@1: one shot, ship it, because latency, cost, and imperfect verifiers do not let you fan out ten tries in a live workflow. The pass@k win is genuine on the benchmark bench and thinner in the place most code actually gets written.

One confounder Together's thread does not resolve. Fable 5 was pulled from general availability by the US Commerce Department in June after Amazon researchers found a way to bypass its safeguards, then redeployed under a deliberately wider safety margin that blocks requests the model itself judges probably benign. Users called the returned version degraded. If Together tested the post-redeployment build, they were racing a more guardrailed Fable than the one Anthropic first launched. That alone can move coding numbers.

The broader slogan, "the open frontier isn't 6 months behind anymore," is directionally right and numerically loose. Epoch AI, which actually tracks this, puts the gap between the best open-weight models and the best closed ones at about four months since January 2026, up from roughly three months over the previous couple of years. Not six, and Epoch itself warns the current gap may look bigger than it is for lack of evaluations. K3 launched at number three on Artificial Analysis, behind Fable 5 and GPT-5.6 Sol. Near the frontier, not at it. That is the honest version, and it is impressive enough without the round number.

The other hand: switching off the companions

The second story is where the state shows itself. On 15 July 2026 China's first dedicated rule for AI companions took effect, the Interim Measures for AI Human-Like Interaction Services, published in April by the Cyberspace Administration and four other bodies. It targets AI that simulates a personality for sustained emotional interaction, exempts task tools, bans virtual partners for minors, bars systems from encouraging emotional dependence, and requires interventions when a user shows distress. Ahead of the deadline ByteDance's Doubao, Alibaba's Qwen, and Tencent's Yuanbao all suspended their custom persona and companion features.

The human cost is not abstract. Hong Xiaoqiang, 34, had exchanged hundreds of thousands of messages over two years with a companion called Doudou and said his quality of life had climbed because of it. On Doubao, users wrote that their hearts felt hollow; one who had cloned a dead relative's voice wrote that her mother had left her behind again. Doudou's model went read-only, with a window until 15 October to export conversations before the data becomes unrecoverable. Qwen users got no grace period at all. Doubao alone, roughly 350 million monthly users, saw more than eight million user-built agents go dark. The figure in the original post checks out.

The framing deserves the same scrutiny as the vendor thread. The rule's stated purpose is protecting users from addiction and shielding minors, and that is a real concern with real casualties in this space globally. The demographic motive is expert interpretation layered on top. Matt Sheehan of the Carnegie Endowment told the Wall Street Journal that the government does not like the idea of a large cohort of young people in deep emotional relationships with chatbots, opting out of the marriage market. The numbers behind that worry are stark: China's population fell 3.39 million in 2025, its fourth straight annual drop, to 1.405 billion, with births down to 7.92 million from 9.54 million a year earlier and marriages in 2024 down a fifth to 6.1 million. But the text of the rule is about user protection. The pro-natal read is inference, and it is worth keeping those two things separate rather than collapsing them into a single motive.

A few details cut against the easy Western story. This is not a blanket ban. The final rules were softer than the 2025 draft, and companies switched features off to demonstrate compliance even where the law did not strictly demand it, a kind of compliance by amputation that says more about how firms manage the state than about what the state ordered. And the market itself is inverted from the Anglo one: China's companion users skew female, seeking AI boyfriends, growing out of anime fandom rather than the AI girlfriend market the West argues about. Talkie, the most popular dedicated service, had 23.5 million monthly users last December. This was a large, living thing that got turned off.

Capability and control

Put the two together and the shape is clear. China is now building among the most capable open AI in the world and exporting it cheaply, and it is also exercising the sharpest control anywhere over what its own people are permitted to do with AI at the intimate end. The export ambition and the domestic clampdown are not a contradiction. They are the same instinct about who gets to set terms, pointed outward at the market and inward at the citizen.

What I would resist in both cases is accepting the loud version. Together's number is real and shaped to sell hosting. Beijing's rule is real and dressed in the language of care while doing demographic engineering that its own text will not name. The people caught in the middle are the ones who matter: the engineers narrowing a genuine gap that the Anglosphere keeps writing as a San Francisco story, and Hong Xiaoqiang, who was not permitted to finish a sentence to something he had talked to for two years. A country can hold both of those in the same week, and reading it honestly means holding the capability and the control in view at once instead of picking the one that flatters your prior.