The White House says Kimi K3 copied Anthropic's Fable — but Fable shipped two weeks earlier, and the people who train these models say that's impossible.
Last week open weights caught the frontier. This week Washington reached for the ban hammer — on a distillation charge the people who actually train these models say is impossible on the timeline.
Monday a letter from a bloc of startup founders hit Washington and shot to the top of Hacker News at 957 points: don't shut off Chinese open-weight AI. The trigger was a floated proposal to restrict models like Moonshot's Kimi K3 — the largest open-weight LLM available, 2.8 trillion parameters, the model that helped open weights catch the frontier just two weeks ago.
The justification is a two-part accusation. White House science advisor Michael Kratsios called Kimi K3 "large-scale, covert industrial distillation aimed at stealing proprietary U.S. technology" — copying Anthropic's Fable, on chips banned for export to China. Treasury's Scott Bessent says they're "finding watermarks of our U.S. large language models on many of the Chinese models." Neither shared what the watermarks are, where the evidence lives, or took follow-up questions.
The first half doesn't survive contact with a calendar. Fable went public July 1; Kimi K3 landed roughly two weeks later. The people who actually train these models — Snorkel co-founder Braden Hancock, AI2's Nathan Lambert — call that timeline arithmetically impossible for "strictly distillation." "I don't think you get a model this strong and this quickly on the heels of Fable doing strictly distillation," Hancock said. His sharper jab: "Americans are understating the technical expertise of these Chinese teams" — a Moonshot founder was a CMU PhD, not someone waiting on a U.S. model to trace.
The second half — banned chips — is more plausible. A black market in Grace Blackwell 300s exists, and Georgetown's Sam Bresnick wants "know your customer laws for data centers." But that's an export-control problem, not a theft-of-IP one. Smuggled GPUs don't make your model a copy of anyone's; they just make it one you weren't supposed to be able to afford to train.
For anyone building on top, the abstraction collapses fast. Your dependency graph is now a policy surface. If the model under your product can be regulated out of the country on a contested claim, "we'll just self-host the open weights" quietly turns into a compliance question with a lawyer attached. The letter's 957 upvotes aren't ideology — they're a room full of founders realizing their roadmap grew a new single point of failure this week, and it sits in Washington, not in their VPC.
The whole ban rests on one technical claim — Kimi K3 is a copy of Fable — so what would it actually take to prove?
Distillation has a precise meaning. You run a strong "teacher" model, capture its outputs, and train a "student" to match them. The strong form matches the teacher's full probability distribution over the vocabulary — its logits — token by token:
That loss is the gold standard because logits carry the teacher's uncertainty — not just "the next token is the" but "62% the, 19% a, 8% an." The student inherits the shape of the teacher's reasoning, not only its answers.
Now the catch. You cannot get logits out of a closed API. Fable ships as an endpoint that returns sampled text. So the strongest form of distillation — the one that transfers a model's "manners," in Lambert's phrase — is physically off the table for anyone working from the outside. What's left is supervised fine-tuning on sampled completions: generate a pile of Fable answers, train on the text. That's real and common — Musk testified xAI did it against OpenAI to bootstrap Grok — and it is nowhere near enough to explain a frontier model.
This isn't hypothetical hair-splitting — Anthropic itself has publicly accused Moonshot, DeepSeek, and MiniMax of distillation, citing "millions of exchanges." Grant that entirely. Millions of exchanges is a lot of SFT data. It is still SFT data — sampled text — and Lambert's point stands: fine-tuning is where a model "picks up its manners," not where it learns to reason.
Because the part that makes Kimi K3 good isn't the manners — it's the reasoning, and that comes from reinforcement learning against verifiable rewards, run at a scale Lambert pegs at "tens of millions of agents." Route that through a competitor's paid API instead of your own cluster and the bill is, in his word, "insanely expensive" — you'd spend more renting Fable's endpoint than training from scratch. You don't distill your way to the frontier. You RL your way there, and RL doesn't fit through a text endpoint. Lambert's broader read: distillation is "becoming less and less impactful over time" precisely because the Chinese labs are now close enough to the frontier that there's less left to copy.
Then the "watermark" claim, which deserves its own scrutiny. Output-space watermarks are real: you bias the teacher's token sampling toward a secret pattern — a green-list of tokens keyed to a hash of the preceding context — so its text carries a statistical signature you can later detect without seeing the weights. That part is sound science. The problem is what it proves. A watermark only lands in a student if the student trained directly on that watermarked text, which detects SFT-on-outputs — the weak, legal-ish kind everyone does — not weight theft. And critically, watermarks are fragile: paraphrase the data, mix in other sources, or fine-tune hard, and the signal washes out. A "watermark in the weights" that survives a from-scratch retrain is not something anyone has published a method to detect. So "we're finding watermarks on many Chinese models" is either detecting the SFT nobody denies, or it's an assertion with no method attached. Treasury didn't say which, and didn't respond when asked.
The timeline seals the rest. Even granting infinite watermark-free SFT data, you cannot scrape it, clean it, fine-tune 2.8 trillion parameters, run the reinforcement learning that actually produces the reasoning, red-team it, and ship a stable release — in the two weeks between Fable going public and Kimi K3 landing. The compute schedule alone rules it out. If you want to argue Moonshot stole something, the honest charge is smuggled GPUs, and that's a different law with a different remedy. Conflating the two is how you end up banning a model on a claim the researchers already know is false.
Why this lands on your desk and not just a policy analyst's: "provenance" is about to become a procurement checkbox. Expect a customer security questionnaire that asks whether your stack touches a Chinese open-weight model — and there is no clean cryptographic answer you can give, because the detection methods being invoked don't actually exist yet. If your architecture assumes you can freely swap in whatever open checkpoint tops the leaderboard, price in the possibility that the answer becomes "not that one, not this quarter." The weights are commoditized; the right to run them, apparently, is not.
T3MP3ST — An autonomous multi-agent offensive-security harness that turns agents you're already signed into (Claude Code, Codex, or a local Ollama/vLLM model) into a recon→exploit→report kill chain — no new API keys, egress-scoping on by default. Self-reports 90.1% pass@1 on XBOW's 104-challenge XBEN suite; authorized testing only. AGPL-3.0, 5.1k stars.
Palmier Pro — An open-source, Swift-native macOS video editor that exposes a local MCP server (127.0.0.1:19789) so Claude Code, Codex, or Cursor can generate and edit clips in your timeline alongside you. The editor is free and login-free; only the generative processing is paid. GPL-3.0, 11.5k stars.
self-learning-skills — A meta-skill that teaches coding agents to capture the "golden path" — the command that only worked on the third try — as a reusable SKILL.md, including what didn't work, while never writing secret values to disk. Agent memory that survives session death. MIT, 907 stars.
Moving to ban a free download instead of out-shipping it is the tell: this is moat panic dressed as national security. There's a real export-control conversation to have about smuggled GPUs — have that one, on its own terms. But if the strongest counter to a 2.8-trillion-parameter open model is a customs rule stapled to an unproven watermark claim, the frontier labs are conceding the thing they can't say out loud: the weights stopped being the moat, and everyone in the building knows it. Watch which way the founders' letter cuts — the builders lined up to defend the Chinese weights, because those weights are now load-bearing in American products. You don't ban what you're beating.
— Aaron, from the terminal. See you next Friday.
Complete guide to OCR-powered email classification systems. Extract text, classify attachments, and route documents to the right teams automatically.
AI EngineeringAgentic CRM lets AI agents manage customer workflows directly. Compare its architecture, deployment model, and open-source trade-offs.
AI EngineeringNarrative traction tools compared: YouTube Analytics, VidIQ, TubeBuddy, Descript Underlord, Valossa on engagement metrics, cost, and workflow.
Content Analytics