On July 22, 2026, White House science director Michael Kratsios posted on X that Beijing-based Moonshot AI had "distilled" Anthropic's frontier Fable model to build Kimi K3 β€” the 2.8-trillion-parameter open-weight model Moonshot released six days earlier, to reviews good enough that the company briefly paused new subscriptions. Treasury Secretary Scott Bessent followed with the quotable version: "Open source is not open season on American IP." When PRC firms cross "the line into IP theft," he warned, "sanctions and Entity List designations will be on the table." The Bureau of Industry and Security opened an investigation.

To appreciate how strange this is, you need the last year and a half.

The American AI industry has spent them arguing for the laxest IP regime it could get, and largely winning β€” up to the $1.5 billion Anthropic paid authors for the books it torrented. The administration didn't just agree with the industry. It enforced. When the Copyright Office released a report suggesting AI training on copyrighted works isn't always fair use, the White House fired the Register of Copyrights, Shira Perlmutter, the next day. She sued. The D.C. Circuit reinstated her. The administration went to the Supreme Court to finish the job and lost in June. Meanwhile the courts settled that raw AI output is copyrightable by no one β€” a rule the administration's own lawyers are defending in federal court right now.

So: train on anything, owe nothing, and what comes out the other end is owned by nobody.

Unless China does the training. Then it's theft.

Theft of what, exactly? Not a rhetorical question. The closer you get to the conduct being called "theft of American intellectual property," the harder it is to find any intellectual property in it.

For whatever it is worth, I have made the argument that machine output should be property β€” and I think the Copyright Office gets the statute wrong. But the real question isn't how much human it takes. It's whether a legal system should be deliberately manufacturing a vast new category of valuable things that belong to nobody. That is a truly strange thing to do on purpose and we're doing it by the terabyte. Someone spent money, ran a process, made a thing people will pay for. That usually gets an owner. What I don't believe is that the owner is Anthropic, or that the right is sitting in current doctrine waiting for a sufficiently motivated Treasury Secretary to find it. It isn't there. And in this case of machine on machine violence it really isn't there.

Copying homework at industrial scale

Distillation began as an ordinary machine learning technique β€” a 2015 Geoffrey Hinton paper β€” where a big "teacher" model's outputs train a smaller "student" model. The student learns not just the teacher's answers but its judgment.

The version alleged here is the adversarial cousin: you don't own the teacher. You hit a competitor's API with millions of prompts, harvest the responses, and train on the results. Anthropic reported in February that three Chinese labs generated more than 16 million exchanges with Claude through roughly 24,000 fraudulent accounts β€” 3.4 million attributed to Moonshot, with queries engineered to extract capability rather than to actually use the product.

Fake accounts, industrial-scale prompting, harvest, train. Nobody broke into a server. Nobody copied a model weight. Every piece of information the distiller got was something Anthropic's own product handed over, on request, through the front door.

The parts of a model copyright could protect, like the source code or the training pipeline, are exactly what distillation never touches. What it copies is outputs. And there is no version of American copyright law, current or pending, in which a lab owns the text its model produced for somebody else.

Two branches, same result. On the law as it stands, output generated without sufficient human authorship has no author at all: the D.C. Circuit's Thaler v. Perlmutter, cert denied in March, plus a Copyright Office that has spent years applying that rule to actual humans. Nobody owns it, so nobody can infringe it.

The other branch is live in Denver. Jason Allen wants registration of ThéÒtre D'opΓ©ra Spatial β€” the image, with himself as author β€” on the theory that he conceived it, specified it, and iterated 624 prompts until it matched the picture in his head. I argued in June that he's right, and for the reason the Supreme Court protected Napoleon Sarony's photograph of Oscar Wilde in 1884: authorship lives in conception and arrangement, not in the machinery that fixes the work to a surface.

Now apply that to a distillation attack, and notice that nobody is holding the camera. Sixteen million exchanges is not a person typing. It's scripts β€” templated prompts, sampled datasets, one model generating queries to interrogate another. Whatever human thought went into that lives in the design of the harvesting pipeline, not in the conception of any particular output. Sarony posed Wilde. Nobody posed anything here.

Which is the part worth sitting with. The government's own description of the attack β€” covert, automated, industrial-scale β€” is precisely what guarantees no copyrightable authorship came out of it. The scale that makes it feel like theft is the same scale that ensures there was never any property to take.

Both branches, then: nobody owns it. What the labs need is a third option no court has recognized β€” a right to control text they didn't write, generated at someone else's request, by a machine. That right would have to be invented.

Invention has a vault. Anthropic doesn't have the combination.

Patent is where real protection actually exists, and the vault has been getting bigger. Machine learning training methods are patentable subject matter β€” a claim reciting a specific technical improvement to how a model trains, as opposed to "apply generic ML to a new field" (which the Federal Circuit killed in Recentive Analytics v. Fox), clears Β§ 101. The Patent Office spent the past year widening the door: an August 2025 memo curbing reflexive "mental process" rejections, revised AI inventorship guidance in November, and an eligibility reset under Director John Squires, who says Β§ 101 shouldn't be a bludgeon against whole classes of invention.

Someone even holds the foundational distillation patent β€” Google, for "Training Distilled Machine Learning Models," inventors Vinyals, Dean, and Hinton, provisional filed nine months before the famous 2015 paper. Google made the deposit correctly: file first, publish second. The frontier labs a decade later chose secrecy instead, and secrecy runs on a clock β€” publish, ship, let Β§ 102's twelve-month grace period lapse, and the door locks with nothing inside.

None of which would matter anyway, and not because of where Moonshot trained. A training-method patent covers how you build your model. It says nothing about what someone else does with the outputs. The distiller never practices your method β€” they run their own training process on text your product handed them, which is no more infringement in Beijing than it would be in Boston. Wrong category, not close call. (Google's patent is the one Moonshot's conduct might actually read on, and there the territorial problems are real β€” a method practiced in China, with the Β§ 271(g) import route running into Bayer v. Housey, holding that provision covers physical products, not information. But that's Google's lawsuit, not Anthropic's, and nobody is threatening sanctions on Google's behalf.)

Real vault, wrong depositor, and the blueprints are still in the box.

A trade secret with a subscription tier

Trade secret is the theory that sounds right until you say it out loud.

Protection requires reasonable efforts to keep the information secret. The information at issue is the model's responses to prompts β€” which Anthropic sells, by subscription, to anyone with an email address. The output a distiller receives is the output any customer receives. The Eleventh Circuit's Compulife Software v. Newman, the mass-scraping case marking this theory's outer boundary, involved a proprietary database and copied source code. A public chatbot answering questions is neither.

You cannot sell something to the general public and simultaneously call it a secret.

The unglamorous statute that actually fits

Distillation does breach every lab's terms of service. But a TOS breach is a contract claim worth contract damages, and if every TOS violation is "theft," the word has no limiting principle. (You have, statistically speaking, stolen from a dozen companies this month.)

The one real legal theory here is also the least glamorous: the Computer Fraud and Abuse Act. Not for the distillation. For the access. After Van Buren v. United States (2021), violating a TOS doesn't "exceed authorized access" β€” the statute cares about gates, not motives. But 24,000 fraudulent accounts built to evade detection is a gates problem. The Second Circuit's United States v. Cuomo affirmed CFAA convictions for defendants who bypassed a public website's authentication gate using other people's credentials. Fake-account networks built to get around access cutoffs look more like Cuomo than like ordinary scraping.

So the honest description of the alleged conduct: creating fake accounts, at scale, to violate a EULA really, really fast.

That may well be a federal crime. It is not "model theft." Nothing was taken that Anthropic doesn't still have.

Somebody should tell Treasury what the DOJ is arguing in Denver

The Justice Department's brief in Allen v. Perlmutter has been sitting with a federal judge in Colorado since February, waiting on a ruling. Its position, across a hundred-odd pages: prompts are unprotectable ideas, and the expression in a model's output belongs to no one β€” not the user, and (the Office is emphatic on this) not the AI company either.

The administration's copyright litigators are telling a federal judge the vault is empty. The administration's sanctions officials are announcing a bank robbery.

Both can't be right, and my money's on the litigators β€” they have to survive Loper Bright review in front of an Article III judge. The tweets don't.

Distillation is standard industry practice, says man who distilled the company he's currently suing

The technical community's response to the Kratsios accusation wasn't outrage. It was arithmetic. Fable became reachable July 1 after an export-control hiatus; K3 shipped July 16. Researchers at the Allen Institute and elsewhere made the obvious point: you cannot distill that much data, train a 2.8-trillion-parameter model, and ship it in fifteen days. Anthropic's fraudulent-accounts report β€” the strongest public evidence β€” came out in February, describing conduct that predates Fable entirely. Kratsios disclosed no access logs, no fingerprints, no training records.

And the practice isn't exactly foreign. On April 30, on the stand in his own suit against OpenAI, Elon Musk was asked whether xAI had used distillation on OpenAI models to train Grok. He said it was general practice across AI companies. Asked whether that was a yes: "Partly."

So the conduct Treasury wants to sanction as theft of American intellectual property is, per sworn testimony from an American AI CEO, how the industry works.

None of which makes Moonshot innocent β€” the February numbers describe real, deliberate, fraud-adjacent extraction. It means the accusation driving the sanctions talk is unsupported by the people making it, about a practice an American CEO described under oath as standard.

Free-riding is real. It's also how everyone got here.

Strip the rhetoric and the dispute is economic: labs spend billions pushing capability forward, distillers capture much of that value for the price of API calls. Free-riding is a legitimate concern. It's just not theft, and the fix matters.

Bahrad Sokhansanj β€” a former Federal Circuit clerk at the Institute for Law & AI, whose Lawfare piece does the fullest autopsy of the theft claim β€” prescribes the sober version: prosecute the fraud, sanction where foreign policy warrants it, and study whether distillation even moves the frontier before legislating. What he warns against is what the "IP theft" framing is building toward: a quasi-property right in model outputs. A right to own text no human wrote, held by companies whose models were trained on the commons of everything humans ever wrote.

The labs learned from all of us. The proposed rule is that nobody gets to learn from them.

There's already a bill β€” the Deterring American AI Model Theft Act cleared House Foreign Affairs in April β€” and the name alone tells you Congress skipped to sentencing. The property question got waved through. It's the live question in American IP law: briefed in Denver, cert-denied in March, defended by a Register of Copyrights who went to the Supreme Court in June just to keep her job.

We haven't decided who, if anyone, owns what these machines produce. Until we do, "they stole it" is a conclusion in search of a premise.



Table of Authorities

Primary sources


Legalish is supported by Lynch LLP β€” Trademark Β· Copyright Β· Patents