Where AI Copying Ends and Theft Begins

Posted by Adam Danyal in AI Risk on

Jensen Huang, chief executive of the chipmaker Nvidia, had never posted on X. His first post shared a letter signed by twenty-five companies, investors and open-technology groups, and the reason for breaking a long public silence sits in one paragraph near the end of the document.

The statement, titled Open Weights and American AI Leadership, carries the names of Microsoft, Meta, IBM, Palantir, Dell Technologies, Hugging Face, Mistral, Mozilla, CrowdStrike, ServiceNow, Andreessen Horowitz and Y Combinator. Most of the text makes a familiar case. Open-weight models, meaning models anyone downloads, inspects, modifies and runs on their own infrastructure, widen access to advanced AI, increase competition, and give customers a way to avoid being locked into a single supplier. The signatories compare the moment to the 1980s, when open-source software was treated as a threat to commercial code and went on to underpin most of the internet.

Then comes the paragraph with law and money attached to it. Policymakers, the signatories write, should be careful not to conflate legitimate model-development techniques with misappropriation.

The technique in question is distillation, the practice of using one model’s outputs to help train or improve another. A team sends questions through an existing model, collects the answers, and uses those answers as teaching material for a newer or smaller system. The letter describes the method as widely used for model improvement, evaluation and validation, and places it inside a long tradition of learning from and building on existing technology.

The other side of the line gets drawn with the same directness. Unlawful efforts to extract value from closed models raise legitimate concerns, the document says, and those concerns belong in targeted legal and commercial frameworks rather than sweeping restrictions on a technique the industry runs on. The wording is deliberate. Washington is weighing limits on open models, and distillation has become the word doing the heaviest lifting in the argument for restriction.

Real misappropriation already has a courtroom record, and the record looks nothing like distillation. In January a federal jury in San Francisco convicted former Google software engineer Linwei Ding on seven counts of economic espionage and seven counts of theft of trade secrets. Over roughly a year he pulled more than two thousand pages of confidential material off Google’s network, covering the architecture of the company’s Tensor Processing Unit chips, its graphics processing systems, and the software linking thousands of chips into a supercomputer for training large models. While still on the payroll he was secretly affiliated with two technology companies based in China, and he told potential investors he could build an AI supercomputer by copying and modifying his employer’s technology. The Justice Department called the verdict the first AI-related economic espionage conviction in American history.

OpenAI, Anthropic and Google are absent from the list of signatories. Two of the three are already in court over material taken from someone else. On Monday a federal judge approved a $1.5 billion copyright settlement under which Anthropic pays authors roughly $3,000 per book for pirated copies used to train Claude, described by plaintiff counsel as the largest known copyright recovery in history. Earlier this month Apple sued OpenAI and two former Apple employees, alleging they carried confidential hardware documents, unreleased product details and supplier information into the AI company. Both disputes turn on material a company never made public.

The distinction the coalition draws rests there. One set of behaviour involves taking documents, files and specifications a company kept behind its own walls. The other involves sending prompts to a model already released to the public and studying what comes back. Regulators drafting a single rule against copying would cover both, and the signatories argue the second belongs in a different category from the first.

The letter binds no one. The signatories also concede open weights carry real risks, noting that once weights are released they move beyond the original developer’s control and modified versions become difficult to trace or reverse. The request is narrower than a defence of openness. Washington is being asked to write rules against theft without writing rules against a method the industry already depends on.


Sources

NVIDIA, Open Weights and American AI Leadership — nvidia.com
X, Jensen Huang’s first post — x.com
Tom’s Hardware, Nvidia and 24 other companies sign open-weights letter as Washington weighs Chinese AI model ban — tomshardware.com
U.S. Department of Justice, Former Google Engineer Found Guilty of Economic Espionage and Theft of Confidential AI Technology — justice.gov
ABC News, Judge approves a $1.5B Anthropic settlement over pirated books used to train the Claude chatbot — abcnews.com
TechCrunch, Apple sues OpenAI over alleged trade secret theft — techcrunch.com