Radar · 22/07/2026 · happened on 20/07/2026

Thompson proposes US law: training as fair use, distillation always permitted

Ben Thompson publishes a proposed US law that makes explicit an asymmetry now embedded in operations. The text asks for two things: declare that collecting public data to train models is fair use, and ban terms of service that prohibit distillation, at least for American companies.

The asymmetry stares back at anyone building. OpenAI and Anthropic built their models on billions of web pages and books without asking anyone’s permission. In their terms of service, though, they write that you cannot use their outputs to train your own model. Anyone evaluating between proprietary and open models lives it every day: you can consume the API, you can build on top, but what you learn stays theirs.

As we reported on July 14, Satya Nadella already voiced the same criticism publicly, calling the labs’ stance unsustainable. Thompson goes further and turns the discomfort into a proposal with two levers: give labs legal certainty on training, so they stop paying billion-dollar lawsuits, and ensure what they learn fuels everyone’s innovation.

The proposal lands as Alibaba just released Qwen 3.8 Max as open weight, 2.4T parameters. Thompson reads the move as alignment with Xi Jinping’s push on open source. If the US closes distillation, China opens it.

In detail

Distillation, in plain terms, is the process where you use answers from a large, expensive model to train a smaller model that mimics its behavior. Technically you do it by querying the starting model’s API on a set of examples, collecting question-answer pairs, and using that dataset to train the target model. OpenAI and Anthropic’s terms of service explicitly forbid this practice: if they discover you’re using their outputs to train a model, they can cut off your access.

Thompson’s point is that stopping distillation is practically impossible. Distilling a model means querying its API, and distinguishing a legitimate user asking many questions from someone building a training dataset is an unsolvable detection problem without spying on user behavior in ways that raise even bigger legal issues. The proposal inverts the logic: instead of banning something you can’t ban, legalize it and leverage it.

The proposal’s structure has two parts that support each other. The first declares that collecting public data for training is fair use. This protects labs from copyright lawsuits, like the $1.5 billion settlement Anthropic just signed on pirated books used to train Claude. The second bans terms of service that prohibit distillation, ensuring knowledge accumulated by frontier models stays available for anyone wanting to build alternatives.

The Chinese context is what gives the proposal urgency. Alibaba released Qwen 3.8 Max, 2.4T parameters, as open weight, reversing May’s decision not to publish it. Thompson links the choice to a recent Xi Jinping speech pushing for open source, collaboration, and sharing. If Chinese open models keep closing the quality gap with American frontier models and distillation stays banned in the US, competitive advantage shifts toward those with fewer constraints.

The proposal’s limits are concrete. It’s a policy suggestion, not a bill introduced to Congress. It doesn’t clarify how to draw the line between fair use and copyright violation for data that isn’t purely public, like copyrighted books or paywalled archives. And it doesn’t address the risk that legalizing training on any public data discourages quality content production if creators see no return. For anyone building AI products, though, the proposal marks a direction: pressure on this asymmetry is mounting, and Alibaba’s move suggests the market will find a way regardless, with or without US law.

Type to search across course, playbooks, skills, papers…