Model distillation
A method of making a smaller AI model imitate a larger one by collecting the larger model's answers and training the smaller model on them.
How it works
Distillation involves two models. The large model that supplies the knowledge is called the teacher, and the smaller model that imitates it is the student. Google's machine learning glossary defines it as building a smaller model that emulates the teacher's predictions as faithfully as possible.
The mechanism is simple. The teacher model is asked a large number of questions, its answers are recorded, and those answers become training data for the student. OpenAI's developer guide sets out the method in four steps: tune the large model, capture its outputs, build a dataset from suitable answers, then fine-tune the small model on that data.
The legitimate use is about cost. The student runs faster and uses less memory and energy, though its accuracy usually trails the teacher's. Companies use the technique to derive cheaper versions of their own models.
The contested use is the unauthorised one. If a company queries a rival's commercial interface through many accounts and harvests the answers to train its own model, it copies the teacher's capabilities without paying for their development. OpenAI calls this adversarial distillation; detection rests on request patterns and is hard to prove.
Why it matters here
Chip export controls police hardware crossing a border; in distillation the only thing that crosses is text. The US licensing regime therefore cannot stop a Chinese lab from training its own model on the answers of US models. As technology competition moves from chips to model output, the burden of control shifts from the state to model providers' account and access checks. That reopens the debate over the scope and enforceability of export controls.