AI Essentials: What is Distillation?

This blog continues our series called “AI Essentials,” which aims to bridge the knowledge gap surrounding AI-related topics. A widely-used industry practice, distillation has recently become a topic of policy debate. What is distillation and what does it mean for startups and the broader AI ecosystem?

It’s often easier to learn from an expert than start yourself from scratch. When it comes to AI model training, the practice of one model learning from a more developed or expert model is called distillation, and it’s one used by industry leaders and startups alike

Think of a master chef. The chef knows unique recipes, how to tell when a steak is exactly medium rare, make a perfect créme sauce, and time a dinner so the food doesn’t get cold. The master chef is akin to a large AI model: enormously capable, but expensive to have on staff, and rightfully so, due to its depth of training. 

Now imagine the master chef begins to train an apprentice. The apprentice chef could follow the same playbook as the master chef: spend years in kitchens across the globe, pore over lessons on technique and obscure cookbooks, or he could watch the master chef cook. As the chef prepares each dish, he doesn't just say "now, add a pinch of salt" he shows the apprentice his reasoning: "Now, I'm adding salt to balance the acidity of the tomatoes, and I'm doing it at this point so it has time to dissolve." In doing so, the apprentice learns both the answer and the master's judgment. The same ritual occurs as hundreds of dishes are prepared, and each time, the apprentice continues to learn.

After this training, the apprentice can’t fully rival the master’s expertise. He can’t invent a complex new dish or explain the deep inner workings of a particular technique. But for the dishes that he’s been trained on, he cooks them almost as well, and he does it faster, without years of training and at a lower cost, since he’s an apprentice chef. 

That’s essentially distillation, when a smaller “student” AI model is trained on not just the outputs, but the reasoning of a larger “teacher” AI model. This works by training the smaller student model to match the probability distributions of the larger teacher model. If, for example, a larger teacher model was asked to predict the next likely word, and it wagers that there is a 50 percent chance of it being “flower,” a 40 percent chance of it being “tree,” and a 10 percent chance of it being “house,” that distribution reveals how the model thinks. “Tree” had a higher percentage chance than the house and was closer to the percentage chance of “flower,” so that must mean the teacher model believes flowers and trees are similar. After millions of similar examples across varying domains and tasks, the student begins to understand how the teacher thinks. The resulting smaller model is then faster, cheaper, and oftentimes just as capable as the larger model. 

Distillation has long been held as a common method for advancing AI research. Many of the largest AI research labs use distillation for their own internal model development—rather than starting from scratch every time. Startups use distillation to make their own more efficient models suited to their needs. A startup without massive amounts of compute could, for instance, distill a large, general model into a smaller one fine-tuned for a specific task, like analyzing medical records or dissecting legal contracts, without having to gather the vast amount of data and compute needed to train a larger model.

But unauthorized, coordinated distillation at scale, carried out in order to copy a competitor’s models and undercut the R&D costs behind its training, is a different kind of act. The recent release of Chinese model Kimi K3—which was nearly as capable and reportedly far cheaper than leading models from OpenAI and Anthropic—renewed focus on the technique of model distillation. Top White House tech advisor Michael Kratsios accused the Chinese AI company Moonshot of covertly distilling Anthropic's Fable model to help build its open-weight Kimi K3 model.  

(Distillation is also often conflated with open source AI, and that has raised the specter of a ban on open models. But both open and closed developers use distillation. Distillation is a model training technique, whereas open source is a distribution method—and generally beneficial for startups and innovation, as we wrote about in an earlier post in this series.)

The Kimi controversy was about unauthorized, coordinated, industrial-scale distillation—extracting a competitor’s model to sidestep the costs behind it—and that warrants scrutiny. The government’s response should recognize distillation itself isn’t an offense. It’s the same technique used by major labs, and the same one that lets a small startup build something innovative that would otherwise be out of reach. Rather than treating all manner of distillation as suspect, policymakers should target covert, coordinated distillation at scale by reaching for other policy levers, like legislation that could clear the way for AI labs to share information and coordinate responses to suspected foul play. Policymakers shouldn’t be considering blanket bans on distillation itself, but rather, seeking to address the harmful conduct of a few bad actors, while preserving the benefits of the method for startups, labs, and individuals. 

Previous
Previous

The Trump tariff rebuild continues and trade wars are back. What does it mean for startups?

Next
Next

Startup News Digest 08/21/26