Why I’m betting on AI distillation — and cheaper models

AI distillation is slashing inference costs, alarming frontier labs and forcing Washington to decide where optimization ends and theft begins.

Why I’m betting on AI distillation — and cheaper models

AI distillation is the generic-drug moment big labs were dreading

Silicon Valley and DC are obsessed with distillation. The old technique caused panic when it crushed prices.

Kimi K3 delivered roughly frontier-level coding for half the price of OpenAI’s GPT-5.6 Sol. I opened a spreadsheet.

I inspect AI pricing like my nonna inspected market tomatoes: suspiciously, ready to reject an insultingly soft San Marzano.

Cheaper inference means better margins to me, national emergency to Washington.

“Teacher-student model compression” sounds technical. The fight is over who can charge a premium for intelligence, and how long.

Frontier labs funded expensive discovery. Distillation reproduces much of its useful behavior without copying weights: AI’s generic-drug moment, minus patents keeping everyone civilized.

Copyright, patents, contracts and trade-secret rules cover pieces of the dispute, but none cleanly fits one model learning from another’s answers.

I understand the labs’ nerves. I would share them.

Google has used this “dangerous trick” for years

Knowledge distillation is over a decade old: a large teacher guides a smaller student to reproduce useful behavior with less computation.

It compresses expertise. The student gets no weights, architecture or training data; it studies behavior.

Google AI chief Jeff Dean described it routinely in a February 2026 podcast:

Through distillation, which is a key technique for making the smaller models more capable, you have to have the frontier model in order to then distill it into your smaller model.

Nvidia distilled its Llama Nemotron models too. Nobody summoned the Senate when Jensen Huang’s company did it; the technique became sinister when Chinese labs mastered it.

The results matter. The BIRD paper, submitted to arXiv on July 17, 2026, used Qwen3-8B. Leichao Dong and co-authors raised MATH-500 accuracy from 86.2% to 92.0% while cutting average responses from 3,099 tokens to 1,115.

Better answers with roughly 64% fewer tokens. Every founder paying inference bills just sat straighter.

BIRD uses self-distillation: the model cleans and shortens its own reasoning. No foreign competitor in a trench coat—just a model realizing it talks too much, like everyone in a 90-minute Zoom.

Reasoning models repeat checks and explore dead ends. Distillation keeps useful capability while removing expensive verbal furniture.

That threatens frontier labs: billion-dollar models can teach cheaper systems valuable tasks. The lab keeps its weights but loses the scarcity premium.

Generic drugs also reproduce results after someone else funded discovery, without copying lab notebooks. AI complicates the law, but the market pressure is here.

Optimization for me, theft for thee

Anthropic has offered evidence worth scrutiny. In February 2026, it said DeepSeek, Moonshot and MiniMax created roughly 24,000 fake accounts and generated 16 million Claude exchanges for industrial-scale distillation.

That is no student experiment. If Anthropic is right, fake identities and access evasion show deliberate conduct. I owe account farms no philosophical loyalty.

White House science adviser Michael Kratsios escalated the Moonshot allegation on July 22:

We have information that Moonshot AI distilled Anthropic’s Fable for the development of its K3 model.

Kratsios also alleged Moonshot built a platform that switched access methods to evade detection. Fortune reported he linked it to Nvidia GB300 systems in Thailand, which current US export controls bar Chinese companies from using.

That allegation—access evasion and possible export-control violations—needs evidence. As of July 25, the public had claims but no technical receipts.

The timeline complicates it. Anthropic’s Fable became public July 1; Kimi K3 launched July 15. Moonshot employee Randy Xian replied with an Italian waiter’s subtlety when asked for ketchup:

Yes, Fable went public on July 1 and K3 launched on July 15. We trained a brand new frontier model in JUST 15 DAYS. Guinness World Record stuff.

Timing disproves nothing: earlier access, other Anthropic models or late-stage post-training could matter. But benchmark similarity proves little.

My enforcement line is conduct: fake-account farms, credential fraud, security circumvention and deliberate contract evasion. Courts and regulators can examine those.

Public answers are harder to own. Anthropic might prove someone broke into the classroom without owning everything learned there.

Selective outrage hurts Silicon Valley. OpenAI and Anthropic trained on vast amounts of human work and face lawsuits over it. CNBC quoted Max Pritt, an attorney for authors suing AI companies, criticizing an administration that vigorously defends tech-company IP while largely ignoring the creators who trained those systems.

Sixteen million fake-account exchanges differ legally from reading a public webpage. Still, labs say investment justifies broad control over outputs. Authors, journalists, artists and programmers said the same; Silicon Valley waved fewer flags.

I sympathize with Anthropic more than my tone suggests. Anger still writes terrible property law.

AI distillation is crushing the price of intelligence

Kimi K3 appears good enough for serious work at a lower price. The Associated Press reported that K3 topped Arena’s front-end coding ranking.

Arena CEO Anastasios Angelopoulos was clear:

This may be the single biggest release of the year.

Bank of America analysts cited by AP estimated K3 costs about half as much as OpenAI’s GPT-5.6 Sol. Executives will notice; most customers do not want to fund Earth’s theoretically smartest model.

They want reliable performance at a sustainable price. Benchmark prestige ranks below “does this break Friday night?”

Hardware taught me painfully: beautiful features mean nothing when cloud costs eat margins or updates break pairing. Technical superiority buys less time than I assumed.

Usually, none.

“Good enough, available and affordable” has buried superior products. Impressive GPU diagrams grant AI no exemption.

SecurityPal founder Pukar Hamal told CNBC he would consider hosting Kimi K3 after checking it for backdoors, because of the significant savings.

That is purchasing: security first; ideology around item 14, after uptime and the finance person compares the token bill with Milan rent.

In a July 22 Axios interview, Jensen Huang said excellent Chinese open models should reach American companies. Nvidia benefits because cheaper models increase usage, chip demand and data-center demand.

Huang put it plainly:

Distillation, learning from AI, learning from other sources of knowledge, is fundamental to intelligence.

Nvidia profits as consumption spreads. Closed labs profit while intelligence stays scarce and API-metered. Kimi means expansion in Santa Clara and margin compression in San Francisco.

When luxury tasting becomes a €12 pasta, the chef calls it commoditization. The investor calls Washington.

A diagram illustrating AI distillation process, showcasing model efficiency and cost reduction in technology development.

Alt text: AI distillation diagram showing synthetic data and capabilities flowing between American and Chinese teacher and student models.

The knowledge flow already runs both ways

Washington’s story of American invention flowing outward has expired. AI is a group chat sharing code, papers, synthetic data and suspiciously familiar ideas.

Mira Murati’s Thinking Machines raised $2 billion, then said its Inkling model used DeepSeek-V3’s architecture. According to Rest of World, post-training also used synthetic data from Moonshot’s Kimi K2.5.

This was no basement operation scraping Hugging Face. Founded by OpenAI’s former chief technology officer, Thinking Machines is among Silicon Valley’s highest-profile AI companies.

San Francisco startup Anysphere acknowledged that a leading Cursor product used Kimi K2.5. AP reported SpaceX plans to acquire Cursor for $60 billion.

Chinese capability already powers an American software company with a proposed valuation exceeding Ford’s market capitalization on many trading days. Purity gets difficult when bankers arrive.

After Chinese regulatory approval, Apple planned Apple Intelligence in China around Alibaba’s Qwen and Baidu’s Ernie. Rest of World reported that the US Department of Defense designated both companies Chinese military-affiliated.

An American iPhone can run approved Chinese AI in China, then return through LAX in someone’s pocket. Draw that on a Cold War map.

Engineers have deadlines. I care about performance, licensing, cost and customization. Nationality matters for legal or security exposure; passports do not improve coding.

OpenAI, Anthropic, Google DeepMind, Meta, Alibaba, Tencent, Moonshot, DeepSeek and Zhipu have published synthetic-data or teacher-student work. The question is which models teach, with what permission.

Distribution defeats blanket restrictions. Governments can block advanced Nvidia chips; quarantining published architectures is harder once weights reach Hugging Face, GitHub, clouds and local machines.

Europe depending on American closed APIs while Chinese open weights set prices makes me deeply uncomfortable.

Launching the AI Continent Action Plan on April 9, 2025, European Commission executive vice-president Henna Virkkunen supplied urgency:

The global race for AI is far from over. It is time to act.

Correct. Fragmented strategies leave Europe renting intelligence from two foreign powers. European companies need capital, compute and a continental market, not country-by-country rebuilding.

Regulation without European champions creates excellent paperwork and strategic dependency. Bravissimo.

If my model is the moat, mamma mia

A startup built only around the smartest API stands on melting ice. Distillation turns premium capability into a cheap dependency.

Faster databases did not kill software companies; they killed “we have a database” as a pitch.

AI will follow. Durable value lies in proprietary workflow data, customer trust, domain evaluations, integrations, permissions and real-use feedback.

Less exciting than benchmark screenshots, these assets survive a 70% model-price fall.

Distillation now exceeds chatbot mimicry. A July 23 paper by Chenhui Gou and four co-authors introduced “Experience Distillation,” turning an agent’s interaction history into reusable behavior without new environment calls.

Across 749 software-engineering tasks and six text-adventure games, it retained at least 64.8% of in-context-learning gains. Direct supervised fine-tuning recovered only 3.8%.

It matched reinforcement-learning baselines with at least 9.6 times fewer environment samples. For agents learning through costly experiments or human feedback, that can decide viability.

OPOD, another July 23 submission, coordinated separate text, image and audio teachers. Across 12 benchmarks and three model sizes, it achieved the highest average score at every scale.

At 30 billion parameters, it beat its base model and a jointly post-trained counterpart on all 12 benchmarks. The specialist teachers could then be discarded, leaving one deployable multimodal model.

Expensive specialists will teach cheaper product models, then leave production. Customers get lower latency and bills. Nobody buys champagne for the teacher.

My founder checklist:

  • Assume model prices will keep falling.
  • Keep the product portable across API providers and open-weight models.
  • Build internal evaluations around customer outcomes, not Arena screenshots.
  • Control sensitive data and user-generated feedback.
  • Do not call “our model” a moat unless the company trained and owns something defensible.

I run Docker on Linux for the same reasons. My Ghost site, ERP, analytics, mail, automations and SvelteKit image interface sit behind infrastructure I control. Self-hosting makes me question life at 1:12 a.m., but portability matters when vendors change prices or policies.

Planning around permanent access to one magical model makes a startup somebody else’s pricing experiment.

Punish the break-in and leave studying alone

The White House is drawing a useful line. In a July 24 Axios report, Kratsios defended authorized distillation for efficient models while condemning covert industrial-scale extraction.

He wrote:

Legitimate AI distillation used to create smaller, more efficient models plays a vital role in this open innovation ecosystem.

Agreed. Target fake accounts, stolen credentials, prohibited automation, access rotation, privacy breaches and circumvention of technical controls.

Sanctions need more than benchmark vibes. Treasury Secretary Scott Bessent threatened sanctions and Commerce Department Entity List designations when Chinese companies cross into IP theft:

When PRC firms conduct covert, industrial-scale distillation attacks that cross the line into IP theft, sanctions and Entity List designations will be on the table.

Fine. Government should disclose enough evidence to distinguish extraction from independent performance—watermarks, account patterns, prompt distributions and access logs—without revealing every defense.

Broad restrictions would hurt American startups first. WIRED reported that more than 200 companies, organized through the Little Tech Association and including Y Combinator, wrote Kratsios and Commerce Secretary Howard Lutnick opposing an outright foreign open-weight ban.

Removing cheap alternatives entrenches a few American labs. Monopoly pricing apparently becomes patriotic after the right lobbying meeting.

Another letter, signed by Nvidia, Microsoft, Meta, Palantir, Box and more than 20 other companies, warned against premature open-weight restrictions and called distillation routine:

Distillation, or the practice of using one model’s outputs to help train or improve another, is a widely used technique for model improvement, evolution, and validation.

Researchers would suffer too. Former Biden White House adviser Suresh Venkatasubramanian told Axios that losing Chinese open-weight models would seriously harm scientific research. Unlike closed APIs, open weights allow inspection and modification.

OpenAI co-founder Greg Brockman treats adversarial distillation as a technical problem. According to Semafor and Axios, OpenAI combines machine learning and human review to detect mass synthetic-data generation, response scoring and attempts to extract reasoning.

Capable labs should invest there. Detection can make abuse costly without blocking research or authorized training.

Huang also told Axios that one closed model creates a single attack and failure point. Downloadable models can run in controlled environments with inspection and restricted network access.

I support sanctions for proven industrial extraction. Vague anti-distillation rules would protect incumbents and weaken the startups Washington praises. Flag pins do not fix competition policy.

The teacher does not get to retire

By 2028, distillation will quietly power nearly every serious AI product. Frontier systems will teach cheaper specialists; agents will absorb costly interactions; multimodal teachers will collapse into deployable models.

Self-distillation will remove wasted reasoning without outside teachers. Marketing may drop the term because customers will expect lower latency and prices.

Frontier labs remain essential: someone must create capabilities worth compressing, and the best teachers set the ceiling. Their advantage will be faster improvement, dependable infrastructure and enough trust to command premium prices.

First place offers no retirement plan. The teacher must keep teaching.

An American AI strategy based on nobody learning from American models has the structural integrity of a wish.

Frequently asked questions

What is AI distillation?

AI distillation is a machine-learning technique in which a teacher model produces answers or guidance that a smaller student model learns to reproduce. The student does not receive the teacher’s original weights, internal architecture or training dataset, allowing useful capabilities to be delivered with less computation.

Why is AI distillation controversial?

AI distillation is controversial because frontier labs invest heavily in developing advanced models while cheaper systems can learn from their outputs. The dispute involves model pricing, intellectual property and access rules, especially when companies allegedly use fake accounts, stolen credentials or technical circumvention to collect outputs at industrial scale.

How does AI distillation reduce AI costs?

AI distillation can reduce costs by teaching smaller models to preserve useful capabilities while using fewer tokens and less computation. It can remove repetitive reasoning, transfer lessons from expensive interactions and combine specialist knowledge into deployable models, producing lower latency and smaller inference bills for customers.

Sources

Related reading

Luca

Luca

Luca by the way is the personal blog of Los Angeles based entrepreneur Luca Capula. A true Italian who lives between Torino and LA.

More posts →