Y Combinators Garry Tan wants US open-weight AI labs to distill frontier models, too
Artificial IntelligenceCurated News 2026-09-11 6 min read

Y Combinators Garry Tan wants US open-weight AI labs to distill frontier models, too

Tan wants smaller, American open-weight AI labs to use the same kind of training techniques on American frontier AI labs, giving the U.S. a more robust set of open-weight options that aren’t Chinese.

Researched and edited by Kiran Ch and the WhatIsFuture editorial team. Reviewed for factual accuracy before publication.

If you’ve been paying attention to the OpenLLM leaderboards over the past six months, you’ve noticed a pattern that should make every Silicon Valley founder sweat. Chinese AI outfits like DeepSeek and Alibaba’s Qwen team aren't just competing with Western open-weight models—they are actively running circles around them. How are they doing it? By aggressively distilling frontier models built by OpenAI, Anthropic, and Google. They take the raw, expensive reasoning outputs of closed American models, feed them into hyper-efficient architecture pipelines, and produce open-weight powerhouses at a fraction of the cost.

Now, Y Combinator CEO Garry Tan is calling for American startups to drop their polite reservations and do the exact same thing. Speaking out on what has long been an open secret in developer channels, Tan wants domestic open-weight labs to leverage American frontier intelligence to build top-tier, American-hosted open models. The signal from TechCrunch’s latest coverage of Tan's comments is crystal clear: if Western builders remain tethered by strict Terms of Service interpretations while global competitors freely harvest synthetic data from state-of-the-art APIs, America risks losing control of the open-source AI infrastructure layer altogether.

Private Community

Join Our Tech Community

Get instant alerts on the most critical AI breakthroughs on our WhatsApp channel. No spam, just signal.

Join Channel Free →

Key Takeaways

  • Garry Tan’s Call to Arms: YC’s chief believes American open-weight AI labs must actively distill proprietary frontier models to stay competitive against aggressively distilled Chinese alternatives.
  • The Regulatory Disconnect: While US labs face tight Terms of Service (ToS) restrictions against using output data for model training, offshore teams operate largely outside Western IP enforcement mechanisms.
  • The Economic Equation: Distillation drops the cost of training high-performing small-to-midsize models by orders of magnitude, making it an existential shortcut for resource-constrained startups.
  • Ecosystem Security Risks: Relying predominantly on foreign open-weight foundation models creates long-term supply chain and sovereign security liabilities for Western software stacks.

The Distillation Paradox: Legal Friction vs. Geopolitical Reality

Model distillation is not magic; it’s high-efficiency knowledge transfer. You take a massive "teacher" model—say, Claude 3.5 Sonnet or GPT-4o—generate millions of complex reasoning traces, and use that high-quality synthetic data to train a lightweight "student" model. The student gets to skip the multi-million-dollar pre-training phase where a model spends months figuring out basic syntax, grammar, and simple logic. Instead, it gets straight to mastering complex reasoning frameworks at a fraction of the compute overhead.

The problem is that Western proprietary AI vendors hate this. If you read the fine print on OpenAI or Anthropic API agreements, there is almost always a explicit clause forbidding users from using their API outputs to train competing models. When Anthropic details distillation campaigns from Alibaba, Moonshot AI, and DeepSeek, it exposes a massive asymmetrical advantage. Asian labs operate largely outside the reach of US civil litigation and corporate enforcement. They scrape, distill, fine-tune, and release world-class open-weight weights to global acclaim while American founders tip-toe around legal landmines.

Garry Tan’s stance is a pragmatic reaction to this imbalance. If the rulebook is only followed by one side, the rulebook becomes a liability. For Silicon Valley to maintain leadership in open weights, American builders need a clear pathway to leverage synthetic data pipelines without being paralyzed by legal threats from closed-source tech giants sitting on hyper-scale compute clusters.

Why Chinese Labs Are Winning the Open-Weight Race

Let's not mince words: DeepSeek-R1 and Qwen 2.5 were not built in a vacuum. They are masterpieces of algorithmic engineering, architecture optimization, and yes, intensive synthetic data distillation. Chinese AI researchers realized early on that competing head-to-head on massive pre-training runs using raw web scrapes required compute budgets they simply couldn't access due to export controls and GPU constraints. Distillation became their ultimate force multiplier.

By extracting reasoning paths, chain-of-thought tokens, and structured code logic from American APIs, non-US labs managed to bypass billions of dollars in trial-and-error pre-training compute. The result? Developers globally—including thousands of startups in San Francisco and New York—are default-deploying models built on foreign open-weight architectures simply because they perform better than domestic open alternatives.

This creates a massive strategic vulnerability. When the default engine for open-source AI agents, enterprise RAG pipelines, and developer tools is anchored overseas, the underlying norms, safety alignments, and architectural choices of those models are set by foreign research priorities. Tan’s push isn't about IP theft; it's about ecosystem survival.

The Compute Wall: Distillation Is an Economic Imperative

There is a harsh financial reality that every startup founder learns early in their funding journey: pre-training a foundation model from scratch is a game reserved for sovereign wealth funds and trillion-dollar hyperscalers. When Jensen Huang explains why Nvidia will grow an astounding 70% next year, he is describing a world where capital expenditure for raw compute is skyrocketing beyond the reach of traditional seed-stage startups.

If an American startup wants to release a specialized open-weight model for robotics, biomedical research, or code generation, spending $50 million buying H100 clusters to pre-train from raw text is suicide. Distillation reduces that capital requirement by 90% or more. It allows small teams to stand on the shoulders of frontier models and spend their precious runway on domain-specific optimization and architectural innovations.

Furthermore, hardware efficiency isn't just a software problem—it's tied to real-world infrastructure constraints. We already know that powering AI is an architecture problem stretching power grids to their limits. Distillation allows developers to produce hyper-dense, 8B-to-70B parameter models that run locally on edge hardware, saving gigawatts of data center power while delivering 95% of the intelligence of a frontier API.

"We are effectively fighting an industrial policy war in software where one side plays strictly by legacy copyright compliance, while the other treats frontier API outputs as a public natural resource. Garry Tan is simply calling out the naked emperor."

What Closed-Source Labs Get Wrong About Ecosystem Wars

The core friction point here is the zero-sum mindset of frontier labs like OpenAI and Anthropic. They view model output harvesting as direct IP theft that cannibalizes their API revenue. But history in open-source software tells a fundamentally different story. Closed platforms rarely win the developer mindshare battle long-term when robust open-weight alternatives exist.

If closed labs aggressively sue or restrict American open-weight developers from distilling their models, they won't stop distillation from happening. They will simply ensure that all the downstream distillation happens off-shore, powering models that they have zero visibility or influence over. By refusing to accommodate a legal or licensed distillation framework for domestic builders, closed labs are inadvertently handing the entire open-source ecosystem to their foreign rivals on a silver platter.

We need a standardized framework—perhaps via API tier licensing or safe-harbor provisions for open-research distillation—that allows domestic developers to tap into frontier synthetic data streams. Expecting Silicon Valley startups to play with one hand tied behind their backs while international teams sprint ahead using distilled reasoning traces is no longer a viable strategy for American tech leadership.

Frequently Asked Questions

What exactly is AI model distillation?

AI model distillation is a technique where a smaller, lightweight model (the "student") is trained using the outputs and reasoning paths generated by a much larger, highly capable model (the "teacher"). This allows the smaller model to achieve near-frontier performance at a small fraction of the training cost and compute requirement.

Why is Garry Tan urging US labs to distill models?

Garry Tan wants American open-weight AI startups to stay competitive with foreign labs (particularly in China) that are aggressively distilling frontier outputs to produce top-tier open models. Without active distillation efforts in the US, Western developers risk relying entirely on foreign-built open-weight software stacks.

Is model distillation legal under current Terms of Service?

Most proprietary AI providers (such as OpenAI and Anthropic) explicitly prohibit using their API outputs to train competing commercial AI models in their Terms of Service. However, enforcing these clauses internationally has proven nearly impossible, creating an regulatory asymmetry between US-based labs and overseas competitors.

This analysis was inspired by a story originally reported by TechCrunch. Read the original report →

Recommended Tool

Supercharge Your Workflow with Claude AI

The AI assistant used by professionals worldwide. Write, code, analyse — all in one place.

Try Claude Free →