\n\n\n\n Copy That Homework, Says Y Combinator's Garry Tan - Agent 101 \n

Copy That Homework, Says Y Combinator’s Garry Tan

📖 5 min read•811 words•Updated Sep 12, 2026

Picture a small team in a rented office somewhere in the U.S. Four engineers, a coffee machine that sounds like it’s dying, and a model they’re trying to train. They don’t have a billion dollars. They don’t have a data center the size of a shopping mall. What they do have is an idea: what if they could teach their small model to think like one of the giant models that already exists?

That’s the picture Garry Tan, CEO of Y Combinator, seems to be sketching out. His argument, reported by TechCrunch, is that smaller American open-weight AI labs should be using distillation techniques on American frontier models, the same way Chinese labs have been doing it. His reasoning is that this would give the U.S. a stronger set of domestic AI options, and better balance between open-weight models and the frontier ones locked behind company walls.

If you’re reading this on agent101.net, you’re probably here because you want to understand AI agents without a computer science degree. So let me unpack what “distillation” actually means, because it’s one of those terms that sounds far more mysterious than it is.

Distillation, explained without the jargon

Imagine you have access to a brilliant professor. She’s read everything, she’s expensive to book, and she lives in a building you can only enter through a very specific door. You can’t take her home with you.

What you can do is ask her thousands of questions, write down all her answers, and then use those answers to train a much cheaper, much smaller tutor. That tutor won’t know everything the professor knows. But on the topics you care about, it can get surprisingly close.

That’s distillation. A large “teacher” model generates outputs. A smaller “student” model learns from those outputs instead of learning from raw data alone. The result is a model that runs cheaper, faster, and often on hardware you could actually afford.

The word “distill” is well chosen. You’re boiling something big down to its essence and throwing away the bulk.

Why this matters if you care about AI agents

Here’s where it connects to the stuff we cover on this site. AI agents are programs that take actions on your behalf: booking things, sorting your inbox, running research tasks, calling other tools. Agents are chatty by nature. They don’t ask a model one question, they ask it dozens in a loop, checking their own work as they go.

That means the cost of the underlying model gets multiplied. A model that’s cheap enough to run for a single chat message might be painfully expensive when an agent hammers it two hundred times to complete one task.

Smaller distilled models change that math. And open-weight models, ones where the actual model file is published and you can download it, change it further, because you can run them on your own hardware without paying per request at all.

So the practical version of Tan’s argument, translated for non-technical readers, is roughly: more small, capable, downloadable American models means more people building useful agents without needing a giant company’s permission or budget.

The uncomfortable part

You may already be squinting at something. If distillation means training your model on another company’s model outputs, isn’t that a bit… cheeky?

That’s exactly the tension in this story. Chinese labs have been credited with using distillation on frontier models to close the gap quickly. Tan’s position is that American open-weight labs should be free to do the same thing on American frontier models. Which means the conversation isn’t purely technical, it’s about what the rules should be, and who gets to set them.

I’m not going to pretend to resolve that here. I’ll just point out what’s interesting about the framing: Tan isn’t asking frontier labs to slow down or open up their weights. He’s arguing that the smaller players should be allowed to learn from them. It’s less “tear down the walls” and more “let people take notes through the window.”

What to actually take away from this

If you’re a non-technical person trying to keep your bearings, three things:

  • Distillation is a compression trick, not magic. Small models learn from big models. That’s it. The technique is why you keep seeing tiny models that punch above their weight.
  • Open-weight matters for agents. Cheap, downloadable models are what let ordinary developers build agents that run constantly without a scary bill.
  • The debate is about permission, not capability. Nobody’s arguing distillation doesn’t work. The argument is over who’s allowed to do it, and to whom.

The takeaway isn’t that a policy question got settled. It’s that a prominent voice in Silicon Valley is publicly asking whether the U.S. open-source side of AI is fighting with one hand tied. For anyone who wants to build agents without a nine-figure budget, that’s a question worth following.

🕒 Published:

🎓
Written by Jake Chen

AI educator passionate about making complex agent technology accessible. Created online courses reaching 10,000+ students.

Learn more →
Browse Topics: Beginner Guides | Explainers | Guides | Opinion | Safety & Ethics
Scroll to Top