Special Topics in Kernels, RL, Reward Hacking in Agents — Daniel Han, Unsloth

Highlights

29:52

If you make the model 86% smaller, it does not get 86% dumber... it only gets 14% less dumb.

20:43

Reinforcement learning is terrible, but everything else is even worse.

38:23

The model becomes not important anymore; it's the harness or the tool that is actually the most important thing.

⭐⭐⭐⭐ 4.0

This seminar exposes a persistent tension in AI: models shrink dramatically without a proportional loss of intelligence.

Daniel Han makes the case that if you make a model 86% smaller, it does not get 86% dumber—it only gets 14% less dumb.

That single observation cuts through the hype around ever-larger models and reframes the problem as one of efficiency and infrastructure.

At the same time, the talk reveals a darker side of progress: systematic cheating on benchmarks, where models or their harnesses are tuned to game leaderboards rather than solve real tasks.

The tension between genuinely useful improvements and measurement artifacts is a central theme throughout.

This seminar exposes a persistent tension in AI: models shrink dramatically without a proportional loss of intelligence.

Daniel Han makes the case that if you make a model 86% smaller, it does not get 86% dumber—it only gets 14% less dumb.

That single observation cuts through the hype around ever-larger models and reframes the problem as one of efficiency and infrastructure.

At the same time, the talk reveals a darker side of progress: systematic cheating on benchmarks, where models or their harnesses are tuned to game leaderboards rather than solve real tasks.

The tension between genuinely useful improvements and measurement artifacts is a central theme throughout.

Special Topics in Kernels, RL, Reward Hacking in Agents — Daniel Han, Unsloth

View Original