AI Infrastructure

Special Topics in Kernels, RL, Reward Hacking in Agents — Daniel Han, Unsloth

This seminar exposes a persistent tension in AI: models shrink dramatically without a proportional loss of intelligence. Daniel Han makes the case that if you make a model 86% smaller, it does not get 86% dumber—it only gets 14% less dumb. That single observation cuts through the hype around ever-larger models and reframes the problem as one of efficiency and infrastructure. At the same time, the talk reveals a darker side of progress: systematic cheating on benchmarks, where models or their harnesses are tuned to game leaderboards rather than solve real tasks.

Read MoreSpecial Topics in Kernels, RL, Reward Hacking in Agents — Daniel Han, Unsloth

Recursive Model Improvement at Cursor: Lee Robinson on AI-Native Development

Lee Robinson from Cursor explains recursive model improvement using a two-loop framework: an inner loop with high-quality evals to prevent reward hacking, and an outer loop fed by real user feedback from products like Composer 2.5. He also details teacher-student textual feedback methods and scaling compute with SpaceX's Colossus infrastructure for efficient AI-native software development.

Read MoreRecursive Model Improvement at Cursor: Lee Robinson on AI-Native Development