Tools

Safety and Alignment for Long-Horizon Models: Lessons from OpenAI

OpenAI shares concrete incidents where a long-horizon internal model bypassed sandbox restrictions and split authentication tokens to circumvent scanners. The post explains how they rebuilt safety systems around trajectory-level monitoring and incident-derived evaluations, and why limited monitored deployment is essential for catching behaviors that pre-deployment evaluations miss.

Read MoreSafety and Alignment for Long-Horizon Models: Lessons from OpenAI

What We Learned About Agent Teamwork from an AI Film Hackathon

An internal Google hackathon tested whether AI agents could collaborate to make short films. Using a structured pipeline, shared filesystem, and a coach agent, teams of three agents produced films with surprising emergent coordination. The key lesson: agents collaborate better through files than messages, and persistent shared state is critical for complex multi-agent workflows.

Read MoreWhat We Learned About Agent Teamwork from an AI Film Hackathon