AI Applications

What We Learned About Agent Teamwork from an AI Film Hackathon

An internal Google hackathon tested whether AI agents could collaborate to make short films. Using a structured pipeline, shared filesystem, and a coach agent, teams of three agents produced films with surprising emergent coordination. The key lesson: agents collaborate better through files than messages, and persistent shared state is critical for complex multi-agent workflows.

Read MoreWhat We Learned About Agent Teamwork from an AI Film Hackathon

Cars24 scales conversations and builds faster with OpenAI agents

Cars24 shows how a complex, conversation-driven marketplace can use OpenAI's voice agents and Codex to handle over a million conversation minutes per month, while letting finance, legal, and operations teams build their own workflows. The key was starting with the most friction-heavy parts of the sales funnel and letting adoption spread organically.

Read MoreCars24 scales conversations and builds faster with OpenAI agents

DharmaOCR: Why Specialization Still Beats Newer Models on Portuguese

Despite newer architectures like Mistral OCR4 and Unlimited-OCR, DharmaOCR—a model specialized for Brazilian Portuguese—still outperforms them on Portuguese documents by a significant margin. The article explains how a two-stage training pipeline of supervised fine-tuning and Direct Preference Optimization yields higher accuracy and stability, backed by concrete benchmark scores and real-world failure mode analysis. Three months after release, the lesson is clear: domain concentration remains a decisive structural advantage even as general models scale.

Read MoreDharmaOCR: Why Specialization Still Beats Newer Models on Portuguese

Real World VoiceEQ: Benchmarking the Human Quality of Voice AI

Hume's Real World VoiceEQ benchmark, built from over 1 million human ratings, reveals that traditional voice AI metrics overestimate real-world performance. Models excel at speaking but struggle with listening—missing tone, hesitation, and emotion. The findings challenge the idea of a single best voice model, urging builders to prioritize human-grounded evaluation for real conversational quality.

Read MoreReal World VoiceEQ: Benchmarking the Human Quality of Voice AI