⭐⭐⭐⭐✨ 4.5
Ethan He, formerly of NVIDIA’s Cosmos team and a builder at xAI, shares the inside story of how Grok Imagine was built from scratch in three months by a small team. He contrasts xAI‘s high-iteration culture with typical processes, emphasizing that iteration speed, infrastructure, and talent mattered more than meetings. Many of the biggest quality gains came from fixing small bugs in data and training pipelines, not algorithmic breakthroughs.
The conversation covers the training pipeline for image and video models: VAEs, diffusion transformers, synthetic captioning made by rewriting user prompts, and the challenge of audio-video alignment. Ethan explains why image models are the foundation for video, and how temporal compression and real-time interactivity require careful trade-offs. The hidden costs of video model training—storage, egress, and GPU hours—are detailed, along with methods like step distillation, consistency models, and GANs for fast inference.
On world models, Ethan defines them as real-time, interactive, long-horizon video systems, not just static generation. He discusses reference-to-video, video extension, and long-context video generation, noting that future progress may depend on language models and agents more than on diffusion alone. Prompt rewriting is critical because video models take instructions literally.
Ethan explores the vision of generative UI—interfaces that go directly from user intent to pixels, exemplified by Flipbook and Neural OS. Video agents are seen as the next frontier, where AI-assisted creation combines editing, generation, and agentic control. He also covers safety topics like AI watermarking with SynthID and detecting generated media.
Finally, he reflects on xAI‘s culture, first-principles thinking, working with Elon, and why research communication often undersells the work. He explains why he left xAI to focus more on LLMs and self-managed context, believing language models will unlock the next wave of video generation and embodied AI.
Ethan He, formerly of NVIDIA’s Cosmos team and a builder at xAI, shares the inside story of how Grok Imagine was built from scratch in three months by a small team. He contrasts xAI‘s high-iteration culture with typical processes, emphasizing that iteration speed, infrastructure, and talent mattered more than meetings. Many of the biggest quality gains came from fixing small bugs in data and training pipelines, not algorithmic breakthroughs.
The conversation covers the training pipeline for image and video models: VAEs, diffusion transformers, synthetic captioning made by rewriting user prompts, and the challenge of audio-video alignment. Ethan explains why image models are the foundation for video, and how temporal compression and real-time interactivity require careful trade-offs. The hidden costs of video model training—storage, egress, and GPU hours—are detailed, along with methods like step distillation, consistency models, and GANs for fast inference.
On world models, Ethan defines them as real-time, interactive, long-horizon video systems, not just static generation. He discusses reference-to-video, video extension, and long-context video generation, noting that future progress may depend on language models and agents more than on diffusion alone. Prompt rewriting is critical because video models take instructions literally.
Ethan explores the vision of generative UI—interfaces that go directly from user intent to pixels, exemplified by Flipbook and Neural OS. Video agents are seen as the next frontier, where AI-assisted creation combines editing, generation, and agentic control. He also covers safety topics like AI watermarking with SynthID and detecting generated media.
Finally, he reflects on xAI‘s culture, first-principles thinking, working with Elon, and why research communication often undersells the work. He explains why he left xAI to focus more on LLMs and self-managed context, believing language models will unlock the next wave of video generation and embodied AI.