Building Closed-Loop Evals for a Multimodal Agent at Scale

Highlights

03:00

Generic visual metrics can steer models away from faithfulness in narrow-domain multimodal evaluation.

03:07

Overly tight safety guardrails homogenize the market, while loose ones lose dish authenticity.

16:40

Offline evaluation must be closed-loop with online business signals to prevent reward hacking.

⭐⭐⭐✨ 3.8

Uber‘s food enhancement agent edits photography for smaller, independent Uber Eats merchants. The system must stay faithful to the original dish, preserve each merchant’s brand and packaging, and avoid homogenizing the marketplace—all without an existing playbook for multimodal evals in a narrow domain. Soumya Gupta and Jai Chopra explain how they navigated reward hacking, built a closed feedback loop combining offline and online signals, and balanced creativity against rigid safety guardrails at scale.

Key challenges included designing evals that capture both visual appeal and faithfulness to the original dish, and avoiding reward hacking where the agent optimizes for eval scores at the expense of real-world quality. The speakers share practical strategies for narrow-domain multimodal evaluations, countering reward hacking, and production feedback loops. ML and applied AI practitioners working on multimodal systems, agentic pipelines, or eval design will take away actionable insights from Uber‘s experience.

Uber‘s food enhancement agent edits photography for smaller, independent Uber Eats merchants. The system must stay faithful to the original dish, preserve each merchant’s brand and packaging, and avoid homogenizing the marketplace—all without an existing playbook for multimodal evals in a narrow domain. Soumya Gupta and Jai Chopra explain how they navigated reward hacking, built a closed feedback loop combining offline and online signals, and balanced creativity against rigid safety guardrails at scale.

Key challenges included designing evals that capture both visual appeal and faithfulness to the original dish, and avoiding reward hacking where the agent optimizes for eval scores at the expense of real-world quality. The speakers share practical strategies for narrow-domain multimodal evaluations, countering reward hacking, and production feedback loops. ML and applied AI practitioners working on multimodal systems, agentic pipelines, or eval design will take away actionable insights from Uber‘s experience.

Building Closed-Loop Evals for a Multimodal Agent at Scale — Soumya Gupta & Jai Chopra, Uber

View Original