
Box Unlocks Multimodal Enterprise Agents with Gemini Embeddings 2

Enterprise content management is undergoing its biggest architectural shift since cloud migration, according to Box. The company argues that traditional text-based search and retrieval-augmented generation (RAG) have successfully unlocked narrative knowledge in enterprise repositories, but the agentic era demands more: capturing the inherently multimodal, spatial, and structured elements that live alongside text. Text embeddings excel at indexing prose, but they lose the strict row-column semantics of financial tables, visual evidence in clinical data, and the spatial layout of multi-page flowcharts.
To address this, Google Cloud and Box are integrating Gemini Multimodal Embeddings 2 into Box‘s Agentic Platform, merging Box‘s Intelligent Content Management platform with Google Cloud’s advanced AI embeddings. The article outlines several benefits of improved embedding: preserving visual and spatial geometry so complex elements like multi-column tables keep their meaning, illuminating visual modality so charts, flowcharts, and product photography become searchable, and connecting hybrid file formats so an agent can cross-reference a PDF policy, a spreadsheet log, and a presentation deck in a unified way.
The architectural solution is Gemini Multimodal Embeddings 2, which introduces a unified multimodal vector space capable of embedding text, raster images, document pages, rendered spreadsheet tables, and visual charts into the same semantic representation. Key capabilities include crossmodal retrieval (text-to-visual and visual-to-text), layout-aware document embedding that preserves visual hierarchies and structural context, and heterogeneous format bridging across .docx, .xlsx, .pdf, .pptx, .png, and .csv without losing modality-specific structural information.
The article then identifies three primary design patterns for multimodal enterprise agents. Pattern 1 is complex financial and analytical reporting. Corporate finance, research, and audit teams analyze structured documents with embedded tables, growth charts, and footnote annotations; text-only indexing separates numbers from context. Multimodal embeddings enable structural alignment, visual trend analysis, and contextual sourcing. Pattern 2 is multimodal clinical decision support and assisted diagnosis. Healthcare data spans physical photos, microscopic pathology slides, and structured risk matrices; multimodal synthesis evaluates symptoms alongside lab evidence, connects niche visual patterns to medical knowledge for rare conditions, and provides risk-aware warnings. Pattern 3 is cross-document multimodal synthesis and data reconciliation. Enterprise information is fragmented across PDFs, Excel charts, PNG flyers, and email threads; multimodal embeddings enable cross-file synthesis, conflict resolution, and visual-to-text auditing.
The article concludes by framing the integration as a move beyond basic search to active, intelligent collaboration. Box‘s Intelligent Content Management platform becomes a governed, semantically indexed reasoning layer where AI agents can interrogate, cross-reference, and act on content with compliance and security controls in place. For financial services, life sciences, and legal operations, multimodal understanding is presented as a competitive requirement. The piece emphasizes that the enterprise data landscape was always multimodal, and now technology exists to make the most of it. Product leaders who adopt multimodal-first architectures, rigorous precision benchmarking, and audit-ready grounding will lead the next wave of productivity. The authors thank Ken Ikeda, Afshaan Mazagonwalla, and Samip Thakkar for their work.


