
Borderless Lakehouse: Google Cloud’s unified data and AI agent platform

Google Cloud announced enhancements to its borderless Lakehouse at Next Tokyo, aiming to transform the data lakehouse from a passive repository into a system of action for autonomous AI agents. Traditional data architectures suffer from high costs, fragmented security, and weak governance, making it difficult for agents to access the full data estate. The borderless Lakehouse, built on open Apache Iceberg, connects on-premises, cross-cloud, and SaaS application data without requiring movement.
The core enabler is catalog federation (preview for AWS Glue, Databricks Unity, and Snowflake Horizon) using the Iceberg REST catalog, allowing instant discovery and query of remote data. Zero-copy integrations with SAP, Salesforce, and Workday let BigQuery query live application data directly, enabling unified analytics across finance, HR, and customer data. Key benefits include zero-copy cross-cloud analytics, bidirectional interoperability (read/write across environments), and unified governance with credential vending at the table level.
To overcome cross-cloud egress fees and latency, the borderless Lakehouse introduces Cross-Cloud Interconnects — private dedicated links with predictable monthly pricing and SLA-backed connections, supporting from 1G to 100G. Intelligent cross-cloud caching temporarily stores remote data fragments in Google Cloud to avoid repeated transfers. Combined with BigQuery’s vectorized processing, AI functions, and Spark Lightning Engine, this allows in-place AI/ML on AWS and Azure data without migration.
For agent trust and accuracy, the Knowledge Catalog acts as an always-on agentic context engine. It ingests metadata from federated catalogs (Glue, Unity, Horizon), translates raw schemas into business terminology, and maintains column-level lineage. This provides a semantic layer and automated governance, ensuring agents respect access permissions and compliance guardrails.
Developers can build and publish custom data agents using the open-source Google Cloud Data Agent Kit and Conversational Analytics API, integrating directly into Gemini Enterprise. Built-in MCP tools connect to BigQuery, Managed Spark, and Cloud Storage, eliminating pipeline code. Business users can then query multi-cloud datasets in natural language.
The economics are addressed via flat-rate cross-cloud interconnects, zero-copy data sharing, and token controls in BigQuery AI. Knowledge Catalog filters context to prevent token bloat, and BigQuery AI can automatically use smaller, distilled models. The article claims customers see 230x reduction in token consumption using BigQuery’s cost-optimized built-in AI functions. Stated limitations include an hourly fee for interconnection service.


