<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:media="http://search.yahoo.com/mrss/"><channel><title>Data Analytics</title><link>https://cloud.google.com/blog/products/data-analytics/</link><description>Data Analytics</description><atom:link href="https://cloudblog.withgoogle.com/blog/products/data-analytics/rss/" rel="self"></atom:link><language>en</language><lastBuildDate>Tue, 06 Oct 2026 16:00:06 +0000</lastBuildDate><image><url>https://cloud.google.com/blog/products/data-analytics/static/blog/images/google.a51985becaa6.png</url><title>Data Analytics</title><link>https://cloud.google.com/blog/products/data-analytics/</link></image><item><title>Managed Apache Iceberg at scale: How Spanner powers Lakehouse runtime catalog</title><link>https://cloud.google.com/blog/products/data-analytics/lakehouse-runtime-catalog-powered-by-spanner/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;As customers modernize to lakehouse architectures, they are standardizing on open formats such as Apache Iceberg to create a shared data estate across compatible engines. This enables you to build AI-native lakehouses that turn data to semantic knowledge, enable proactive action, and operate at an agentic scale.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;One of the key components of a lakehouse is the catalog, and in the Apache Iceberg environment, that usually means the Iceberg REST Catalog. An Apache Iceberg catalog is responsible for maintaining table pointers, handling atomic commits, and serving as the single source of truth for table locations. But as large enterprise organizations modernize to lakehouses, they have begun to realize that they need a highly scalable and available managed catalog as a part of their lakehouse architecture. This becomes even more important as querying scales with agents. To support a large amount of repeated small queries from agents, you will need to build on top of a managed catalog that provides atomicity, consistency, availability and concurrency at massive scale. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In this blog, we explore the challenges a managed catalog faces in modern cloud environments at agent scale, and show you how Google Cloud’s serverless &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/biglake-metastore-now-supports-iceberg-rest-catalog?e=a"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Lakehouse runtime catalog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; can help address them. Powered by &lt;/span&gt;&lt;a href="https://cloud.google.com/spanner"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Spanner&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, Google Cloud’s always-on database with virtually unlimited scale, and built to meet the open Apache Iceberg REST catalog specification, the Lakehouse runtime catalog is the highly scalable and available foundation you need for the agentic era.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Challenges a managed catalog needs to solve in the Lakehouse&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When speaking with data engineers and infrastructure leads running production analytics at scale in a Lakehouse, the following core pain points consistently emerge and need to be solved by a managed catalog: &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Atomic commits and concurrency control:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Iceberg guarantees ACID transactions via optimistic concurrency control (OCC). A managed catalog must implement a bulletproof atomic compare-and-swap (CAS) operation to swap the current metadata pointer.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;High availability and operational maintenance:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Because queries fail immediately if the catalog is down, a managed catalog becomes a critical Tier-1 service.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Scaling the database backing the catalog:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Catalog architects typically face a difficult trade-off when choosing a backing database for table metadata and state. Traditional scale-up relational databases provide SQL and ACID transactions, but hit vertical CPU, memory, storage and connection limits under heavy concurrent read/write loads unless manually sharded which incurs a huge operational overhead; while scale-out database systems are either eventually consistent, hard to manage, not enterprise-ready, or all of the above. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Table maintenance coordination:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; A catalog alone does not optimize data; you must build and operate ancillary pipelines for compaction, snapshot expiration, manifest rewriting, and orphan file cleanup.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Governance and security:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; A catalog acts as the security gatekeeper. The catalog must implement and maintain:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Authentication protocols (e.g., OAuth2 token exchange, IAM federation)&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Access control down to namespace and table levels&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Vended storage credentials (e.g., generating short-lived tokens so query engines don't need broad, direct storage credentials)&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Lakehouse runtime catalog&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To solve these challenges, we built the Lakehouse runtime catalog (GA) with support for Iceberg Rest Catalog. The Lakehouse runtime catalog is a fully serverless, highly available, and unified metadata registry designed from the ground up to support modern open table formats like&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Apache Iceberg. By natively implementing the Apache Iceberg REST Catalog specification&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;,&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; the Lakehouse runtime catalog decouples metadata discovery from compute engines, helping ensure multiple Iceberg-compatible engines can access a shared data estate and enabling you to take your workloads to production sooner. We’ve helped many customers streamline the migration of their managed catalogs. For example, &lt;/span&gt;&lt;a href="https://cloud.google.com/customers/etsy-lakehouse"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Etsy&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; migrated its catalog to Lakehouse runtime catalog, joining data in place to accelerate pipeline queries by 60%.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_bmVFk8w.max-1000x1000.jpg"
        
          alt="1"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This approach offers a number of architectural benefits:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Open APIs&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Support for Iceberg Rest Catalog enables different teams to use their preferred analytics tools on a single, unified dataset.&lt;/span&gt;&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Multi-engine interoperability&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Once registered, tables are immediately discoverable and queryable across Google Cloud Managed Service for Apache Spark, BigQuery, and open-source engines via standard REST interfaces. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/the-future-of-data-lakehouse-for-the-agentic-era?e=0"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Read/write interoperability&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt; for Iceberg tables: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Leverage Iceberg-compatible engines such as BigQuery, Managed Spark to write to Iceberg tables registered in the Lakehouse runtime catalog. Customers can also use Managed Spark to write to Iceberg tables in external catalogs.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Fully managed Iceberg storage with enterprise-grade features: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Use Google's &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/unveiling-new-bigquery-capabilities-for-the-agentic-era?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;differentiated infrastructure&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to run analytics with performance on Iceberg tables. This gives you the benefits of open-source flexibility plus performance, scale, governance, and multimodal processing. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Zero data copy&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Table definitions point directly to your existing data in the underlying object store. You do not move, rewrite, or duplicate your underlying data.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Bi-directional catalog federation&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;across clouds:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Access data from &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/lakehouse/docs/set-up-cross-cloud-connection-databricks"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Databricks Unity&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/lakehouse/docs/set-up-cross-cloud-connection-snowflake"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Snowflake Horizon&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/lakehouse/docs/set-up-cross-cloud-connection-aws-glue"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;AWS Glue&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; with support for vended credentials and OIDC token exchange. This lets you bring Google AI &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/introducing-the-borderless-lakehouse?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;directly to your AWS and Azure data&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Secure access using credential vending&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The catalog supports multiple authorization mechanisms, letting you choose between &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/lakehouse/docs/credential-vending"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;credential vending&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and end-user credentials. This means that you can access tables with modern mechanisms such as credential vending without needing direct access to the files in the underlying object store (Cloud Storage, AWS S3, Azure Blob Storage). &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;AI-powered context and governance: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;The Lakehouse runtime catalog integrates directly with &lt;/span&gt;&lt;a href="https://cloud.google.com/products/knowledge-catalog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Knowledge Catalog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://cloud.google.com/products/iam?e=0"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud IAM&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, allowing you to define trusted context for your agents, and apply table-level security consistently across all compute engines.&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Get out-of-the-box search, lineage, and insights for Iceberg tables in the catalog. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Atomic commits and concurrency control, high availability and scalability: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Backed by Google’s planet-scale infrastructure and Spanner, you get the high availability, concurrency and scale you need for your metadata to scale with your data. Support for Cloud Storage dual-region and multi-region buckets enables failover use cases. It also provides reduced TCO due to serverless and no-ops environments, and scalability for any workload size.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;What powers the Lakehouse runtime catalog?&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The Lakehouse runtime catalog is a highly available, concurrent and scalable catalog with strong consistency guarantees because it is built on top of &lt;/span&gt;&lt;a href="https://cloud.google.com/spanner"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Spanner&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. Unlike traditional scale-up relational databases that hit vertical single-node ceilings, Spanner combines full relational SQL semantics and multi-table ACID transactions with the horizontal scale-out elasticity for both reads and writes of a NoSQL system. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;Spanner makes the Lakehouse runtime catalog highly available through Spanner’s &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/spanner/docs/instance-configurations#regional-configurations"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;regional configurations&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; with up to 99.99% availability. Spanner also delivers out-of-the-box scalability for  Lakehouse runtime catalog: As a horizontally scalable database, Spanner does not require manual sharding and scales compute and storage independently and transparently. Spanner dynamically monitors data volume and query load, splitting and redistributing data ranges across nodes. Compute nodes scale dynamically based on CPU utilization and storage thresholds. Spanner automatically detects split-level overload and moves heavy splits away from overloaded nodes. Spanner also lets Lakehouse runtime catalog users eliminate the &lt;/span&gt;&lt;a href="https://dl.acm.org/doi/abs/10.1145/2491245?__cf_chl_tk=oi5rkZKtX5q2CR93K9mrIHsAp_oqRxwaGqqpqq99v9Y-1790923307-1.0.1.1-pNUQvRr4OBosWiOcfxGuTWTD8I3JW_8V0i5xLuCsErU" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;traditional trade-offs&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; between relational consistency and distributed scalability, delivering Lakehouse runtime catalog’s industry-leading &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/spanner/docs/true-time-external-consistency"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;consistency guarantees&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. Lakehouse transactions are serializable — the order of transactions within the database is the same as the order in which clients observe the transactions to have been committed. &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;This foundation allows Lakehouse users to operate at agent-scale. &lt;/span&gt; &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Then, to further power agentic use cases, the Lakehouse runtime catalog integrates directly with Knowledge Catalog to easily discover lakehouse Iceberg tables and provide trusted context to agents. Knowledge Catalog leverages an efficient combination of &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/spanner/docs/full-text-search"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;full-text search&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and native&lt;/span&gt;&lt;a href="https://docs.cloud.google.com/spanner/docs/vector-search-overview"&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;vector search&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; provided by Spanner; this approach enables better recall for search retrieval, pairing lexical keyword searches with semantic embeddings in a single query. Because both index types are built on the identical base dataset, they update with strict, transactional ACID consistency alongside base table DML operations. This removes operational overhead such as managing sync pipelines, and external-vector and full-text search systems. In short, the Lakehouse runtime catalog provides faster time-to-market for your agentic use cases.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Combine analytical and operational workloads&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With Google Cloud’s borderless Lakehouse based on Apache Iceberg, you can combine your analytical data with your operational workloads. Use cases span combining data assets from your Lakehouse with OLTP data (from Spanner) for analytics, to low-latency serving applications where your Lakehouse assets are accessible in an operational database such as Spanner, to conversational analytics in first-party and third-party agents. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Below, in an example, you can see the Lakehouse runtime catalog with Spanner in action. Here, we combine analytical data for taxi trips in Manhattan (backed by Apache Iceberg) with operational data for taxi zones in Spanner to find the most congested traffic routes. The example also shows that you can also use a Conversational Analytics Agent to access the same data and get second-order insights.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/2_TvlKrOP.gif"
        
          alt="2"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="cjj42"&gt;Find out most congested traffic routes in Manhattan&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Modernize to the borderless Lakehouse&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Modernizing to Google Cloud’s&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://cloud.google.com/solutions/data-lakehouse"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Lakehouse&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;minimizes data silos across analytics engines and agents, unifies multi-engine governance, provides trusted context to your agents and slashes operational TCO. Leveraging the Spanner-based Lakehouse &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/lakehouse/docs/set-up-lakehouse-iceberg-rest-catalog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;runtime catalog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; helps prepare your modern cloud environments to operate at agent-scale. To learn more and get started with a free trial, visit the &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://cloud.google.com/products/lakehouse?e=a"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Lakehouse&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; web&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;page and learn more about Spanner &lt;/span&gt;&lt;a href="https://cloud.google.com/spanner"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Tue, 06 Oct 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/lakehouse-runtime-catalog-powered-by-spanner/</guid><category>Databases</category><category>Data Analytics</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Managed Apache Iceberg at scale: How Spanner powers Lakehouse runtime catalog</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/lakehouse-runtime-catalog-powered-by-spanner/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Vinod Ramachandran</name><title>Product Manager</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Mina Mikhail</name><title>Software Engineering Manager</title><department></department><company></company></author></item><item><title>Accelerating analytics: PayPal’s journey with Managed Service for Apache Spark</title><link>https://cloud.google.com/blog/products/data-analytics/paypals-journey-with-managed-service-for-apache-spark/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In a data-driven world, PayPal’s ability to deliver timely and actionable insights is central to staying ahead. At PayPal, data powers everything from fraud detection to user experience enhancements. Data is also central to unleashing the potential of agentic solutions and experiences. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Over time, though, our analytics environment had become a complex ecosystem of various technologies and solutions assembled on-premise to address growing demands. While this approach supported our needs at the time, it began presenting new challenges to scale and maintain.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Navigating a challenging analytics landscape&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Due to expedited growth and acquisitions, our data analytics platform gradually turned into an uneven landscape. Each new platform or integration addressed a specific business need, but together, they increased operational overhead and introduced performance blockages. Scalability became increasingly difficult, and time-to-insight slowed as processes grew more complex. &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Complexity breeds stagnation&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;PayPal’s &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/databases/paypals-historic-data-migration-is-the-foundation-for-its-gen-ai-innovation"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;legacy data analytics platform&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; was powerful—handling petabytes daily—but it was also increasingly rigid following rapid growth. Scaling up during peak retail events or global launches meant months of planning, slow manual provisioning of hardware, and too often, a compromise between speed and cost.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;As PayPal continued to scale globally, we recognized the need for a streamlined, unified infrastructure to drive data efficiency and accelerate innovation.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;The solution: Unified, cloud-native analytics&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To overcome these obstacles, we migrated our analytics workloads from legacy Hadoop on-premise platforms to Google’s &lt;/span&gt;&lt;a href="https://cloud.google.com/products/managed-service-for-apache-spark"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Managed Service for Apache Spark&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Key reasons for this choice included:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Rapid provisioning and elastic scaling: Managed Spark enabled us to deploy clusters in minutes and scale based on processing needs, eliminating lengthy setup and idle resource costs. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Unified infrastructure: Standardizing on Apache Spark created consistency across teams while leveraging Managed Service for Apache Spark and other managed services reduced operational complexity.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Seamless integration: Native hooks into &lt;/span&gt;&lt;a href="https://cloud.google.com/storage"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Storage&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (GCS), &lt;/span&gt;&lt;a href="https://cloud.google.com/bigquery"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;BigQuery&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and other Google Cloud services streamlined end-to-end data movement.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This move enabled PayPal to modernize our data processing capabilities, leveraging the flexibility, scalability, and reliability of cloud-native solutions. By consolidating previously disparate workflows and batch jobs that run on multiple platforms onto a single cloud-based analytics platform, we reduced data silos and built a unified data foundation that provides faster, richer insights. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This empowered developers and application teams to focus on delivering business value rather than being limited by infrastructure. Crucially, this shift was about more than re-platforming. We fostered a new culture of experimentation, enabling teams to test, tune, and deploy analytics workloads quickly in response to changing business needs.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;The results: Faster insights, lower overhead&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The impact of our modernized Google Cloud-based ecosystem leveraging Managed Spark has been profound:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Processing times for core analytics workloads improved by 25%, enabling near real-time insights for key business operations.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;SLA adherence rose substantially by 30%, even during traffic surges such as seasonal sales events.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Operational costs dropped as we consolidated tooling and reduced manual maintenance.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;But perhaps most importantly, our engineers now spend less time firefighting and more time innovating, rapidly prototyping new analytics capabilities that deliver value to customers and partners.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Transitioning from a fragmented environment to a cohesive, cloud-native platform has fundamentally strengthened PayPal’s analytics capabilities. As business needs evolve, investing in a scalable, unified data foundation ensures that we can deliver insights with speed, precision, and impact—driving continued innovation for customers worldwide. Our journey with Managed Service for Apache  Spark is an important step in building that modern analytics foundation.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;Learn more about how you can get started with &lt;/span&gt;&lt;a href="https://cloud.google.com/products/managed-service-for-apache-spark"&gt;&lt;span style="font-style: italic; text-decoration: underline; vertical-align: baseline;"&gt;Managed Service for Apache Spark&lt;/span&gt;&lt;/a&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://cloud.google.com/bigquery"&gt;&lt;span style="font-style: italic; text-decoration: underline; vertical-align: baseline;"&gt;BigQuery&lt;/span&gt;&lt;/a&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt; to build your &lt;/span&gt;&lt;a href="https://cloud.google.com/data-cloud"&gt;&lt;span style="font-style: italic; text-decoration: underline; vertical-align: baseline;"&gt;Agentic Data Cloud&lt;/span&gt;&lt;/a&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt; today.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Thu, 01 Oct 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/paypals-journey-with-managed-service-for-apache-spark/</guid><category>Financial Services</category><category>Data Analytics</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/paypal-apache-spark.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Accelerating analytics: PayPal’s journey with Managed Service for Apache Spark</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/paypal-apache-spark.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/paypals-journey-with-managed-service-for-apache-spark/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Vinod Ganesan</name><title>Director, Analytics Reliability Engineering, PayPal</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Raghu Agani</name><title>Sr. Manager, Big Data Platforms Engineering, PayPal</title><department></department><company></company></author></item><item><title>Empower your agents with the Google Cloud CLI remote MCP server</title><link>https://cloud.google.com/blog/products/ai-machine-learning/google-cloud-cli-remote-mcp-server-in-preview/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Today, we’re expanding our ecosystem of managed remote MCP servers by introducing the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/sdk/use-gcloud-mcp"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud CLI remote MCP server&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; in preview.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Powered by the popular &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/sdk/gcloud"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;gcloud&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/reference/bq-cli-reference"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;bq (BigQuery)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; command-line tools, this new server gives AI agents immediate, broad access to command-line operations for managing Google Cloud infrastructure and working with advanced BigQuery workflows securely and seamlessly. &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Why CLI matters for AI agents&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Agents are increasingly performing complex cloud operations, but standardizing how they interact with backend systems remains a challenge. The Google Cloud CLI remote MCP server bridges this gap by packaging the versatility of hundreds of gcloud and bq commands into one single MCP server. This results in two strong benefits for the agent:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Higher-level abstractions:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; CLI commands package complex multi-step workflows, validation checks, and high-level operations into unified commands rather than requiring multi-step API orchestration.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Leverages model training:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; LLMs are heavily pre-trained on public command-line documentation, syntaxes, and usage examples, making CLI invocation intuitive and highly accurate for models.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;The benefits of putting CLI behind remote MCP&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Managing cloud infrastructure with AI agents traditionally requires installing and maintaining Google Cloud CLI binaries inside agent execution environments. The Cloud CLI remote MCP server bridges CLI capabilities with MCP benefits by providing an isolated execution sandbox on Google Cloud infrastructure. This solves key infrastructure challenges:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Simplified dependency and runtime management:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; For teams building custom agents, maintaining local CLI versions and dependencies across dev, test, and production environments creates operational overhead. Remote MCP eliminates local installations and runtime maintenance.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Access for web-based agent endpoints:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Web-hosted agent platforms and web interfaces (such as Gemini Enterprise and other hosted enterprise agent platforms) run in environments where users cannot control or install local packages. Remote MCP enables secure, managed access to Google Cloud CLI operations directly from these surfaces.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Enterprise-grade security and governance&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Connecting an AI agent to your infrastructure requires strict, enterprise-ready safeguards. This remote server leverages Google Cloud's standard identity and governance frameworks to keep your environments secure:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Zero ambient credentials:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The server isolates execution in a network-restricted proxy boundary with no ambient credentials. Authentication and authorization are handled through &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/iam/docs/agent-identity-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Agent Identity&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://developers.google.com/identity/protocols/oauth2" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;OAuth 2.0&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/iam/docs"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Identity and Access Management (IAM)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Strict policy enforcement:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Every command executed through the remote MCP server is run with the permissions of the authenticated caller identity. Both standard IAM permissions and organization policy service constraints are strictly enforced against downstream target resources.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Advanced protection with Model Armor:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; To minimize the risks associated with AI tool calling, the Cloud CLI remote MCP server integrates with &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/model-armor/model-armor-mcp-google-cloud-integration"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Model Armor&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. You can proactively screen LLM prompts and responses to protect against risks like prompt injection and malicious inputs.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Cloud audit logging:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The Cloud CLI remote MCP server can be configured to log every tool invocation to Audit Logs (Data Access logs under cloudcli.googleapis.com/mcp). Security teams can gain full visibility into caller identities, OAuth clients, and IAM authorization decisions (mcp.googleapis.com/tools.call) without exposing sensitive command payloads or personally identifiable information (PII).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Connecting to the Google Cloud CLI Remote MCP Server&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Integrating cloud management into your agents no longer requires packaging Google Cloud CLI binaries, managing local execution runtimes, or maintaining dependencies inside agent container images. Because the Google Cloud CLI remote MCP server implements the standard Model Context Protocol, any MCP-compatible agent platform or orchestration runtime can connect immediately via standard configuration:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;{\r\n  &amp;quot;mcpServers&amp;quot;: {\r\n    &amp;quot;google-cloud-cli&amp;quot;: {\r\n      &amp;quot;uri&amp;quot;: &amp;quot;https://cloudcli.googleapis.com/mcp&amp;quot;,\r\n      ...\r\n    }\r\n  }\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca352f1010&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Your agent immediately gains access to execute gcloud and bq commands in a secure, network-isolated cloud sandbox.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Authentication is handled via keyless Agent Identity for hosted Google Cloud platforms, or standard OAuth 2.0 for external runtimes. For authentication options, see the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/mcp/set-up-authentication-mcp-servers"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;MCP Authentication Guide&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Bringing infrastructure management to agents&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;a href="https://docs.cloud.google.com/sdk/use-gcloud-mcp"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;The Cloud CLI remote MCP server&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; exposes two powerful tools, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;run_gcloud_command&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;run_bq_command&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, giving your AI agents broad, immediate access to Google Cloud operations through natural language.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Managing cloud infrastructure with &lt;/strong&gt;&lt;code&gt;&lt;strong style="vertical-align: baseline;"&gt;run_gcloud_command&lt;/strong&gt;&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;run_gcloud_command&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, agents can execute the full breadth of gcloud operations to manage, diagnose, and secure your Google Cloud environment. An example follows:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Observability and incident diagnostics&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: An agent streamlines incident diagnostics by automating command execution and reducing context-switching across tools.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/1_IPA9ALa.gif"
        
          alt="1"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Extending BigQuery operations with the run_bq_command tool&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;While the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/use-bigquery-mcp"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;BigQuery MCP server&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; already helps organizations analyze and explore data using AI agents, with the introduction of &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;run_bq_command&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, agents can now tackle advanced BigQuery tasks such as resource allocation, job monitoring, and task scheduling by unlocking the full scope of &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/sdk/reference/mcp#mcp-tools"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;bq CLI&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; functionality. Key capabilities include:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Automating scheduled queries&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: An agent utilizes BigQuery Data Transfer Service configurations to schedule queries automatically. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Job and resource management&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Gain deep insight into query execution details, including processed data volume, slot usage, and execution plans, as well as managing reservations.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Access and permissions control:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Data administrators and owners can inspect and update table permissions directly through the agent.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/2_JA5HWbH.gif"
        
          alt="2"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Pricing and availability&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The Google Cloud CLI MCP server is available today in public preview. There is no additional charge to use the MCP server itself. You pay only for the GCP resources you create and any applicable data transfer costs.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://docs.cloud.google.com/sdk/use-gcloud-mcp"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;To get started&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, enable the Cloud CLI Execution API (`&lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;cloudcli.googleapis.com&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;`) in your Google Cloud project, grant the required &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;MCP Tool User&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; (`&lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;roles/mcp.toolUser&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;`) IAM role to your agent or user identity, and configure your MCP client to connect to `cloudcli.googleapis.com/mcp`.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/sdk/use-gcloud-mcp"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Use the Google Cloud CLI Remote MCP Server Guide&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/sdk/reference/mcp#mcp-tools"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud CLI MCP Reference&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/mcp/overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Remote MCP Servers Overview&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/mcp/supported-products"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Full List of Google OneMCP Servers&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/mcp/set-up-authentication-mcp-servers"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Set up authentication to Google and Google Cloud MCP servers&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/gemini-enterprise-agent-platform/govern/agent-identity-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise Agent Platform &amp;amp; Agent Identity Overview&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://www.youtube.com/watch?v=-fb0ycu4kiU" rel="noopener" target="_blank"&gt;&lt;span data-rich-links='{"fple-t":"Automate Google Cloud with Cloud CLI Remote MCP Server","fple-u":"https://www.youtube.com/watch?v=-fb0ycu4kiU","fple-mt":null,"type":"first-party-link"}' style="text-decoration: underline; vertical-align: baseline;"&gt;Automate Google Cloud with Cloud CLI Remote MCP Server&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Wed, 30 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/ai-machine-learning/google-cloud-cli-remote-mcp-server-in-preview/</guid><category>Data Analytics</category><category>Developers &amp; Practitioners</category><category>AI &amp; Machine Learning</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Empower your agents with the Google Cloud CLI remote MCP server</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/ai-machine-learning/google-cloud-cli-remote-mcp-server-in-preview/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Prosper Nwankpa</name><title>Senior Engineering Manager, Google Cloud</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Adam Hwang</name><title>Software Engineering Manager, Google Cloud</title><department></department><company></company></author></item><item><title>Introducing Ask, a new Google Earth Engine feature to accelerate geospatial coding</title><link>https://cloud.google.com/blog/products/data-analytics/accelerate-geospatial-coding-with-ai-in-google-earth-engine/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Whether you are mapping global forest cover, detecting changes in the built environment, or monitoring agricultural yields, writing scripts in &lt;/span&gt;&lt;a href="https://earthengine.google.com/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Earth Engine&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is a powerful way to develop these insights. This platform is part of &lt;/span&gt;&lt;a href="https://ai.google/earth-ai/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Earth AI&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, our collection of geospatial models and datasets designed to help you transform planetary information into actionable intelligence, and we’ve launched a new feature, Ask, that makes accessing that planetary intelligence even easier. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We know that translating complex geospatial logic into code takes time, and memorizing specific Google Earth Engine API functions, searching documentation, and debugging syntax or memory errors can interrupt your flow. Ask is designed to help you get to actionable insights faster, by integrating Gemini capabilities directly into the Google Earth Engine Code Editor.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Starting today, you can use your own Gemini API key to write, debug, understand, and optimize geospatial queries without ever leaving the Code Editor.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_hhaY4vo.max-1000x1000.jpg"
        
          alt="1"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="c2mdz"&gt;Figure 1: The new Ask panel resides on the right side of the Code Editor, providing chat-based AI assistance tailored to your active script.&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;How it works: Context-aware assistance&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;You can ask a question directly in the Code Editor, and get a response based on context from your workspace, such as:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;The full text of your active script&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Your imported assets and geometries&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Your active session chat history&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Because Gemini capabilities in Google Earth Engine automatically understand your work, you don’t have to add code or explain your datasets. It already knows these details, so it can provide more relevant, helpful answers. Additionally, you can automatically view the diff and merge it into your script. No more copy and pasting!&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;You can also personalize Ask to make responses more comprehensive and suited to your needs:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Access the latest Gemini models:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Choose the Gemini model that meets your needs — at launch, the available models are Gemini 3 Flash Preview, Gemini 3.1 Pro Preview, and Gemini 3.5 Flash.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Docs search:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Use AI to search the official &lt;/span&gt;&lt;a href="https://developers.google.com/earth-engine" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;documentation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for up-to-date syntax and functions.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Dataset search:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Allow AI to search the &lt;/span&gt;&lt;a href="https://developers.google.com/earth-engine/datasets" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Earth Engine Data Catalog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to find and reference the exact datasets you need.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Google Search:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Grounds response in the latest public web results using Grounding with Google Search (mutually exclusive with docs search and dataset search).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Quick-start examples&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Not sure where to start? Here are four ways you can use Ask in your workflows:&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;1. Generate code from natural language&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Need to calculate NDVI (“Normalized Difference Vegetation Index”), perform a cloud mask, or build a chart, but can't remember the exact syntax? Just ask and Earth Engine can generate JavaScript code for you. You can review the code directly in the chat and click &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Insert&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; to add it to your script, or &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Copy&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; it to your clipboard. If your editor is not empty, inserting code displays a side-by-side diff view so you can review changes before accepting them.&lt;/span&gt;&lt;/p&gt;
&lt;p style="padding-left: 40px;"&gt;&lt;strong style="vertical-align: baseline;"&gt;Prompt example:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;"Write a script to load Sentinel-2 imagery for 2025 over Boulder, CO, apply a cloud mask, calculate NDVI, and add the median composite to the map."&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;2. Explain complex code&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If you’re working with a script written by a colleague or adapting an example from the community, you can use Gemini capabilities to explain it. Simply ask, &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;"Explain what this script does,"&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; and receive a step-by-step breakdown of the logic and Earth Engine functions being used.&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;3.  Perform one-click troubleshooting and debugging&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Debugging is an inevitable part of coding. When your script throws an error in Google Earth Engine, the Console displays an error message. Now, those messages include a &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;"Troubleshoot"&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; button. Clicking it automatically populates the Ask panel with a prompt that contains the error message and a request for help to diagnose the issue and suggest a fix.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_qed5vZZ.max-1000x1000.jpg"
        
          alt="2"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="gvxcy"&gt;Figure 2: The "Troubleshoot" button in the Console makes debugging errors fast and frictionless.&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;4. Optimize queries&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If your script is, for example, running slowly, consuming resources inefficiently, throwing "computation timed out" or "too many concurrent aggregations" errors, ask for optimization tips. Gemini capabilities can suggest best practices like early filtering, reducing computation steps, or converting client-side loops into server-side operations.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Get started in two steps&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Ask is available globally today. To get started, you just need a Gemini API key:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Get an API key:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; If you don’t already have a Gemini API key, head to &lt;/span&gt;&lt;a href="https://aistudio.google.com/app/apikey" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google AI Studio&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and create one (there are free options; if you choose a paid-tier API key, you will be billed for your usage). Learn more about &lt;/span&gt;&lt;a href="https://ai.google.dev/terms" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini API terms&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Add the API key to GEE:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Open the Google Earth Engine Code Editor. On the right-hand panel, click the &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Ask&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; panel. Click the key icon in the bottom left corner and enter your API key.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We want to hear from you! Please use the Code Editor Feedback button to share your feedback and help us improve the experience.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Mon, 28 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/accelerate-geospatial-coding-with-ai-in-google-earth-engine/</guid><category>Maps &amp; Geospatial</category><category>Data Analytics</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Introducing Ask, a new Google Earth Engine feature to accelerate geospatial coding</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/accelerate-geospatial-coding-with-ai-in-google-earth-engine/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Joel Conkling</name><title>Product Lead</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Katie Friis</name><title>Software Engineer</title><department></department><company></company></author></item><item><title>Maximizing Apache Spark availability: Mitigating compute stockouts with flexible VMs and other best practices</title><link>https://cloud.google.com/blog/products/data-analytics/maximize-apache-spark-availability-with-flexible-vms/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The surge in AI development has created unprecedented demand for compute capacity around the globe. This can have negative implications for data processing and pipelines with Apache Spark. Whether you are managing your own Spark infrastructure or using a managed service, you can face availability constraints. However, a significant advantage of using Google’s &lt;/span&gt;&lt;a href="https://cloud.google.com/products/managed-service-for-apache-spark"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Managed Service for Apache Spark&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is the availability of &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/managed-spark/docs/concepts/configuring-clusters/flexible-vms"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;flexible VMs,&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; which provide a targeted mechanism to adopt a dynamic, resource-agnostic philosophy and ensure your pipelines remain operational, even during regional or zonal capacity stockouts.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Understanding capacity stockouts&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Capacity stockouts occur when demand for a specific machine family (such as N2 or N2D) exceeds available capacity in a target zone or region. For time-sensitive analytics pipelines, rigid single-VM requirements transform standard provisioning into a single point of failure which can result in cluster creation delays, failed executions, and potentially compromised business SLAs.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Flexible VMs&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Flexible VMs fundamentally overhaul how a Managed Spark cluster requests compute resources. Rather than binding a cluster to a rigid instance type, flexible VMs allow teams to establish an ordered list of acceptable machine families for master, primary worker, and secondary worker nodes.&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Key features&lt;/span&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Multi-family blending:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Mix nodes across diverse machine types and generations, combining Gen2 families (e.g., N2, N2D) with Gen4 families (e.g., N4, C4) in a single configuration.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Mixed storage support:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Broaden available capacity pools by allowing storage options to dynamically adapt to the underlying host family's supported disk types.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Comprehensive cluster coverage:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Apply flexible rules to primary workers, secondary (preemptible/spot) workers, and master nodes to guarantee cluster provisioning end-to-end.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Ranked configuration: A strategy for success&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;A successful flexible VM implementation relies on intentional ranking. By defining a clear hierarchy of options, Managed Spark clusters automatically attempt provisioning, systematically mitigating stockout risks without requiring manual intervention. To improve the availability of  suitable VMs, we recommend specifying at least two machine families in the highest priority (Rank 0) flexible VM list.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;As an example, for production pipelines standardizing on &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;n2d-standard-16&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; shapes, the following tiering strategy provides robust resilience against capacity constraints:&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;/p&gt;
&lt;div align="left"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;&lt;table style="width: 98.9556%;"&gt;&lt;colgroup&gt;&lt;col style="width: 23.9583%;"/&gt;&lt;col style="width: 41.0156%;"/&gt;&lt;col style="width: 35.026%;"/&gt;&lt;/colgroup&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th scope="col" style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Rank&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Machine family examples&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Storage recommendation&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Rank 0 (Primary)&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;n2d-standard-16, n2-standard-16&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span data-rich-links='{"dde_di":"kix.ggbi3pqb29vd","dde-fdv":"Standard Local SSD or PD","dde-sii":"dropdownItem.w6kf4s2yvnl8","ddefe-ddi":{"cv":{"op":"set","opValue":[{"di-id":"dropdownItem.64xdhfe3knvm","di-v":"Hyperdisk Balanced","di-dv":"Hyperdisk Balanced","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.w6kf4s2yvnl8","di-v":"Standard Local SSD or PD","di-dv":"Standard Local SSD or PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.9kbsce39m365","di-v":"Standard PD","di-dv":"Standard PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}}]}},"ddefe-t":"Storage Recommendation","type":"dropdown"}' style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span data-rich-links='{"dde_di":"kix.ggbi3pqb29vd","dde-fdv":"Standard Local SSD or PD","dde-sii":"dropdownItem.w6kf4s2yvnl8","ddefe-ddi":{"cv":{"op":"set","opValue":[{"di-id":"dropdownItem.64xdhfe3knvm","di-v":"Hyperdisk Balanced","di-dv":"Hyperdisk Balanced","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.w6kf4s2yvnl8","di-v":"Standard Local SSD or PD","di-dv":"Standard Local SSD or PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.9kbsce39m365","di-v":"Standard PD","di-dv":"Standard PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}}]}},"ddefe-t":"Storage Recommendation","type":"dropdown"}' style="vertical-align: baseline;"&gt;Standard Local SSD or PD&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Rank 1&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;n4-standard-16, n4d-standard-16&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span data-rich-links='{"dde_di":"kix.ggbi3pqb29vd","dde-fdv":"Hyperdisk Balanced","dde-sii":"dropdownItem.64xdhfe3knvm","ddefe-ddi":{"cv":{"op":"set","opValue":[{"di-id":"dropdownItem.64xdhfe3knvm","di-v":"Hyperdisk Balanced","di-dv":"Hyperdisk Balanced","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.w6kf4s2yvnl8","di-v":"Standard Local SSD or PD","di-dv":"Standard Local SSD or PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.9kbsce39m365","di-v":"Standard PD","di-dv":"Standard PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}}]}},"ddefe-t":"Storage Recommendation","type":"dropdown"}' style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span data-rich-links='{"dde_di":"kix.ggbi3pqb29vd","dde-fdv":"Hyperdisk Balanced","dde-sii":"dropdownItem.64xdhfe3knvm","ddefe-ddi":{"cv":{"op":"set","opValue":[{"di-id":"dropdownItem.64xdhfe3knvm","di-v":"Hyperdisk Balanced","di-dv":"Hyperdisk Balanced","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.w6kf4s2yvnl8","di-v":"Standard Local SSD or PD","di-dv":"Standard Local SSD or PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.9kbsce39m365","di-v":"Standard PD","di-dv":"Standard PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}}]}},"ddefe-t":"Storage Recommendation","type":"dropdown"}' style="vertical-align: baseline;"&gt;Hyperdisk Balanced&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Rank 2&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;c4-standard-16, c3-standard-22&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span data-rich-links='{"dde_di":"kix.ggbi3pqb29vd","dde-fdv":"Hyperdisk Balanced","dde-sii":"dropdownItem.64xdhfe3knvm","ddefe-ddi":{"cv":{"op":"set","opValue":[{"di-id":"dropdownItem.64xdhfe3knvm","di-v":"Hyperdisk Balanced","di-dv":"Hyperdisk Balanced","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.w6kf4s2yvnl8","di-v":"Standard Local SSD or PD","di-dv":"Standard Local SSD or PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.9kbsce39m365","di-v":"Standard PD","di-dv":"Standard PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}}]}},"ddefe-t":"Storage Recommendation","type":"dropdown"}' style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span data-rich-links='{"dde_di":"kix.ggbi3pqb29vd","dde-fdv":"Hyperdisk Balanced","dde-sii":"dropdownItem.64xdhfe3knvm","ddefe-ddi":{"cv":{"op":"set","opValue":[{"di-id":"dropdownItem.64xdhfe3knvm","di-v":"Hyperdisk Balanced","di-dv":"Hyperdisk Balanced","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.w6kf4s2yvnl8","di-v":"Standard Local SSD or PD","di-dv":"Standard Local SSD or PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.9kbsce39m365","di-v":"Standard PD","di-dv":"Standard PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}}]}},"ddefe-t":"Storage Recommendation","type":"dropdown"}' style="vertical-align: baseline;"&gt;Hyperdisk Balanced&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Rank 3 &lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;e2-standard-16&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span data-rich-links='{"dde_di":"kix.ggbi3pqb29vd","dde-fdv":"Standard PD","dde-sii":"dropdownItem.9kbsce39m365","ddefe-ddi":{"cv":{"op":"set","opValue":[{"di-id":"dropdownItem.64xdhfe3knvm","di-v":"Hyperdisk Balanced","di-dv":"Hyperdisk Balanced","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.w6kf4s2yvnl8","di-v":"Standard Local SSD or PD","di-dv":"Standard Local SSD or PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.9kbsce39m365","di-v":"Standard PD","di-dv":"Standard PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}}]}},"ddefe-t":"Storage Recommendation","type":"dropdown"}' style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span data-rich-links='{"dde_di":"kix.ggbi3pqb29vd","dde-fdv":"Standard PD","dde-sii":"dropdownItem.9kbsce39m365","ddefe-ddi":{"cv":{"op":"set","opValue":[{"di-id":"dropdownItem.64xdhfe3knvm","di-v":"Hyperdisk Balanced","di-dv":"Hyperdisk Balanced","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.w6kf4s2yvnl8","di-v":"Standard Local SSD or PD","di-dv":"Standard Local SSD or PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.9kbsce39m365","di-v":"Standard PD","di-dv":"Standard PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}}]}},"ddefe-t":"Storage Recommendation","type":"dropdown"}' style="vertical-align: baseline;"&gt;Standard PD&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;gcloud dataproc clusters create $CLUSTER_NAME \\\r\n--num-workers=10 \\\r\n--zone=&amp;quot;&amp;quot; \\\r\n--region=us-east1 \\\r\n--worker-instance-selection=\&amp;#x27;{&amp;quot;machineTypes&amp;quot;:[&amp;quot;n2d-standard-16&amp;quot;,&amp;quot;n2-standard-16&amp;quot;],&amp;quot;rank&amp;quot;:0,&amp;quot;diskConfig&amp;quot;:{&amp;quot;bootDiskType&amp;quot;:&amp;quot;pd-standard&amp;quot;,&amp;quot;bootDiskSizeGb&amp;quot;:400}}\&amp;#x27; \\\r\n--worker-instance-selection=\&amp;#x27;{&amp;quot;machineTypes&amp;quot;:[&amp;quot;n4-standard-16&amp;quot;,&amp;quot;n4d-standard-16&amp;quot;],&amp;quot;rank&amp;quot;:1,&amp;quot;diskConfig&amp;quot;:{&amp;quot;bootDiskType&amp;quot;:&amp;quot;hyperdisk-balanced&amp;quot;,&amp;quot;bootDiskSizeGb&amp;quot;:400}}\&amp;#x27; \\\r\n--worker-instance-selection=\&amp;#x27;{&amp;quot;machineTypes&amp;quot;:[&amp;quot;c4-standard-16&amp;quot;,&amp;quot;c3-standard-22&amp;quot;],&amp;quot;rank&amp;quot;:2,&amp;quot;diskConfig&amp;quot;:{&amp;quot;bootDiskType&amp;quot;:&amp;quot;hyperdisk-balanced&amp;quot;,&amp;quot;bootDiskSizeGb&amp;quot;:400}}\&amp;#x27; \\\r\n--worker-instance-selection=\&amp;#x27;{&amp;quot;machineTypes&amp;quot;:[&amp;quot;e2-standard-16&amp;quot;],&amp;quot;rank&amp;quot;:3, &amp;quot;diskConfig&amp;quot;:{&amp;quot;bootDiskType&amp;quot;:&amp;quot;pd-ssd&amp;quot;,&amp;quot;bootDiskSizeGb&amp;quot;:400}}\&amp;#x27; \\\r\n--master-instance-selection=\&amp;#x27;{&amp;quot;machineTypes&amp;quot;:[&amp;quot;n4-standard-16&amp;quot;,&amp;quot;n4d-standard-16&amp;quot;],&amp;quot;rank&amp;quot;:0,&amp;quot;diskConfig&amp;quot;:{&amp;quot;bootDiskType&amp;quot;:&amp;quot;hyperdisk-balanced&amp;quot;,&amp;quot;bootDiskSizeGb&amp;quot;:400}}\&amp;#x27;&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca3430f950&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For pipelines standardizing on legacy &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;n1-standard-16&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; shapes, the following tiering strategy helps transition workloads toward newer, more available architectures while preserving operational stability:&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;/p&gt;
&lt;div align="left"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;&lt;table style="width: 99.7389%;"&gt;&lt;colgroup&gt;&lt;col style="width: 25.625%;"/&gt;&lt;col style="width: 36.875%;"/&gt;&lt;col style="width: 37.3438%;"/&gt;&lt;/colgroup&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th scope="col" style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Rank&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Machine family examples&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Storage recommendation&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Rank 0 (Primary)&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;n1-standard-16&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;n2-standard-16&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span data-rich-links='{"dde_di":"kix.ggbi3pqb29vd","dde-fdv":"","dde-sii":"dropdownItem.w6kf4s2yvnl8","ddefe-ddi":{"cv":{"op":"set","opValue":[{"di-id":"dropdownItem.64xdhfe3knvm","di-v":"Hyperdisk Balanced","di-dv":"Hyperdisk Balanced","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.w6kf4s2yvnl8","di-v":"Standard Local SSD or PD","di-dv":"Standard Local SSD or PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.9kbsce39m365","di-v":"Standard PD","di-dv":"Standard PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}}]}},"ddefe-t":"Storage Recommendation","type":"dropdown"}' style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span data-rich-links='{"dde_di":"kix.ggbi3pqb29vd","dde-fdv":"","dde-sii":"dropdownItem.w6kf4s2yvnl8","ddefe-ddi":{"cv":{"op":"set","opValue":[{"di-id":"dropdownItem.64xdhfe3knvm","di-v":"Hyperdisk Balanced","di-dv":"Hyperdisk Balanced","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.w6kf4s2yvnl8","di-v":"Standard Local SSD or PD","di-dv":"Standard Local SSD or PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.9kbsce39m365","di-v":"Standard PD","di-dv":"Standard PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}}]}},"ddefe-t":"Storage Recommendation","type":"dropdown"}' style="vertical-align: baseline;"&gt;Standard Local SSD or PD&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Rank 1&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;n2d-standard-16&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span data-rich-links='{"dde_di":"kix.ggbi3pqb29vd","dde-fdv":"","dde-sii":"dropdownItem.w6kf4s2yvnl8","ddefe-ddi":{"cv":{"op":"set","opValue":[{"di-id":"dropdownItem.64xdhfe3knvm","di-v":"Hyperdisk Balanced","di-dv":"Hyperdisk Balanced","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.w6kf4s2yvnl8","di-v":"Standard Local SSD or PD","di-dv":"Standard Local SSD or PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.9kbsce39m365","di-v":"Standard PD","di-dv":"Standard PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}}]}},"ddefe-t":"Storage Recommendation","type":"dropdown"}' style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span data-rich-links='{"dde_di":"kix.ggbi3pqb29vd","dde-fdv":"","dde-sii":"dropdownItem.w6kf4s2yvnl8","ddefe-ddi":{"cv":{"op":"set","opValue":[{"di-id":"dropdownItem.64xdhfe3knvm","di-v":"Hyperdisk Balanced","di-dv":"Hyperdisk Balanced","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.w6kf4s2yvnl8","di-v":"Standard Local SSD or PD","di-dv":"Standard Local SSD or PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.9kbsce39m365","di-v":"Standard PD","di-dv":"Standard PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}}]}},"ddefe-t":"Storage Recommendation","type":"dropdown"}' style="vertical-align: baseline;"&gt;Standard Local SSD or PD&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Rank 2&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;n4-standard-16&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;n4d-standard-16&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span data-rich-links='{"dde_di":"kix.ggbi3pqb29vd","dde-fdv":"","dde-sii":"dropdownItem.64xdhfe3knvm","ddefe-ddi":{"cv":{"op":"set","opValue":[{"di-id":"dropdownItem.64xdhfe3knvm","di-v":"Hyperdisk Balanced","di-dv":"Hyperdisk Balanced","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.w6kf4s2yvnl8","di-v":"Standard Local SSD or PD","di-dv":"Standard Local SSD or PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.9kbsce39m365","di-v":"Standard PD","di-dv":"Standard PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}}]}},"ddefe-t":"Storage Recommendation","type":"dropdown"}' style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span data-rich-links='{"dde_di":"kix.ggbi3pqb29vd","dde-fdv":"","dde-sii":"dropdownItem.64xdhfe3knvm","ddefe-ddi":{"cv":{"op":"set","opValue":[{"di-id":"dropdownItem.64xdhfe3knvm","di-v":"Hyperdisk Balanced","di-dv":"Hyperdisk Balanced","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.w6kf4s2yvnl8","di-v":"Standard Local SSD or PD","di-dv":"Standard Local SSD or PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.9kbsce39m365","di-v":"Standard PD","di-dv":"Standard PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}}]}},"ddefe-t":"Storage Recommendation","type":"dropdown"}' style="vertical-align: baseline;"&gt;Hyperdisk Balanced&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Rank 3&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;e2-standard-16&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span data-rich-links='{"dde_di":"kix.ggbi3pqb29vd","dde-fdv":"","dde-sii":"dropdownItem.9kbsce39m365","ddefe-ddi":{"cv":{"op":"set","opValue":[{"di-id":"dropdownItem.64xdhfe3knvm","di-v":"Hyperdisk Balanced","di-dv":"Hyperdisk Balanced","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.w6kf4s2yvnl8","di-v":"Standard Local SSD or PD","di-dv":"Standard Local SSD or PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.9kbsce39m365","di-v":"Standard PD","di-dv":"Standard PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}}]}},"ddefe-t":"Storage Recommendation","type":"dropdown"}' style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span data-rich-links='{"dde_di":"kix.ggbi3pqb29vd","dde-fdv":"","dde-sii":"dropdownItem.9kbsce39m365","ddefe-ddi":{"cv":{"op":"set","opValue":[{"di-id":"dropdownItem.64xdhfe3knvm","di-v":"Hyperdisk Balanced","di-dv":"Hyperdisk Balanced","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.w6kf4s2yvnl8","di-v":"Standard Local SSD or PD","di-dv":"Standard Local SSD or PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}},{"di-id":"dropdownItem.9kbsce39m365","di-v":"Standard PD","di-dv":"Standard PD","di-ts":{"ts_bd":false,"ts_fs":11,"ts_ff":"Arial","ts_it":false,"ts_sc":false,"ts_st":false,"ts_tw":400,"ts_un":false,"ts_va":"nor","ts_bgc2":{"clr_type":0,"hclr_color":null},"ts_fgc2":{"clr_type":0,"hclr_color":null},"ts_bd_i":false,"ts_fs_i":false,"ts_ff_i":false,"ts_it_i":false,"ts_sc_i":false,"ts_st_i":false,"ts_un_i":false,"ts_va_i":false,"ts_bgc2_i":false,"ts_fgc2_i":false},"di-cv":{"dicv_v":0,"dicv_ft":0}}]}},"ddefe-t":"Storage Recommendation","type":"dropdown"}' style="vertical-align: baseline;"&gt;Standard PD&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Leveraging Hyperdisk Balanced&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Unlocking maximum availability with flexible VMs often requires adopting modern storage architectures like &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/compute/docs/disks/hd-types/hyperdisk-balanced"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Hyperdisk Balanced&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. Newer instance families (including N4 and C4) rely on Hyperdisk to deliver predictable performance across variable VM sizes. Starting with default IOPS and throughput settings typically provides a reliable baseline for the majority of distributed Spark jobs.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Trade-offs and key considerations&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;While flexible VMs  dramatically improve cluster provisioning success, aligning them with enterprise requirements involves evaluating several architectural and financial factors:&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;1. Resource quotas&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;It is no longer enough to have one specific machine (e.g., N2) quota. You need to ensure you have sufficient compute and disk quotas allocated for all specific machine types and disks (including Hyperdisk) defined in their flexible VM lists.&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;2. Compute flexible Committed Use Discounts (CUDs)&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Traditional, resource-based CUDs are tied to specific machine families, which limits flexibility. Adopt &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/compute/docs/instances/committed-use-discounts-overview#spend_based"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Compute flexible Committed Use Discounts (CUDs)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to apply savings across multiple VM families and regions.&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;3. Performance Characteristics&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Performance can vary between machine generations, as well as between Local SSD and Hyperdisk. While the Managed Spark team maintains &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/compute/docs/machine-resource"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;internal benchmarks for these comparisons&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, actual outcomes are workload-dependent. Testing your specific Spark jobs across these families is essential for understanding SLA impacts.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Additional recommendations&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In addition to implementing flexible VMs, there are several other key architectural and scheduling strategies to improve resource availability and workload stability:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;AutoZone:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Implement &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/managed-spark/docs/concepts/configuring-clusters/auto-zone"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;AutoZone&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; routing to allow Managed Spark to automatically select the zone best suited to execute the job based on current capacity.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Smaller machine shapes:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Avoid high in demand, large-core shapes. Design workloads and YARN containers to utilize &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/compute/docs/general-purpose-machines"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;smaller machine shapes&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (such as 4, 8, or 16 cores). These smaller shapes are much easier to fulfill from the available GCE on-demand pool.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Autoscaling:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Deploy cluster &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/managed-spark/docs/concepts/configuring-clusters/autoscaling"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;autoscaling&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; with reasonable &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;maxInstances&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; to manage capacity effectively for bursty or unpredictable workloads without relying on rigid, massive upfront provisioning.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/managed-spark/docs/guides/create-partial-cluster"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Partial cluster creation&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Configure a minimum acceptable number of primary workers. This allows clusters to spin up under resource constraints and begin executing, while autoscaling can dynamically add remaining workers as resources become available.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Establish regional fallbacks:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Some regions, such as &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;us-central1,&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; can experience  high demand. Setting up fallbacks to other regions reduces capacity stockout risks.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Keep your Spark jobs running with flexible VMs&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Managing your own Apache Spark infrastructure can be complex, especially when capacity stockouts disrupt your data processing. Utilizing a managed service like Managed Service for Apache Spark provides unique advantages — including built-in platform resilience and access to flexible VMs. By adopting a prioritized fallback strategy with flexible VMs, you can protect your workloads from regional hardware shortages and keep your critical pipelines running.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Ready to improve your Spark workload resilience? Start configuring&lt;/span&gt; &lt;a href="https://docs.cloud.google.com/managed-spark/docs/concepts/configuring-clusters/flexible-vms"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;flexible VMs&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for your Managed Spark clusters today.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Mon, 21 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/maximize-apache-spark-availability-with-flexible-vms/</guid><category>Open Source</category><category>Data Analytics</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Maximizing Apache Spark availability: Mitigating compute stockouts with flexible VMs and other best practices</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/maximize-apache-spark-availability-with-flexible-vms/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Sravani Bobbala</name><title>Software Engineering Manager</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Qiqi Wu</name><title>Product Manager</title><department></department><company></company></author></item><item><title>Accelerating the borderless Lakehouse: Announcing preview of cross-cloud caching</title><link>https://cloud.google.com/blog/products/data-analytics/borderless-lakehouse-cross-cloud-caching-and-connections/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Today, we are excited to announce enhancements to the &lt;/span&gt;&lt;a href="https://cloud.google.com/solutions/data-lakehouse?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;borderless Lakehouse&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;,&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; our answer to how data engineers, data scientists, and increasingly, AI agents, can query governed data directly where it lives.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To reason accurately and automate complex enterprise workflows, agents and data consumers of all types need fast, unified access to an organization's complete data estate, joining customer records, transaction logs, and operational telemetry across clouds. However, modern enterprise data is rarely confined to a single location; data estates often span Amazon S3, Azure Data Lake Storage (ADLS), Google Cloud Storage, operational databases, and SaaS platforms like Salesforce, SAP, and Workday. Historically, uniting these distributed datasets required brittle ETL pipelines, duplicated storage, and prohibitive cross-cloud data transfer costs.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/introducing-the-borderless-lakehouse?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;We introduced the &lt;/span&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;borderless Lakehouse&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; earlier this year to let organizations query and activate data in place across clouds. By adopting the Apache Iceberg REST catalog specification, we federate directly to catalogs such as Databricks Unity Catalog, AWS Glue, and Snowflake Horizon. We also introduced &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Partner Cross-Cloud Interconnect &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;to establish high-bandwidth, private links to other cloud providers, lowering per-gigabyte transfer costs compared to the public internet. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Today, we are taking multi-cloud efficiency a step further by optimizing &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;how much data needs to be transferred across the wire in the first place&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We are excited to announce two new features to help further reduce costs of querying cross-cloud data.  First, the preview of &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/lakehouse/docs/about-borderless-lakehouse#intelligent-caching"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;cross-cloud caching&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for Lakehouse transparently accelerates cross-cloud queries in BigQuery and cuts remote transfer costs by caching frequently accessed data locally in Google Cloud. &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Combining standard Iceberg columnar compression with cross-cloud caching means you often only need to transfer under 5% of the data you process across clouds,&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; which helps lower the Total Cost of Ownership (TCO) to make cross-cloud analytics and AI viable at enterprise scale. In addition, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;BigQuery cross-cloud connections&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; are also available in preview to query non-Iceberg data in other clouds and accelerate workloads.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;How cross-cloud caching works&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Cross-cloud caching meets enterprise performance and security requirements with no knobs to turn or storage to manage to accelerate your queries. Some of the mechanisms used under the hood are:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Sub-file block granularity:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Instead of transferring entire multi-gigabyte files across clouds when a query touches only a few columns, cross-cloud caching operates at the sub-file block level for columnar formats like Apache Parquet. BigQuery caches only the specific column chunks and dictionary pages projected by the query. On a cache miss, BigQuery fetches the needed data from the remote cloud to answer the query, and saves a local copy in the cache for future queries, drastically cutting network transfer and latency on repeated workloads.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Default encryption at rest:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Cached data blocks are encrypted at rest by default using Google-managed encryption keys (GMEK) so that temporary cache storage maintains the same enterprise-grade security posture as native BigQuery storage without extra overhead.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Tenant and regional isolation:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Cache entries are strictly partitioned by project and catalog boundaries to help prevent cross-tenant data exposure. Lakehouse anchors both the local cache and query execution strictly to the configured Google Cloud region (e.g., &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;us-east4&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;) to support compliance with regional data residency requirements when querying remote clouds.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Freshness checks:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Multi-cloud caching often forces a trade-off between speed and freshness. To avoid stale reads, BigQuery fetches remote object metadata before using cached data to ensure the data hasn’t changed and the user still has access. Any upstream table modification prompts BigQuery to fetch new files, while unreferenced cached blocks expire automatically, delivering local query speed with single-source-of-truth accuracy.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For more details on caching mechanics, statistics counters, and regional considerations, see the Lakehouse &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/lakehouse/docs/about-borderless-lakehouse#intelligent-caching"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;intelligent caching documentation&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Cross-cloud caching in action&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;So how does this work in day-to-day operations? Consider an e-commerce team querying a 10 TiB Iceberg sales table (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;aws_lakehouse_catalog.sales.web_sales&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;) in Amazon S3, federated into Lakehouse from Databricks Unity Catalog. During evening promotional drops (8:00–9:00 PM), analysts query historical transactions to identify which storefronts drive peak volume and revenue among high-intent demographics:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;SELECT w.web_name, hd.hd_buy_potential, COUNT(*) AS total_transactions, ROUND(SUM(ws.ws_sales_price), 2) AS total_sales\r\nFROM `aws_lakehouse_catalog.sales.web_sales` ws\r\n-- Joins household_demographics, time_dim (8:00-9:00 PM), and web_site.\r\nGROUP BY w.web_name, hd.hd_buy_potential;&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca36838ed0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Initial execution: Cold columnar retrieval&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;On this initial cold run, the local cache is empty (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;cacheBytesRead: "0"&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;). BigQuery applies partition pruning and column projection to transfer only the required Parquet byte ranges from Amazon S3 over Partner Cross-Cloud Interconnect:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;{\r\n  &amp;quot;totalBytesProcessed&amp;quot;: &amp;quot;230343464114&amp;quot;,\r\n  &amp;quot;objectStorageStats&amp;quot;: [\r\n{&amp;quot;cloudProvider&amp;quot;: &amp;quot;AWS&amp;quot;, \r\n&amp;quot;objectStorageBytesRead&amp;quot;: &amp;quot;25834740486&amp;quot;, \r\n&amp;quot;cacheBytesRead&amp;quot;: &amp;quot;0&amp;quot;}]\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca350f9190&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Logical data processed:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; BigQuery processes &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;214.5 GiB&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; across the 10 TiB dataset.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Standard Iceberg compression efficiency:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; BigQuery reads &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;24.1 GiB&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; from S3 thanks to standard Iceberg columnar compression with Zstandard (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;zstd&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;) — an &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;8.9:1 compression ratio&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;. As these sub-file Parquet blocks arrive in Google Cloud, BigQuery populates the regional cache.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Follow-on exploration: Adding a dimension&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In practice, analysts and agents rarely run the exact same query twice in a row. To drill deeper into fulfillment methods, the analyst modifies the query by adding the shipping method dimension (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;sm.sm_type&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;):&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;SELECT w.web_name, sm.sm_type, hd.hd_buy_potential, COUNT(*) AS total_transactions, ROUND(SUM(ws.ws_sales_price), 2) AS total_sales\r\nFROM `aws_lakehouse_catalog.sales.web_sales` ws\r\nJOIN `aws_lakehouse_catalog.sales.ship_mode` sm ON ws.ws_ship_mode_sk = sm.sm_ship_mode_sk\r\n-- Reuses existing joins on household_demographics, time_dim, and web_site.\r\nGROUP BY w.web_name, sm.sm_type, hd.hd_buy_potential;&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca350f8b90&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Job statistics for this follow-on query show:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;{\r\n  &amp;quot;totalBytesProcessed&amp;quot;: &amp;quot;287928766472&amp;quot;,\r\n  &amp;quot;objectStorageStats&amp;quot;: [\r\n{&amp;quot;cloudProvider&amp;quot;: &amp;quot;AWS&amp;quot;, \r\n&amp;quot;objectStorageBytesRead&amp;quot;: &amp;quot;1426587648&amp;quot;, \r\n&amp;quot;cacheBytesRead&amp;quot;: &amp;quot;25834740486&amp;quot;}]\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca350f9ed0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;94.8% cache hit rate:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; BigQuery serves &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;24.1 GiB&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; of previously queried columns directly from local cache.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Granular remote retrieval:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; BigQuery transfers only &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;1.33 GiB&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; from S3 for the new &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;ws_ship_mode_sk&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; column and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;ship_mode&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; table.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Sub-file flexibility:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Modifying a query reuses cached column chunks and transfers only newly required bytes.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Compounding efficiency at enterprise scale&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When thinking about TCO of cross-cloud queries, the top two factors to account for are:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Compression ratio: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;when using default compression algorithms (Zstandard/zstd) on Iceberg, columnar data is highly compressible. If you assume that your data achieves a compression ratio of 8:1, it means every 1 TiB of logical data processed only requires ~128 GiB of data to move over the network.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Cache hit rates: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;when data is retrieved from cache rather than across the network because it was recently accessed, a network transit is avoided. Assuming 80% of your data results in a cache hit it means for every 100 GiB of physical data accessed only 20 GiB moves over the network.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Taking both factors and assumptions into account, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;for every 1 TiB of data your organization processes, you only need to transfer ~26 GiB across the network (under 3% of total data processed)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;. Combining this reduction with Partner Cross-Cloud Interconnect lowers TCO enough to make cross-cloud analytics and AI cost-effective at petabyte scale.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;BigQuery cross-cloud connections now in preview&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Alongside cross-cloud caching, the preview of &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;BigQuery cross-cloud connections&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; lets organizations connect BigQuery directly to open-format data in Amazon S3 and Azure Storage. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Understanding when to use catalog federation versus cross-cloud connections is straightforward:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;BigQuery cross-cloud connections (for raw files):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; For standalone files (CSV, JSON, ad-hoc Parquet) without an Iceberg catalog, cross-cloud connections let you create BigQuery external tables referencing remote bucket paths directly.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Lakehouse catalog federation (for Iceberg):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; For Iceberg data managed by catalogs like Databricks Unity, AWS Glue, or Snowflake Horizon, Lakehouse automatically synchronizes schemas and table snapshots to simplify the user experience and ensure users are always querying the latest data.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Cross-cloud connections serve as the modern architectural evolution by using standard BigQuery compute workers in Google Cloud regions rather than compute workers in other clouds. This approach helps unlock &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;global region availability&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; and provides &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;full BigQuery feature parity &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;— including with BigQuery AI and Gemini on remote files.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The cross-cloud caching capabilities for Lakehouse applies to data queried from BigQuery cross-cloud connections as well as Lakehouse catalog federation. To learn how to create connections and query external bucket paths, see the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/cross-cloud-connections"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;BigQuery cross-cloud connections setup documentation&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Fri, 18 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/borderless-lakehouse-cross-cloud-caching-and-connections/</guid><category>BigQuery</category><category>Data Analytics</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Accelerating the borderless Lakehouse: Announcing preview of cross-cloud caching</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/borderless-lakehouse-cross-cloud-caching-and-connections/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Will Ochandarena</name><title>Group Product Manager</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Jason Ganetsky</name><title>Staff Software Engineer</title><department></department><company></company></author></item><item><title>The future of orchestration: Pine59’s journey to Airflow 3 on Google Cloud</title><link>https://cloud.google.com/blog/topics/supply-chain-logistics/the-future-of-orchestration-pine59s-journey-to-airflow-3-on-google-cloud/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Operating large data pipelines requires an orchestration layer that scales smoothly as workloads expand. When your pipelines process millions of complex data points every day to feed predictive models, staying up-to-date with your technology stack is a strategic necessity.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://www.pine59.com/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Pine59&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; provides location intelligence data through data pipelines that produce analytical metrics on cadences ranging from hourly to quarterly. One of the company’s most data-intensive metrics, Daily Foot Traffic, computes data for as many as 14 million distinct locations in a single job. To handle this massive volume, Pine59’s system runs entirely on Google Cloud, with the heavy lifting in &lt;/span&gt;&lt;a href="https://cloud.google.com/bigquery"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;BigQuery&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and all of it orchestrated by &lt;/span&gt;&lt;a href="https://cloud.google.com/products/managed-service-for-apache-airflow"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Managed Service for Apache Airflow&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (formerly Cloud Composer) running Apache Airflow 3.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;As the company’s volume of data and number of machine learning workloads scaled up, Pine59 decided to modernize its monorepo, which contains hundreds of directed acyclic graphs (DAGs). Here is a look at how that transition improved Pine59’s MLOps capabilities, developer workflow, and pipeline speed.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Proactive modernization for growth&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Pine59 has long relied on a shared monorepo with code and tooling spanning multiple projects to run its metric production pipelines. As it considered its infrastructure’s future, the company wanted to help its data pipelines run faster and more reliably.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;That’s why it decided to stress-test production workloads against the newly available Managed Airflow (Gen 3) architecture running Airflow 3. The initial results were unambiguous: the Gen 3 environment delivered immediate and significant processing speed, task scheduling, and overall stability improvements. Recognizing the clear potential for performance gains, Pine59 initiated a full transition to the new environment.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/Vertical_Version_yyTMhsG.max-1000x1000.jpg"
        
          alt="1 - Pine59 Google Cloud Architecture Vertical Version"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Orchestrating advanced MLOps&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Pine59’s pipelines don’t just move data; they drive complex ML models, so a core aspect of its migration was optimizing the orchestration of its ML inference workloads.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Previously, Pine59 had used standard Kubernetes operators for these tasks. By moving to Managed Airflow (Gen 3), which features a highly optimized and abstracted infrastructure layer, the company’s engineering team refined its MLOps architecture. They did so by setting up a dedicated &lt;/span&gt;&lt;a href="https://cloud.google.com/kubernetes-engine"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Kubernetes Engine&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (GKE) cluster that was specifically optimized for model inference and integrated it into the Pine59 pipelines.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This clear separation of orchestration and heavy ML execution compute allows data processing and model inference to run efficiently, showcasing Managed Airflow as a resilient, scalable backbone for enterprise MLOps.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Supporting developers with custom extensibility&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Beyond infrastructure improvements, Pine59 was also able to immediately capitalize on Airflow 3’s delivery of a vastly improved developer workflow and user interface. Indeed, managing hundreds of interconnected DAGs requires excellent observability, and Pine59 found Airflow 3’s plugin authoring system remarkably easy to use.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To improve internal developer velocity, the company quickly built a number of custom plugins that it integrated directly into its new Airflow UI:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;BigQuery Auto-linkify:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; A tool that automatically detects internal BigQuery table references within the Airflow Logs and XCom tabs, dynamically generating direct links to BigQuery Studio for faster debugging (available as a &lt;/span&gt;&lt;a href="https://gist.github.com/jan-hajny-unacast/74e1e504e3e3c8765323bd019a87fb30" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;public GitHub gist&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;)&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;DAG Run Configuration Search:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; A custom search form added directly to the DAG overview page. It allows Pine59 engineers to query specific key-value pairs within DAG run payloads (configs) and instantly surface matching runs. This in turn drastically reduces troubleshooting time.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In addition, the team also deployed a compatibility shim layer within its monorepo. This “compat” module dynamically abstracts logic between Airflow versions, streamlining operator migration across versions.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Faster, more reliable pipelines&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For Pine59, migrating to Managed Airflow (Gen 3) with Airflow 3 has yielded clear, quantifiable results.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The most important improvement was the speed of its DAG runs. In the company’s previous setup, tasks often got stuck in a queued state during peak processing surges. With Gen 3, queue latency has dropped dramatically, allowing tasks to start running almost immediately.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Consider the comparison below of total aggregated “queued” &amp;amp; “running” time of more than 300 runs of the same DAG between Managed Airflow (Gen2) with Airflow 2.11 vs. Managed Airflow (Gen3) with Airflow 3.1 below. As we can readily see, the difference in queued time is significant.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/image2_fCfKA2m.max-1000x1000.png"
        
          alt="image2"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Coupled with internal DAG optimizations made during the transition, the performance gains are also highly tangible. For example, the Daily Foot Traffic pipeline previously took nearly 38 minutes to complete. With the new instance, the same workload now takes less than 26 minutes —nearly 32% less processing time.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Today, Pine59 processes all its production workloads on its new Managed Airflow (Gen 3) instance. By moving to this next generation orchestration, the company improved its MLOps capabilities, equipped its developers with better tools, and built a faster, more resilient foundation for future workloads.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If your engineering team spends more time managing infrastructure than delivering value, consider a similar transition and discover how it can help you move from maintaining servers to building the future of your data and AI pipelines today.&lt;/span&gt;&lt;/p&gt;
&lt;hr/&gt;
&lt;p&gt;&lt;sup&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;Special thanks to the following contributor to this post: Alexandre Crespo-Perez&lt;/span&gt;&lt;/sup&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Thu, 17 Sep 2026 17:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/supply-chain-logistics/the-future-of-orchestration-pine59s-journey-to-airflow-3-on-google-cloud/</guid><category>Data Analytics</category><category>Infrastructure Modernization</category><category>Customers</category><category>Supply Chain &amp; Logistics</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>The future of orchestration: Pine59’s journey to Airflow 3 on Google Cloud</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/supply-chain-logistics/the-future-of-orchestration-pine59s-journey-to-airflow-3-on-google-cloud/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Piotr Wieczorek</name><title>Lead Senior Product Manager, Google</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Jan Hajný</name><title>Senior Data Engineer, Pine59</title><department></department><company></company></author></item><item><title>Scaling Telco Autonomy: Leveraging GNNs with Distributed GraphFlow</title><link>https://cloud.google.com/blog/products/databases/run-gnns-at-scale-with-ease-introducing-distributed-graphflow/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The telecommunications industry is currently undergoing a paradigm shift, moving from traditional manual human-driven operations to fully &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/topics/telecommunications/the-autonomous-network-operations-framework-for-csps?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Autonomous Network Operations.&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; Modern networks have grown increasingly complex, heterogeneous, and large-scale, making handcrafted rules-based methods and traditional Machine Learning (ML) approaches alone insufficient to automate network operations. While ML methods can identify subtle patterns and make fine predictions from large amounts of structured data, they lack the ability to understand, reason about the data and the system it represents, and ultimately make the kind of decision a human operator would.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The growth of AI agents and their ability to reason is a promising solution to this shortcoming. However, in the same way a human operator is not capable of directly ingesting the statistical information spread across the billions of data points created in a large network, AI agents also lack the ability to operate at this scale. To address this challenge, telecommunications companies are adopting Graph Neural Networks (GNNs), a modern form of machine learning designed to operate natively on massive volumes of temporal and relational data. By integrating GNNs with AI agents, operators can combine advanced diagnostics such as root cause analysis, capacity planning, traffic forecasting, what-if simulations, and real-time anomaly detection with the reasoning power required to interpret these insights and execute justified actions. This powerful combination enables networks to safely move towards Level 5 Autonomy as &lt;/span&gt;&lt;a href="https://www.tmforum.org/missions/autonomous-networks" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;defined by TM Forum&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, where the system operates autonomously. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In this post, we present the three components (Data, ML, and AI) that will power Google Cloud’s Autonomous Network Operations framework.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_qK2rt5p.max-1000x1000.jpg"
        
          alt="1"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_rvvQ1TV.max-1000x1000.png"
        
          alt="2"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="n2lgl"&gt;Google Autonomous Network Operations framework architecture&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Foundation: Digital Twin on Spanner Graph&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;At the heart of Google Cloud’s Autonomous Network Operations framework is the network digital twin: a highly detailed, virtual replica that continuously mirrors its living telecommunications network in real time. Rather than being a static model, it is represented as a dynamic, temporal network graph that captures the evolving state and relations of its components over time. This architectural approach allows operators to "go back" in time to train and evaluate ML models on historical data, while providing AI agents with the foundational operational knowledge required to achieve Level 5 Autonomy. By simulating the impact of proposed network changes within this digital environment, the Digital Twin establishes a critical layer of trust, enabling AI agents to confidently design future states and automatically resolve network issues.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Google Cloud’s &lt;/span&gt;&lt;a href="https://cloud.google.com/products/spanner/graph?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Spanner Graph&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is well suited to host this digital twin:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Scalability and Availability&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Spanner Graph provides a no compromise foundation for modern applications, offering virtually unlimited scaling that grows as the network grows, along with 0-RPO/0-RTO and five 9s of availability.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Multi-Model Support&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Supports multiple data models (Relational, Graph, Vector, and Full-Text Search) in a single platform allowing developers to build complex compositions such as graph transversals combined with nearest neighbor vector search.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Global Consistency&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Spanner provides a globally consistent view of the network, simplifying system development.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The next figure illustrates a network topology with four node types: routers, interfaces (the physical ports), VPNs (L3VPN service instances), and flows (active traffic sessions). These are connected by directed edge types capturing the full network stack: physical containment (router-interface), physical links (interface-interface), control-plane peering (router-router via OSPF/iBGP), service membership (router-VPN), and traffic anchoring (flow-interface, flow-VPN).&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/3_E5yMTVW.max-1000x1000.jpg"
        
          alt="3"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="n2lgl"&gt;High Level network topology&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;The ML layer: Distributed Graph Flow (DGF)&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To predict how a network will behave and react, the digital twin leverages an ML layer powered by &lt;/span&gt;&lt;a href="https://dgf.readthedocs.io/" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Distributed Graph Flow&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt; (DGF)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;. By training on the vast volumes of structured historical data hosted within Spanner Graph, this layer uncovers critical predictive insights that enable human operators and AI agents to manage networks proactively rather than reactively.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;DGF is a recently open-sourced Python library designed to manage the entire end-to-end lifecycle of GNN modeling. Developed by Google CoreML and Google Research, it brings a decade of internal Google-scale tools and expertise directly to Google Cloud enterprise clients. To accommodate different engineering needs, the library offers high-performance, composable, low-level primitives for advanced teams, alongside a simple API for rapid development that requires no prior GNN expertise.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For instance, training and evaluate a GNN model in GraphFlow with the high level API can be as simple as writing 5 lines of code:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;import dgf\r\n\r\n# Fetch the data from Spanner Graph\r\ngraph, schema = dgf.io.read_spanner_graph(...)\r\n\r\n# Train a node attribute prediction model\r\nmodel = dgf.learning.train_node_model(graph, schema, target_column=&amp;quot;risk_score&amp;quot;)\r\n\r\n# Evaluate the model\r\nmodel.evaluate()\r\n# Make predictions\r\nmodel.predict(graph, seed_node_idxs=[0, 1, 2])\r\n\r\n# Save the model for later\r\nmodel.save(&amp;quot;/tmp/model&amp;quot;)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca2fff3210&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The DGF provides high-level concepts that map directly to Autonomous Network Operations requirements:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/4_87R4Pjc.max-1000x1000.png"
        
          alt="4"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Use cases&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By leveraging DGF and GNNs, telcos can move from reactive maintenance to proactive prevention through several advanced use cases:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Anomaly detection&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: GNNs generate node and edge embeddings that encapsulate historical patterns and current health. Any anomalous embeddings are flagged for review before they lead to service degradation.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Root cause analysis (RCA)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: DGF can output specific subgraphs containing only the relevant network instances related to an incident, such as "Attach Failures" in a specific ZIP code. This allows troubleshooting agents to perform high-speed analysis without scanning the entire global network.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Predictive maintenance&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The system can predict the likelihood of device failures or edge breaks, such as "handover failures" for fast-moving equipment, enabling proactive load balancing or rerouting. Furthermore, by combining agents, remedial actions can be automated by adopting a ‘human-on-the-loop’/’human-in-the-loop’.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;What-if analysis&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: GNNs enable Telcos to simulate scenarios like fiber cuts,  or traffic surges or device configuration changes. By modeling topological dependencies, GNNs can predict how these local changes propagate across the entire network, allowing engineers to test resilience and evaluate mitigation strategies in a risk-free digital environment.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Scenario: Root cause analysis with GNNs and DGF&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Once you have created a digital twin (&lt;/span&gt;&lt;a href="https://github.com/GoogleCloudPlatform/cloud-spanner-samples/tree/main/telco-and-csp/ano-gnn" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;example code&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;), a straight-forward 5-step process can be used to implement Root Cause Analysis(RCA) detection using GNNs and DGF. &lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Connect to the Digital Twin&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Use the DGF Spanner Graph connector (&lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;dgf.io.read_spanner_graph&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;) to load the network topology directly from Spanner Graph's Digital Twin into the DGF environment.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Train a Supervised Node (or Edge) Prediction model&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Depending on the training data and objective, you will train a supervised node prediction model to predict a target node feature or an edge prediction model to predict an edge between the root cause entity node and the affected entity node. For the given sample data you will use the high-level &lt;/span&gt;&lt;code style="font-style: italic; vertical-align: baseline;"&gt;dgf.learning.train_node_model&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; API to train a supervised node prediction model.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Use the node prediction model to predict root cause node&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The node prediction model can be directly used to predict the impact score on the node with the anomaly. Entity nodes affected by the anomaly with highest predicted impact score will be the top candidates for root cause.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Deploy to &lt;/strong&gt;&lt;a href="https://cloud.google.com/blog/products/ai-machine-learning/introducing-gemini-enterprise-agent-platform"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise Agent Platform&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt; (formerly Vertex AI)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Export the model and host it on a Gemini Enterprise endpoint to enable scalable, low-latency predictions.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Real-time Inference&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Make prediction calls to the inference endpoint with the anomaly date as input. The endpoint will return the predicted root cause Entity nodes. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Get started today&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The integration of GNN using Distributed Graph Flow into network operations is more than just a technical upgrade; it is a critical evolution for the telco industry. By moving towards a GNN-powered autonomous framework, operators can significantly shorten outage times, optimize capacity in real-time, and ultimately deliver a superior customer experience through improved operational efficiency.&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To start building your own intelligent network applications, check out the &lt;/span&gt;&lt;a href="https://github.com/google/distributed_graph_flow" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Distributed GraphFlow (DGF)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; library, which provides the essential primitives for scalable GNN training and inference. For a hands-on experience, follow our step-by-step &lt;/span&gt;&lt;a href="https://github.com/GoogleCloudPlatform/cloud-spanner-samples/tree/main/telco-and-csp/ano-gnn" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;code sample&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. You can also explore our recent award-&lt;/span&gt;&lt;a href="https://www.tmforum.org/catalysts/awards?moonshotsOnly=false" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;winning Moonshot project&lt;/span&gt;&lt;/a&gt; &lt;span style="vertical-align: baseline;"&gt;on &lt;/span&gt;&lt;a href="https://www.tmforum.org/catalysts/projects/C26.0.965/businessaware-gnnhealing-networks" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Business-aware GNN-healing networks&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and dive deeper into our approach on self-optimizing autonomous networks by &lt;/span&gt;&lt;a href="https://services.google.com/fh/files/misc/self_optimizing_autonomous_networks_white_paper.pdf" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;reviewing this whitepaper.&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Tue, 15 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/databases/run-gnns-at-scale-with-ease-introducing-distributed-graphflow/</guid><category>BigQuery</category><category>Data Analytics</category><category>Databases</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Scaling Telco Autonomy: Leveraging GNNs with Distributed GraphFlow</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/databases/run-gnns-at-scale-with-ease-introducing-distributed-graphflow/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Brian Naughton</name><title>Senior Principal Architect, Telecommunications</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Mathieu Guillame-Bert</name><title>Software Engineer</title><department></department><company></company></author></item><item><title>Agent-ready analytics: Unlocking insights with BigQuery augmented analytics</title><link>https://cloud.google.com/blog/products/data-analytics/bigquery-augmented-analytics-tvfs/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;BigQuery now features a suite of augmented analytics Table-Valued Functions (TVFs) designed to automate complex data analysis at scale. Augmented analytics combines AI, ML and statistical methods to automate insight discovery and pattern explanation. These functions allow you to diagnose why metrics changed, uncover underlying trends and relationships across the data, and even isolate the true impact of business decisions. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;These TVFs run directly where your data lives, which helps speed up analysis and reduces the need to export data into external tools. In addition, since these functions are compact and yield structured SQL outputs, they can easily be integrated as skills for AI agents, which easily enables automated, conversational data investigation workflows. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We are introducing six new augmented analytics functions in BigQuery, each created to address a specific analytical challenge:&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;/p&gt;
&lt;div align="left"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;&lt;table&gt;&lt;colgroup&gt;&lt;col/&gt;&lt;col/&gt;&lt;col/&gt;&lt;/colgroup&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;TVF Function&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;What It Helps You Find&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Real World Question It Answers&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;AI.KEY_DRIVERS&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Identifies the top drivers behind an increase or drop in a metric between two time periods or groups. &lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Why did revenue spike this quarter compared to last quarter?&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;AI.CAUSAL_EFFECT&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Quantifies the impact of an action or event by comparing the observed results to an expected baseline.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;How much of the revenue lift came from our pricing update rather than organic growth?&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;ML.CORRELATION&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Evaluates the direction and strength of the relationship between pairs of numeric metrics. &lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Does increased user session duration correlate with higher lifetime customer value?&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;ML.DETECT_CHANGE_POINTS&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Identifies specific dates or intervals where a metric experiences a shift compared to surrounding patterns.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;During which time periods did our platform latency experience persistent, structural shifts?&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;ML.TREND&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Separates the underlying growth or decline from short-term fluctuations or noise. &lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;What are the underlying trends of my revenue over the past year, abstracting away the outlying spikes and drops?&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;ML.SEASONALITY&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Discovers predicable repeated cycles across hours, days, weeks, months or quarters.  &lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Which days of the week consistently experience the highest server load?&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;As we show in the next section, these functions can be easily chained together. The output of one function, such as a detected time window, can directly parameterize the next analytical step.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;A step-by-step example of chaining insights&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Consider a case where there is a shift in a metric, and you need to diagnose the underlying cause and measure the business lift. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To diagnose, we can chain ML.DETECT_CHANGE_POINTS, AI.KEY_DRIVERS and AI.CAUSAL_EFFECT using the Austin Bikeshare sample dataset (bigquery-public-data.austin_bikeshare.bikeshare_trips). This dataset contains historical trip volume and demographic data for the city’s bikesharing program. &lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Step 1: Detect change points&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;ML.DETECT_CHANGE_POINTS automatically identifies statistically significant structural shifts or level changes in your time-series data. While this example demonstrates the analysis  in a single aggregate metric, this function is highly scalable and is capable of running across millions of individual time series. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To find these shifts,  we run the following query across the daily baseline:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;quot;WITH daily_trips AS (\r\n SELECT\r\n   TIMESTAMP_TRUNC(start_time, DAY) AS trip_day,\r\n   COUNT(*) AS total_trips\r\n FROM `bigquery-public-data.austin_bikeshare.bikeshare_trips`\r\n GROUP BY 1\r\n)\r\nSELECT\r\n begin_timestamp,\r\n end_timestamp,\r\n metrics.avg AS avg_daily_trips,\r\n metrics.min AS min_daily_trips,\r\n metrics.max AS max_daily_trips,\r\n metrics.count AS duration_days\r\nFROM ML.DETECT_CHANGE_POINTS(\r\n (SELECT * FROM daily_trips),\r\n data_col =&amp;gt; &amp;#x27;total_trips&amp;#x27;,\r\n timestamp_col =&amp;gt; &amp;#x27;trip_day&amp;#x27;\r\n);&amp;quot;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca35890650&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The output identifies the exact time intervals where the baselines have shifted over the company’s history:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/image5_syrvVHj.png"
        
          alt="1"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If we look at the raw daily session counts, this aligns with shifts over time. We highlight the two change points with the longest durations below:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_lQiu1DF.max-1000x1000.png"
        
          alt="2"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The shift in February 2018 aligns with the day the Austin City Council passed the “Dockless Mobility Pilot Program”, to transform the transit ecosystem, integrating shared electric scooters and bikes into the public. &lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Step 2: Key drivers attribution&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We can input the February 2018 slice found directly to AI.KEY_DRIVERS to determine the particular factors (i.e. bike_type, subscriber_type, etc) driving the surge. AI.KEY_DRIVERS can scan through millions of rows of multi-dimensional data in seconds. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We define the &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;interest group &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;as the slice of time after the shift occurs and compare it against the time period before the shift as the &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;reference group&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;quot;WITH daily_segments AS (\r\n  SELECT \r\n    start_station_name,\r\n    end_station_name,\r\n    subscriber_type,\r\n    bike_type,\r\n    1 AS trip_count,\r\n    -- We use the precise breakpoint identified by Change Points\r\n    IF(EXTRACT(DATE FROM start_time) &amp;gt;= &amp;#x27;2018-02-11&amp;#x27;, TRUE, FALSE) AS after_shift\r\n  FROM `bigquery-public-data.austin_bikeshare.bikeshare_trips`\r\n  -- Equidistant ~30 day window around the event\r\n  WHERE start_time BETWEEN &amp;#x27;2018-01-12&amp;#x27; AND &amp;#x27;2018-03-13&amp;#x27;\r\n)\r\nSELECT \r\n  drivers,\r\n  metric_interest,\r\n  metric_reference,\r\n  difference,\r\n  relative_difference,\r\n  unexpected_difference,\r\n  contribution\r\nFROM AI.KEY_DRIVERS(\r\n  (SELECT * FROM daily_segments),\r\n  metric_col =&amp;gt; &amp;#x27;trip_count&amp;#x27;,\r\n  interest_label_col =&amp;gt; &amp;#x27;after_shift&amp;#x27;,\r\n  dimension_cols =&amp;gt; [&amp;#x27;start_station_name&amp;#x27;, \r\n                     &amp;#x27;end_station_name&amp;#x27;, \r\n                     &amp;#x27;subscriber_type&amp;#x27;, \r\n                     &amp;#x27;bike_type&amp;#x27;],\r\n  top_k =&amp;gt; 10\r\n);&amp;quot;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca3470e310&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;AI.KEY_DRIVERS isolates the top contributing dimension values.  Each row contains a &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;segment&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, which represents a slice of data identified by a specific combination of dimension values (e.g., subscriber_type = 'UT Student' and bike_type = 'classic'). &lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/3_BD47nN6.max-1000x1000.png"
        
          alt="3"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The analysis reveals that the overall trip count increased +374.7% (+40,159 trips) between the reference and interest time windows. The massive growth was overwhelmingly concentrated in U.T. Student Memberships (+7,167.1%) and trips ending at the 21st &amp;amp; Speedway @PCL station (+20,739.1%).&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This aligns with Austin Bikeshare’s response to the Dockless Mobility Pilot Program. In early February, the bikeshare program launched a large promotional partnership with the University of Texas that offered free annual memberships to all UT students.&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Step 3: Causal effect&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;While we know what drove the surge and when it started, we need to isolate the true return on investment over organic expectations. AI.CAUSAL_EFFECT can construct an &lt;/span&gt;&lt;a href="https://arxiv.org/pdf/2510.24452" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;ARIMA_PLUS&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; counterfactual to measure what the volume would have been had the program never launched. &lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;quot;WITH daily_trips AS (\r\n  SELECT \r\n    TIMESTAMP_TRUNC(start_time, DAY) AS trip_day, \r\n    COUNT(*) AS total_trips\r\n  FROM `bigquery-public-data.austin_bikeshare.bikeshare_trips`\r\n  -- Training on the 6-month baseline leading up to the intervention\r\n  WHERE start_time BETWEEN &amp;#x27;2017-08-11&amp;#x27; AND &amp;#x27;2018-04-11&amp;#x27;\r\n  GROUP BY 1\r\n)\r\nSELECT \r\n  *\r\nFROM AI.CAUSAL_EFFECT(\r\n  (SELECT * FROM daily_trips),\r\n  data_col =&amp;gt; &amp;#x27;total_trips&amp;#x27;,\r\n  timestamp_col =&amp;gt; &amp;#x27;trip_day&amp;#x27;,\r\n  -- We inject the breakpoint found in Step 1 as our intervention\r\n  intervention_timestamp =&amp;gt; &amp;#x27;2018-02-11 00:00:00&amp;#x27;,\r\n  output_time_series =&amp;gt; TRUE\r\n);&amp;quot;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca3470ee90&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If we graph the predicted and actual trips per day, we can see the surge compared to the counterfactual.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/4_X3SlXmT.max-1000x1000.png"
        
          alt="4"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If we set the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;output_time_series =&amp;gt; &lt;/code&gt;&lt;code style="vertical-align: baseline;"&gt;FALSE&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, we can see a summary of the lift&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/5_4o5pYGA.max-1000x1000.jpg"
        
          alt="5"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;AI.CAUSAL_EFFECT reveals that the program caused a +358% volume surge above organic baseline projections, resulting in an estimated 89,775 incremental trips (with 99.9% probability of causal effect).&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Connecting augmented analytics to Conversational Analytics&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Conversational Analytics lets you chat with agents about your data using natural language. All new BigQuery augmented analytical functions are now available in &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/conversational-analytics"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Conversational Analytics&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. Since these TVFs can execute complex analytics at BigQuery-scale in seconds, Conversational Analytics can orchestrate multi-step investigative workflows based on a given prompt. Below we show two examples:&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Example 1: Chicago taxi trips&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Here is an example using the &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Chicago Taxi Trips &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;(`bigquery-public-data.chicago_taxi_trips.taxi_trips`).&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Prompt:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; What metric has the strongest correlation with drivers getting tipped? Then run an attribution analysis to tell me which categorical dimensions (like location and payment type) most disproportionately drive that specific metric.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/Example_1.max-1000x1000.png"
        
          alt="6"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The results here used ML.CORRELATION in combination with AI.KEY_DRIVERS.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Credit card payments serve as the primary positive driver of trip distance, adding +1.65M due to longer travel routes and automated digital tip tracking. Trips originating from O'Hare International Airport (Community Area 76) represent another major positive factor, contributing an additional +1.10M miles among tipped credit card rides. In contrast, cash transactions act as a significant negative driver (-652.96K miles), reflecting that cash is predominantly used for shorter journeys rather than extended airport travel.&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Example 2: Iowa liquor dataset&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Here is an example using the Iowa liquor dataset (`bigquery-public-data.iowa_liquor_sales.sales`) that uses both ML.TREND in combination with ML.SEASONALITY.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Prompt:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Find the historical trend for bottles sold. Then, describe the yearly seasonality patterns.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/Example_2.max-1000x1000.png"
        
          alt="7"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The results show that liquor sales in Iowa show persistent long-term growth, rising from 1.3–1.5 million bottles in 2012 before stabilizing around 2.6 million in recent years. There are strong seasonal cycles, particularly during October and December as well as May and June. There is a drop in sales around January and February.  &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The skills for these TVFs are now available at the &lt;/span&gt;&lt;a href="https://github.com/google/skills" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Skills Github&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; repository. The BQ AI/ML skills can be found &lt;/span&gt;&lt;a href="https://github.com/google/skills/tree/main/skills/cloud/bigquery-ai-ml" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Take the next step&lt;/span&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Documentation:&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;ul&gt;
&lt;li aria-level="2" style="list-style-type: circle; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/reference/standard-sql/bigqueryml-syntax-ai-key-drivers"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;AI.KEY_DRIVERS &lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="2" style="list-style-type: circle; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/reference/standard-sql/bigqueryml-syntax-causal-effect"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;AI.CAUSAL_EFFECT&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="2" style="list-style-type: circle; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/reference/standard-sql/bigqueryml-syntax-correlation"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;ML.CORRELATION&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="2" style="list-style-type: circle; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/reference/standard-sql/bigqueryml-syntax-seasonality"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;ML.SEASONALITY&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="2" style="list-style-type: circle; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/reference/standard-sql/bigqueryml-syntax-trend"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;ML.TREND&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="2" style="list-style-type: circle; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/reference/standard-sql/bigqueryml-syntax-detect-change-points"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;ML.DETECT_CHANGE_POINTS&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/conversational-analytics#bigquery-ml-support"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;BigQuery AI/ML support in Conversational Analytics&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;hr/&gt;
&lt;p&gt;&lt;sub&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;We would like to extend our sincere thanks to Katelin Amann, Shirley Fu, Chaoyi Shen, Haiyang Qi, Zheng Zhang, Xi Cheng and the wider engineering team for their feedback and contributions of this work.&lt;/span&gt;&lt;/em&gt;&lt;/sub&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Mon, 14 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/bigquery-augmented-analytics-tvfs/</guid><category>BigQuery</category><category>Data Analytics</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Agent-ready analytics: Unlocking insights with BigQuery augmented analytics</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/bigquery-augmented-analytics-tvfs/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Jenny Ortiz</name><title>Senior Software Engineer, Google</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Haoming Chen</name><title>Staff Software Engineer, Google</title><department></department><company></company></author></item><item><title>Announcing Pause/Resume and NVIDIA RTX PRO 6000 Blackwell GPU support in Dataflow</title><link>https://cloud.google.com/blog/products/data-analytics/new-dataflow-features-to-enable-large-scale-ai-workloads/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Overview&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;As enterprises scale their AI and agentic workflows, they require serverless platforms that make data preparation for model training, evaluation, and inference effortless and efficient. &lt;/span&gt;&lt;a href="https://cloud.google.com/products/dataflow"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Dataflow&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is a critical component of Google Cloud’s AI stack. It enables our customers to create batch and streaming pipelines that support a variety of analytics and AI use cases. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Today, we’re delivering significant enhancements to Dataflow that directly address your top challenges: maximizing compute efficiency for long-running batch jobs and delivering extra inference power for your most demanding AI workloads. We’re thrilled to announce the general availability of Pause/Resume for Dataflow batch jobs as well as support for G4 VMs powered by NVIDIA RTX PRO 6000 Blackwell Server Edition GPUs. With these features, you can accelerate your AI development lifecycle and optimize your costs.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Recover wasted compute and increase developer productivity with Pause/Resume for Dataflow batch jobs&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Dataflow customers frequently run large batch workloads that sometimes run for a few days. When these jobs fail, Dataflow users currently cannot access the data that was already processed before the job failure. Instead, they have to retry the entire job, leading to wasted compute resources and decreased engineering productivity.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In addition to addressing failures from large jobs, Dataflow customers with AI workloads sometimes want to increase the utilization of accelerated compute resources like GPUs and TPUs by dynamically re-allocating them from already running, lower priority Dataflow batch jobs to higher priority workloads like feature engineering and AI inference. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To better support these use cases, we are announcing the GA launch of Pause/Resume for Dataflow batch jobs. Powered by internal Google innovation, this feature enables Dataflow customers to resume their failed long running jobs instead of starting from scratch. It also allows customers to pause and resume their Dataflow batch jobs based on their respective business requirements.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For more details, see &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/dataflow/docs/guides/pause-job#console"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;manually pause a Dataflow job&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/image1_k7tia8a.max-1000x1000.png"
        
          alt="image1"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Accelerate AI inference workloads with NVIDIA RTX PRO 6000 GPUs&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;While Dataflow already supports a wide variety of GPUs and TPUs for accelerating AI inference workloads, we’re taking things a step further by announcing support for G4 VMs powered by &lt;/span&gt;&lt;a href="https://cloud.google.com/dataflow/docs/gpu/gpu-support#availability"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;NVIDIA RTX PRO 6000 Blackwell GPUs&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The NVIDIA RTX PRO 6000 Blackwell GPU delivers significant performance gains compared to the NVIDIA L4 GPU, bringing 96GB vGPU memory and 1.6 TB/s of bandwidth. This means that you can perform AI inference right within your Dataflow job using up to 70B+ parameter models. You can do this while continuing to take advantage of native Dataflow ML capabilities like &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/dataflow/docs/machine-learning/runinference-best-practices"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;RunInference&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/dataflow/docs/guides/right-fitting"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;right fitting&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://cloud.google.com/dataflow/docs/guides/tune-horizontal-autoscaling#parallelism-hint"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GPU-enabled autoscaling&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; which make it easy for you to onboard and scale your AI inference jobs without having to manage underlying infrastructure or manually deal with hard problems like tuning and autoscaling. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Take the next step&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Together, Pause/Resume and RTX PRO 6000 Blackwell GPUs help you optimize your batch job costs while running demanding AI workloads. We’re incredibly excited about Dataflow’s capabilities and the possibilities they unlock for our customers. &lt;/span&gt;&lt;a href="https://cloud.google.com/dataflow/docs/machine-learning"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Get started with Dataflow&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; today and use these features to solve your hardest AI challenges. We cannot wait to see what you build.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Mon, 14 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/new-dataflow-features-to-enable-large-scale-ai-workloads/</guid><category>Data Analytics</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Announcing Pause/Resume and NVIDIA RTX PRO 6000 Blackwell GPU support in Dataflow</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/new-dataflow-features-to-enable-large-scale-ai-workloads/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Efesa Origbo</name><title>Product Manager, Google Cloud</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Danny McCormick</name><title>Software Engineer, Google Cloud</title><department></department><company></company></author></item><item><title>What’s new with Google Data Cloud</title><link>https://cloud.google.com/blog/products/data-analytics/whats-new-with-google-data-cloud/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;September 7 - September 10&lt;/h3&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Pub/Sub SMTs can now AI Inference your Gemini Enterprise Agent Platform models!&lt;br/&gt;&lt;/strong&gt;&lt;a href="https://docs.cloud.google.com/pubsub/docs/smts/ai-inference-smt" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Pub/Sub AI Inference SMTs&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;allow you to apply inference on an incoming stream of events using models hosted in Gemini Enterprise Agent Platform. The model’s prediction is appended to your event, making it available for downstream processing in your data warehouse (like BigQuery) or operational database (like BigTable). This feature, now generally available, can dramatically simplify or enhance anomaly detection systems you are operating. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;PostgreSQL Source Connector is now generally available in Managed Service for Apache Kafka!&lt;br/&gt;&lt;/strong&gt;Managed Service for Apache Kafka’s PostgreSQL connector allows customers to capture changes from their PostgreSQL database and ingest them into their Kafka infrastructure with low latency. This source connector is compatible with &lt;a href="https://docs.cloud.google.com/managed-service-for-apache-kafka/docs/connect-cluster/create-cloud-sql-postgres-source-connector" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud SQL for Postgres&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/managed-service-for-apache-kafka/docs/connect-cluster/create-generic-postgres-source-connector" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;AlloyDB, and self-managed PostgreSQL&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; databases. Try this along with our entire portfolio of &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/managed-service-for-apache-kafka/docs/kafka-connect-overview" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;managed connectors, including MirrorMaker 2.0, BigQuery, Cloud Storage, and Pub/Sub&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;! E-mail &lt;/span&gt;&lt;a href="mailto:kafka-hotline@google.com" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;kafka-hotline@google.com&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; if you have questions or feedback!&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Pause-on-failure for Dataflow batch jobs is GA&lt;br/&gt;&lt;/strong&gt;&lt;a href="https://docs.cloud.google.com/dataflow/docs/guides/pause-job" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Dataflow pause-on-failure&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; enables you to preserve the state of a batch Dataflow job before it fails. By pausing your Dataflow job, you can address issues that are external to the pipeline and resume processing without losing completed work. This helps you better manage resource costs and improve job reliability when you face temporary outages or capacity constraints.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;The &lt;/strong&gt;&lt;strong style="vertical-align: baseline;"&gt;insertAll&lt;/strong&gt;&lt;strong style="vertical-align: baseline;"&gt; API is now the BigQuery Storage Write API (REST)&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;The legacy insertAll streaming API is now rebranded as the BigQuery Storage Write API (REST). By dropping the "legacy" label, developers can confidently build long-term HTTP-based streaming workflows. This stateless JSON-over-HTTPS endpoint offers a lightweight alternative to heavy gRPC libraries—ideal for serverless web apps, IoT telemetry, and AI logging. The transition is seamless for existing users, requiring zero code changes and offering 100% backward compatibility. However, the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/write-api-grpc" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Storage Write API (gRPC)&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; version remains the recommended standard for high-throughput, continuous pipelines.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;August 31 - September 4&lt;/h3&gt;
&lt;ul&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Stateful processing is available in BigQuery continuous queries in Preview&lt;/strong&gt;&lt;br/&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/continuous-queries-introduction#supported_stateful_operations"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Stateful operations&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; significantly expand what’s possible with BigQuery continuous queries. This feature allows users to leverage functions like &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;JOIN&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;s, aggregations, and windowing functions directly in their streaming queries. Now you can calculate metrics over time (for example, a 30-minute average) to power your downstream applications and AI agents with much richer, real-time signals.&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Try out our feature &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/continuous-query-joins"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and share your feedback with bq-continuous-queries-feedback@google.com!&lt;/span&gt;&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Synthetic data generator tool is available for Managed Service for Kafka&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;You’ve launched your first Kafka cluster. Now what? The next thing to do is to produce some data to the cluster, but that involves modifying a client application somewhere or spinning up a virtual machine. The synthetic data generator tool, now generally available, can start sending mock data to your cluster in 3 clicks, and will get data streaming into your cluster in less than two minutes. The perfect utility for those moments you just want to test your cluster and new features. Try &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/managed-service-for-apache-kafka/docs/quickstart-synthetic-data"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;our quickstart&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; today!&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Dataflow pipeline updates are faster &amp;amp; more flexible&lt;br/&gt;&lt;/strong&gt;&lt;a href="https://docs.cloud.google.com/dataflow/docs/guides/upgrade-guide"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Dataflow pipeline updates&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;can now stop-and-replace pipelines, a major addition to the existing in-place-update feature. The new parallel pipeline option accelerates the migration between the old &amp;amp; new pipeline, resulting in reduced disruption to your business. You can also set a timeout on drains that prevents runaway costs for your pipeliness in the event of stuck processing. This feature is generally available. Try it &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/dataflow/docs/guides/updating-a-pipeline"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;!&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;July 6 - July 10&lt;/span&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;New Lakehouse managed tables now in preview &lt;br/&gt;&lt;/strong&gt;&lt;a href="https://docs.cloud.google.com/lakehouse/docs/manage-tables" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Lakehouse tables for Apache Iceberg&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; are now in preview and available &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;in the console&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;. By using Google-managed Apache Iceberg tables in Lakehouse, you can eliminate the overhead of maintaining duplicate data pipelines and complex synchronization logic between BigQuery and open-source engine&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;s&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;. This unified table format delivers native, multi-engine read and write interoperability, allowing you to run concurrent DML/DDL operations across diverse analytics tools on a single, shared storage layer.  Built-in automated table management handles painful background optimization tasks like compaction and partition tuning, freeing up your team to focus on building rather than managing storage maintenance.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;June 1 - June 5&lt;/span&gt;&lt;/span&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Beyond the Query: Powering AI Agents with Bigtable, Firestore &amp;amp; Memorystore &lt;br/&gt;&lt;/strong&gt;&lt;span style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;Discover the latest advancements in Google Cloud's NoSQL Database portfolio, including Bigtable, Firestore, and Memorystore. This series is designed for a broad audience: whether you are exploring these databases for the first time or are an existing user looking to leverage the new capabilities announced at Next '26. &lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;a href="https://rsvp.withgoogle.com/events/beyond-the-query-powering-ai-agents-with-bigtable-firestore-memorystore" rel="noopener" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Register here to secure your spot!&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Cloud Engineer's AI Toolkit Workshops: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Solve data-driven challenges with &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;BigQuery, AlloyDB&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Gemini&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; and more. &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Hosted by Google Cloud Labs, this highly technical event is built specifically for Platform Engineers, SREs, and cloud infrastructure teams ready to bridge the gap between AI prototypes and production-grade deployments. Look out for more locations coming soon&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Toronto&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; - June 25 (Data Cloud) | &lt;/span&gt;&lt;a href="https://rsvp.withgoogle.com/events/google-cloud-labs-data-cloud-toronto" rel="noopener" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;RSVP Here&lt;/span&gt;&lt;/a&gt;&lt;br/&gt;&lt;strong style="vertical-align: baseline;"&gt;Chicago&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; - June 30 (Data Cloud) | &lt;/span&gt;&lt;a href="https://rsvp.withgoogle.com/events/google-cloud-labs-data-cloud-chicago" rel="noopener" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;RSVP Here&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;strong style="vertical-align: baseline;"&gt;Start a 10-day &lt;/strong&gt;&lt;a href="https://cloud.google.com/bigtable"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Bigtable&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt; free trial with a 1 node SSD cluster and up to 500GB of storage capacity. &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;W&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;ith no credit card required to start, you can easily ingest workloads and manage workloads that require low-latency, high-throughput, and predictable access. &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Plus, new Google Cloud customers get &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/sql/docs/mysql/create-free-trial-instance"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;$300 in free credits&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; on signup.&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;May 11 - May 15&lt;/span&gt;&lt;/span&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;strong style="vertical-align: baseline;"&gt;Managed Service for Apache Airflow&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; has launched a wave of new features, including the general availability of Airflow 3.1, AI-powered agentic troubleshooting, a new managed Airflow MCP Server for custom agent integration, and declarative YAML-based orchestration pipelines—discover all the details in the&lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/managed-apache-airflow-scaling-data-and-ai-workloads"&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;full blog post&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;April 20 - April 24&lt;/span&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;strong style="vertical-align: baseline;"&gt;Google-built ODBC Driver for BigQuery is now available in Preview&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;We are excited to announce the launch of the new, Google-built ODBC driver for BigQuery. This new open-source driver provides a direct, high-performance connection for applications to BigQuery and is developed entirely in-house by Google. &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/odbc-for-bigquery"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Download a new driver and connect your application to BigQuery&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;April 13 - April 17&lt;/span&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;We announced &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/looker-studio-is-data-studio"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;we are reintroducing Data Studio&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to play a significant role in the AI era, expanding from data visualizations and reports to host BigQuery conversational agents and data apps built in Colab notebooks.&lt;/span&gt;&lt;/li&gt;
&lt;li&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;We announced &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/introducing-bigquery-graph"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;BigQuery Graph is now available in preview&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, offering an easy-to-use, highly scalable graph analytics solution, empowering data professionals to model, analyze and visualize massive-scale relationships in an entirely new way. &lt;/span&gt;&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;April 6 - April 10&lt;/span&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;We introduced &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/business-intelligence/looker-embedded-adds-conversational-analytics"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Conversational Analytics for Looker Embedded environments&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, enabling users to add natural language experiences to their own custom data-driven applications, powered by Gemini. &lt;/span&gt;&lt;/li&gt;
&lt;li&gt;&lt;span style="vertical-align: baseline;"&gt;We expanded Looker’s capabilities for faster ad-hoc analysis, with the &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/business-intelligence/looker-self-service-explores"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;introduction of self-service Explores&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, enabling you to bring your own data to Looker’s semantic layer and gain instant access to insights in a governed data environment.&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;March 23 - March 27&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;We showed you how you can &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/databases/cloudsql-read-pools-support-autoscaling"&gt;&lt;span style="vertical-align: baseline;"&gt;scale your reads with Cloud SQL autoscaling read pools.&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; This feature allows you to provision multiple read replicas that are accessible via a single read endpoint and to dynamically adjust your read capability based on real-time application needs. &lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;Our customers are leveraging the full power of Conversational Analytics and Looker to drive major business and technical breakthroughs in the AI era. Companies like &lt;/span&gt;&lt;a href="https://cloud.google.com/customers/telenor-looker"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Telenor&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/customers/petcircle-looker"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Pet Circle&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/customers/fluent-commerce"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Fluent Commerce&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/customers/lighthouse"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Lighthouse Intelligence&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/customers/wego"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Wego&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and &lt;/span&gt;&lt;a href="https://cloud.google.com/customers/roller"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;ROLLER&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; are turning data into insights and actions, grounded by Looker’s semantic layer.&lt;/span&gt;&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;March 16 - March 20&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;We introduced &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/gemini-supercharges-the-bigquery-studio-assistant"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;an enhanced Gemini assistant in BigQuery Studio&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, transforming the agent from a code assistant into a fully context-aware analytics partner.&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;February 23 - February 27&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;We introduced &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/databases/managed-mcp-servers-for-google-cloud-databases"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;managed and remote MCP support for Google Cloud databases&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, including AlloyDB, Spanner, Cloud SQL, Bigtable and Firestore, to power the next generation of agents. This announcement extends the ability for AI models to plan, build, and solve complex problems, connecting to the database tools our customers leverage daily as the backbone of their work environment.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;We outlined how you can &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/build-data-agents-with-conversational-analytics-api"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;build a conversational agent in BigQuery using the Conversational Analytics API&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to help you build context-aware agents that can understand natural language, query your BigQuery data, and deliver answers in text, tables, and visual charts.&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;February 16 - February 20&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;Our customers are leveraging the full power of Looker to drive major business and technical breakthroughs. Companies like &lt;/span&gt;&lt;a href="https://cloud.google.com/customers/arrive"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Arrive&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/customers/audika"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Audika&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/customers/looker-carousell"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Carousell&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/customers/framebridge"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Framebridge&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/customers/gumgum"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GumGum&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/customers/intel-looker"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Intel&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/customers/overdose-digital"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Overdose Digital&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/customers/one-looker"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Ocean Network Express&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/customers/subskribe"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Subskribe&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://cloud.google.com/customers/promevo-looker"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Promevo&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; are leveraging Looker’s newest AI-driven capabilities, including Conversational Analytics, to transform data to insights and actions, and empower their entire organization with a single source of truth, powered by Looker’s semantic layer.&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;February 2 - February 6&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;Join us on March 4 for our webinar, Win Your AI Strategy with Cloud SQL Enterprise Plus, to learn how to power your generative AI workloads with 3x higher performance and 99.99% availability. &lt;/span&gt;&lt;a href="https://rsvp.withgoogle.com/events/win-your-ai-strategy-with-cloud-sql-enterprise-plus" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Register today&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to discover how to build a scalable, enterprise-grade foundation for your most demanding AI applications.&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;January 26 - January 30&lt;/span&gt;&lt;/span&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;We introduced &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/introducing-conversational-analytics-in-bigquery"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Conversational Analytics in BigQuery&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;, which allows users to analyze data using natural language.&lt;/span&gt;&lt;/a&gt; &lt;span style="vertical-align: baseline;"&gt;Conversational Analytics in BigQuery is an intelligent agent that generates, executes and visualizes answers grounded in your business context directly in BigQuery Studio, making data insights for data professionals more conversational.&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;We outlined how &lt;/span&gt;&lt;a href="https://cloud.google.com/transform/from-asset-to-action-how-data-products-have-become-the-foundation-for-ai-agents"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;data products have become the foundation for AI agents&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, providing the context needed to make autonomous agents reliable and trusted for real business use, backed by organized business logic and semantic understanding.&lt;/span&gt;&lt;/li&gt;
&lt;li&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;We highlighted how &lt;/span&gt;&lt;a href="https://cloud.google.com/use-cases/data-analytics-agents"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;you can supercharge data analytics workflows&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and outlined Google Cloud’s AI agent offerings for data engineering, data science, and development tools, so you can integrate agentic workflows in your applications, empower your teams and speed discovery.&lt;/span&gt;&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;January 19 - January 23&lt;/span&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;We have fundamentally reimagined &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/new-firestore-query-engine-enables-pipelines"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Firestore with pipeline operations for Enterprise edition&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Experience a powerful new engine featuring over a hundred new query features, index-less queries, new index types, and observability tooling to improve query performance. Seamlessly migrate using built-in tools and leverage Firestore’s existing differentiated serverless foundation, virtually unlimited scale, and industry-leading SLA. Join a community of 600K developers to craft expressive applications that maximize the benefits of rich queryability, real-time listen queries, robust offline caching, and cutting-edge AI-assistive coding integrations.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://www.mssqltips.com/sqlservertip/11578/introducing-google-cloud-sql/" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Introducing Google Cloud SQL on MSSQLTips&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; We are highlighting a new technical guide published on MSSQLTips titled "Introducing Google Cloud SQL." This article serves as an essential resource for SQL Server administrators and developers exploring Google Cloud's fully managed database service. It provides a detailed overview of Cloud SQL capabilities, including high availability, security integration, and the seamless transition of on-premises SQL Server workloads to the cloud, making it an ideal resource for those planning their migration strategy.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="vertical-align: baseline;"&gt;We are excited to announce the &lt;/span&gt;&lt;strong&gt;&lt;a href="https://medium.com/google-cloud/bridging-the-identity-gap-microsoft-entra-id-integration-with-cloud-sql-for-sql-server-a30207d63035" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Public Preview of Microsoft Entra ID&lt;/span&gt;&lt;/a&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; (formerly Azure Active Directory) integration with Cloud SQL for SQL Server. Designed to tackle the challenge of identity sprawl in multi-cloud environments, this integration allows organizations to govern database access using their existing Microsoft identity infrastructure. Key benefits include centralized identity management, enhanced security features like Multi-Factor Authentication (MFA), and simplified user administration through direct group mapping. This feature is available for SQL Server 2022 and supports both public and private IP configurations.&lt;/span&gt;&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;January 12 - January 16&lt;/strong&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Google-built JDBC Driver for BigQuery is now available in Preview&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;We are excited to announce the launch of the new, Google-built JDBC driver for BigQuery. This new open-source driver provides a direct, high-performance connection for Java applications to BigQuery and is developed entirely in-house by Google. &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/jdbc-for-bigquery"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Download a new driver and connect your Java application to BigQuery&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong style="vertical-align: baseline;"&gt;Troubleshoot Airflow tasks instantly with Gemini Cloud Assist investigations:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Cloud Composer just got smarter. We are excited to announce that &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Gemini Cloud Assist investigations &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;are now available directly within&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; Cloud Composer 3&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;. Instead of manually sifting through raw logs, you can now simply click "Investigate" on a failed Airflow task. Gemini analyzes logs and task metadata to identify failure patterns—such as resource exhaustion or timeouts—and provides actionable recommendations driven by Gemini Cloud Assist to resolve the issue. This integration shifts the debugging experience from manual toil to automated root cause analysis, significantly reducing the time required to restore your pipelines.&lt;/span&gt; &lt;a href="https://docs.cloud.google.com/composer/docs/composer-3/troubleshooting-dags#investigations"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Learn more about AI-assisted troubleshooting&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-related_article_tout"&gt;





&lt;div class="uni-related-article-tout h-c-page"&gt;
  &lt;section class="h-c-grid"&gt;
    &lt;a href="https://cloud.google.com/blog/products/data-analytics/whats-new-with-google-data-cloud-2025/"
       data-analytics='{
                       "event": "page interaction",
                       "category": "article lead",
                       "action": "related article - inline",
                       "label": "article: {slug}"
                     }'
       class="uni-related-article-tout__wrapper h-c-grid__col h-c-grid__col--8 h-c-grid__col-m--6 h-c-grid__col-l--6
        h-c-grid__col--offset-2 h-c-grid__col-m--offset-3 h-c-grid__col-l--offset-3 uni-click-tracker"&gt;
      &lt;div class="uni-related-article-tout__inner-wrapper"&gt;
        &lt;p class="uni-related-article-tout__eyebrow h-c-eyebrow"&gt;Related Article&lt;/p&gt;

        &lt;div class="uni-related-article-tout__content-wrapper"&gt;
          &lt;div class="uni-related-article-tout__image-wrapper"&gt;
            &lt;div class="uni-related-article-tout__image" style="background-image: url('https://storage.googleapis.com/gweb-cloudblog-publish/images/whats_new_data_cloud_fWg4bKK.max-500x500.png')"&gt;&lt;/div&gt;
          &lt;/div&gt;
          &lt;div class="uni-related-article-tout__content"&gt;
            &lt;h4 class="uni-related-article-tout__header h-has-bottom-margin"&gt;What’s new with Google Data Cloud - 2025&lt;/h4&gt;
            &lt;p class="uni-related-article-tout__body"&gt;Recent product news and updates from our data analytics, database and business intelligence teams.&lt;/p&gt;
            &lt;div class="cta module-cta h-c-copy  uni-related-article-tout__cta muted"&gt;
              &lt;span class="nowrap"&gt;Read Article
                &lt;svg class="icon h-c-icon" role="presentation"&gt;
                  &lt;use xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="#mi-arrow-forward"&gt;&lt;/use&gt;
                &lt;/svg&gt;
              &lt;/span&gt;
            &lt;/div&gt;
          &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/a&gt;
  &lt;/section&gt;
&lt;/div&gt;

&lt;/div&gt;</description><pubDate>Thu, 10 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/whats-new-with-google-data-cloud/</guid><category>Databases</category><category>Business Intelligence</category><category>Data Analytics</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/whats_new_data_cloud_fWg4bKK.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>What’s new with Google Data Cloud</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/original_images/whats_new_data_cloud_fWg4bKK.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/whats-new-with-google-data-cloud/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>The Google Cloud Data Analytics, BI, and Database teams </name><title></title><department></department><company></company></author></item><item><title>Agentic analytics with the Data Agent Kit</title><link>https://cloud.google.com/blog/products/data-analytics/agentic-analytics-with-the-data-agent-kit/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Imagine your director sends you a chat message Monday morning: &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;Our average order value dropped 7% in January, but total revenue stayed flat. Why?&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If you’re a data practitioner, you know why these types of questions can be tough. They’re totally open ended. There’s not a single root cause dashboard you can open. Was there an error in the web logs? Was a promo code misconfigured? You won’t know until you start digging, and you rarely find the answer in just one place.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Each piece of the answer lives somewhere different in your environment:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Sales history&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; (orders and line items) sits in a data warehouse&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Live customer records&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; are in a production PostgreSQL instance&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Marketing campaign rules&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; are raw JSON files in an object store&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Writing any one of these queries is easy. You’ll write the same one a dozen times, tweaking &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;WHERE&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; clauses or adding subqueries to find the answer. Then you’ll bounce to the next system and start again with a different dialect. Before you know it, you have ten browser tabs open and a whole afternoon gone, all to answer one question.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Data Agent Kit&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/data-agent-kit?utm_campaign=CDR_0xaea1deef_default_b548660329&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Data Agent Kit&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is built to solve this issue. It is a set of MCP servers and agent skills that helps data developers run data workflows from their IDEs. It’s available both as an extension for VS Code forks (Antigravity IDE, Cursor) and as a &lt;/span&gt;&lt;a href="https://github.com/gemini-cli-extensions/data-agent-kit-starter-pack" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;plugin&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for other tools (Antigravity 2.0, Antigravity CLI, Claude Code, Codex), so you don’t need to leave your IDE to get answers.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The Data Agent Kit relies on two core mechanisms:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Model Context Protocol (MCP):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; an open standard that connects your agent to tools, databases, and remote cloud infrastructure.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Skills:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; markdown files that augment your agent’s knowledge, teaching it how to interact with your specific stack.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Instead of generating SQL snippets and copy-pasting them into a console, Data Agent Kit lets agents run the queries and read the results on your behalf.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Let’s see what this looks like in practice applied to the average order value scenario. In this setup, the data warehouse is &lt;/span&gt;&lt;a href="https://cloud.google.com/bigquery?utm_campaign=CDR_0xaea1deef_default_b548660329&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;BigQuery&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, the Postgres instance is &lt;/span&gt;&lt;a href="https://cloud.google.com/sql?utm_campaign=CDR_0xaea1deef_default_b548660329&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud SQL&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and the campaign rules sit in &lt;/span&gt;&lt;a href="https://cloud.google.com/storage?utm_campaign=CDR_0xaea1deef_default_b548660329&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Storage&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_dak_architecture.max-1000x1000.png"
        
          alt="1_dak_architecture"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="0ubmi"&gt;Data Agent Kit sample architecture&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Finding out what happened&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The investigation begins in the IDE’s chat pane with the following natural language prompt to confirm the baseline numbers:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;Calculate our monthly average order value from August 2025 through January 2026 using the orders and order items tables in BigQuery.&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca35a44bd0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Checking its work&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The agent processes your prompt, invokes relevant skills, and prepares to start querying your data. But before it can execute anything, the IDE  pauses to ask for permissions to use the necessary MCP tools (e.g. &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;execute_sql_readonly&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;). You can allow it once for auditing, or select “always allow” to keep the workflow moving. Once approved, the agent sends off the queries.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/2_skill_tool_use.gif"
        
          alt="2_skill_tool_use"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="0ubmi"&gt;Invoking skills and BigQuery MCP from chat&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Agentic IDEs allow you to inspect the execution trail, which reveals items like each MCP tool call or the raw SQL sent to BigQuery. It’s important to keep an eye on generated code, though reading a query can take much less time than writing one against schemas you’re unfamiliar with.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Breaking down the numbers&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The numbers showed that average order value remained around $110 from August to December, but dropped to $103 in January. To find out why, ask the agent to drill down:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;quot;Break down January&amp;#x27;s AOV by order type to see what&amp;#x27;s going on&amp;quot;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca35a36390&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The results point to a skewed average instead of a business decline. Online and Offline orders stayed healthy (~$110). A new channel called &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;B2B-Wholesale&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; appeared in January with an AOV of just ~$75. Nothing declined, but the product mix changed.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Crossing into Cloud SQL&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;You know &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;what&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; led to lower AOV. Next, you need to figure out &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;who&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; the wholesale buyers are. The customer records are stored in a Cloud SQL Postgres operational database, and you can continue in the same chat thread:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;Who are these B2B customers? Check our Cloud SQL database for their account details and creation dates.&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca340cfd90&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The agent switches to the Cloud SQL MCP and inspects the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;customers&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; table for you. All 100 wholesale accounts are brand-new business entities created within the last 30 days. None of them existed in December.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/3_b2b_customers.gif"
        
          alt="3_b2b_customers"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="0ubmi"&gt;Querying operational customer records in Cloud SQL&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Dropping into the terminal&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;A quick glance at the B2B orders in BigQuery shows that 92% applied &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;promo_code = BIGORDER25&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;. You can then ask the agent to track that code back to the campaign files, and it will use the Google Cloud Storage MCP server to access the file.  &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The marketing campaign shows a 25% discount code led to a huge number of low-priced wholesale orders, which reduced the blended AOV while total revenue remained flat.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In a single chat session, the agent queried analytical data (BigQuery), operational records (Cloud SQL), and unstructured metadata (Cloud Storage) to find the root cause.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Updating the director&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Now, you can prompt the agent to return a short executive summary for your director.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/4_executive_summary.max-1000x1000.png"
        
          alt="4_executive_summary"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="0ubmi"&gt;Agent-generated executive summary&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;And voilà! With a few natural language prompts straight from your IDE, you've answered the director's open ended question.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Root cause analysis is only part of the job. The next time this issue occurs, you won't want to run through the same situation. Instead, you can turn this investigation into a reproducible data model.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Build a reproducible pipeline&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Ask the agent to turn your ad-hoc analysis into a persistent dbt project:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;Build a dbt project that joins our BigQuery staging models with our Cloud SQL customer and pet profile attributes. Add a uniqueness test on order_id and run dbt build.&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca3521b590&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;From a single prompt, the agent creates a virtual Python environment with &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;dbt-bigquery&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; and writes project models and tests. But then &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;dbt build&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; fails. The uniqueness test catches duplicates on &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;order_id&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Customers can own more than one pet. The first version of the model attached those profiles directly to each order, so an order from a three-pet household became three rows (not unique).&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The agent reads its own terminal output and catches the failure. It then rewrites the dbt logic and reruns it until the build passes.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This introduces an important note about agentic workflows. Agents are capable of writing mountains of code - but you'll still need to apply data quality checks to your pipeline (fortunately, an agent can write those too).&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The next time leadership asks why average order value moved, you'll have a dbt model ready to answer it.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Wrap up&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;An agentic IDE keeps you from bouncing between your warehouse, your databases, your object store, and your terminal.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By pairing open standards like MCP and modular (and editable!) agent skills, the Data Agent Kit removes the friction between question and answer. Combing through unfamiliar schemas, translating between dialects, writing the joins you’ve written a hundred times: that becomes the agent’s job. You’re in charge of directing the investigation.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Try it yourself&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The Data Agent Kit is in preview and works natively in Antigravity (2.0, CLI, IDE), Claude Code, Codex, Cursor, and other popular tools.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Try the Scenario:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; walk through the full setup in the &lt;/span&gt;&lt;a href="https://codelabs.developers.google.com/dak-analytics-eng-antigravity-ide#0" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Analytics with Data Agent Kit and Antigravity IDE Codelab&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Read the Docs:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; learn more at the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/data-cloud-extension?utm_campaign=CDR_0xaea1deef_default_b548660329&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Data Agent extension documentation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Explore the Plugin:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; check out the skills and tools in the open-source repository on &lt;/span&gt;&lt;a href="https://github.com/gemini-cli-extensions/data-agent-kit-starter-pack" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GitHub&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Tue, 08 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/agentic-analytics-with-the-data-agent-kit/</guid><category>Data Analytics</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/0_hero_image_5no6K6G.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Agentic analytics with the Data Agent Kit</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/0_hero_image_5no6K6G.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/agentic-analytics-with-the-data-agent-kit/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Jeff Nelson</name><title>Developer Advocate, Google</title><department></department><company></company></author></item><item><title>How Yahoo optimizes resources with flexible VMs in Managed Service for Apache Spark</title><link>https://cloud.google.com/blog/products/data-analytics/how-yahoo-optimizes-apache-spark-with-flexible-vms/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;As a global media and technology company connecting hundreds of millions of users to finance, sports, and entertainment platforms, Yahoo operates a massive data infrastructure where analytics workloads must run continuously at high speed. In deadline-driven data environments, relying on fixed virtual machine (VM) configurations creates a brittle system; if a specific machine shape faces a regional capacity constraint, cluster provisioning in &lt;/span&gt;&lt;a href="https://cloud.google.com/products/managed-service-for-apache-spark"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Managed Service for Apache Spark&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (formerly Dataproc) can experience delays and stall critical data pipelines.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Yahoo utilizes &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/managed-spark/docs/concepts/configuring-clusters/flexible-vms"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;flexible VMs&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; in &lt;/span&gt;&lt;a href="https://cloud.google.com/products/managed-service-for-apache-spark"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Managed Service for Apache Spark&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; clusters to automatically absorb these resource fluctuations by defining a ranked list of acceptable VM shapes. This allows the system to dynamically search regional zones and maintain pipeline execution without manual intervention. To search for capacity across a region, teams must also enable &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/managed-spark/docs/concepts/configuring-clusters/flexible-vms"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Auto-Zone placement&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This optimization builds on Yahoo's broader data modernization journey, which involved &lt;/span&gt;&lt;a href="https://www.youtube.com/watch?v=_7Oz1V1-ZiE" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;migrating on-premises Hadoop and big data estates&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; directly to Google Cloud. By transitioning those legacy workloads, the team established a cloud foundation capable of running high-scale batch and streaming analytics with dynamic resource flexibility.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-video"&gt;



&lt;div class="article-module article-video "&gt;
  &lt;figure&gt;
    &lt;a class="h-c-video h-c-video--marquee"
      href="https://youtube.com/watch?v=_7Oz1V1-ZiE"
      data-glue-modal-trigger="uni-modal-_7Oz1V1-ZiE-"
      data-glue-modal-disabled-on-mobile="true"&gt;

      
        

        &lt;div class="article-video__aspect-image"
          style="background-image: url(https://storage.googleapis.com/gweb-cloudblog-publish/images/maxresdefault_iMaqL8o.max-1000x1000.jpg);"&gt;
          &lt;span class="h-u-visually-hidden"&gt;Hadoop pioneer to cloud innovator: Yahoo’s data lake modernization journey&lt;/span&gt;
        &lt;/div&gt;
      
      &lt;svg role="img" class="h-c-video__play h-c-icon h-c-icon--color-white"&gt;
        &lt;use xlink:href="#mi-youtube-icon"&gt;&lt;/use&gt;
      &lt;/svg&gt;
    &lt;/a&gt;

    
  &lt;/figure&gt;
&lt;/div&gt;

&lt;div class="h-c-modal--video"
     data-glue-modal="uni-modal-_7Oz1V1-ZiE-"
     data-glue-modal-close-label="Close Dialog"&gt;
   &lt;a class="glue-yt-video"
      data-glue-yt-video-autoplay="true"
      data-glue-yt-video-height="99%"
      data-glue-yt-video-vid="_7Oz1V1-ZiE"
      data-glue-yt-video-width="100%"
      href="https://youtube.com/watch?v=_7Oz1V1-ZiE"
      ng-cloak&gt;
   &lt;/a&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This post provides a technical blueprint for configuring flexible VM instance rankings in &lt;/span&gt;&lt;a href="https://cloud.google.com/products/managed-service-for-apache-spark"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Managed Service for Apache Spark&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to automatically manage capacity constraints and maintain pipeline execution.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Operational trade-offs of static configurations&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Configuring clusters with a single, fixed machine type in a specific zone introduces constraints when regional zonal capacity fluctuations occur, potentially impacting cluster provisioning. Rather than manage these capacity variations through custom retry logic or manual intervention, using flexible configurations allows your infrastructure to automatically adapt. By accepting multiple VM shapes and searching across zones in the selected region, flexible configurations help streamline provisioning to better support high-scale analytics workloads.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Rules for configuring flexible clusters&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Deploying flexible configurations requires aligning several connected design choices:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Enable auto-zone placement:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; You must pass a region(&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;--region=${REGION}&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;) or an empty zone string (&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;--zone=""&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;) so Managed Spark can search for available capacity across the entire region.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Maintain core and memory symmetry:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; If your Managed Spark cluster uses &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/managed-spark/docs/concepts/configuring-clusters/autoscaling"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;autoscaling&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, all machine types in your flexible list must share a similar core count and memory size, even if they come from different VM families. A uniform CPU-to-memory ratio across primary and secondary workers prevents performance degradation, as the smallest ratio determines your effective container sizing.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Align component properties:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Managed Spark calculates system properties based on VM cores and memory. When mixing machine shapes, you may need explicit property overrides to keep YARN and Spark resource allocations aligned with your expected worker behavior.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Two ways flexible VMs support massive workloads&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For large-scale data environments, flexible configurations support operations in two ways:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Higher cluster creation success:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Instead of failing when a preferred VM type is out of stock, Managed Spark selects from a ranked list to keep provisioning moving.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Better regional resource use:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Auto-zone placement searches the entire region to find capacity, which reduces provisioning friction during high-demand periods.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;gcloud example&lt;/strong&gt;&lt;/h3&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;gcloud dataproc clusters create analytics-cluster \\\r\n  --region=us-central1 \\\r\n  --zone=&amp;quot;&amp;quot; \\\r\n  --num-workers=10 \\\r\n  --master-instance-selection=\&amp;#x27;{&amp;quot;machineTypes&amp;quot;:[&amp;quot;e2-standard-8&amp;quot;],&amp;quot;rank&amp;quot;:0}\&amp;#x27; \\\r\n  --master-instance-selection=\&amp;#x27;{&amp;quot;machineTypes&amp;quot;:[&amp;quot;n2-standard-8&amp;quot;],&amp;quot;rank&amp;quot;:1}\&amp;#x27; \\\r\n  --worker-instance-selection=\&amp;#x27;{&amp;quot;machineTypes&amp;quot;:[&amp;quot;e2-standard-8&amp;quot;],&amp;quot;rank&amp;quot;:0}\&amp;#x27; \\\r\n  --worker-instance-selection=\&amp;#x27;{&amp;quot;machineTypes&amp;quot;:[&amp;quot;n2-standard-8&amp;quot;],&amp;quot;rank&amp;quot;:1}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca358933d0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;API example&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;You can also build this capacity policy into your automated pipelines or &lt;/span&gt;&lt;a href="https://cloud.google.com/products/managed-service-for-apache-airflow"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Managed Service for Apache Airflow&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; DAGS using the &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;instanceFlexibilityPolicy&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; field in the ‘Dataproc’ API:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;{\r\n  &amp;quot;projectId&amp;quot;: &amp;quot;PROJECT_ID&amp;quot;,\r\n  &amp;quot;clusterName&amp;quot;: &amp;quot;analytics-cluster&amp;quot;,\r\n  &amp;quot;config&amp;quot;: {\r\n    &amp;quot;gceClusterConfig&amp;quot;: {\r\n      &amp;quot;zoneUri&amp;quot;: &amp;quot;&amp;quot;\r\n    },\r\n    &amp;quot;secondaryWorkerConfig&amp;quot;: {\r\n      &amp;quot;numInstances&amp;quot;: 8,\r\n      &amp;quot;instanceFlexibilityPolicy&amp;quot;: {\r\n        &amp;quot;instanceSelectionList&amp;quot;: [\r\n          {\r\n            &amp;quot;machineTypes&amp;quot;: [&amp;quot;n2-standard-8&amp;quot;],\r\n            &amp;quot;rank&amp;quot;: 0\r\n          },\r\n          {\r\n            &amp;quot;machineTypes&amp;quot;: [&amp;quot;e2-standard-8&amp;quot;, &amp;quot;t2d-standard-8&amp;quot;],\r\n            &amp;quot;rank&amp;quot;: 1\r\n          }\r\n        ]\r\n      }\r\n    }\r\n  }\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca350f8850&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This API policy achieves the same goal: it establishes your preferred shape, documents valid fallbacks, and lets Managed Spark resolve resource constraints without breaking your automation scripts.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Establishing an infrastructure policy&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Managing data at this scale requires standardizing a clear resource policy rather than relying on a single rigid machine type. Your configuration standards should outline:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Preferred and fallback VM families for secondary workers.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Default auto-zone placement to enable flexible provisioning.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Identical core and memory configurations when using autoscaling.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Uniform CPU-to-memory ratios across all worker groups to maintain predictable container sizing.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Explicit YARN or Spark property overrides to guarantee consistent runtime behavior across different machine lines.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Shuffle-safe patterns for Spark workloads running on Spot or highly elastic capacity.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By adopting flexible configurations, you turn infrastructure scarcity into a predictable fallback plan, keeping your critical data pipelines up and running.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Yahoo impact and results&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By implementing flexible VMs in Managed Service for Apache Spark, Yahoo successfully reduced cluster provisioning failures by 85% which were caused by regional capacity stockouts. This flexible configuration allows their data infrastructure to automatically handle capacity constraints and successfully provision resources without requiring manual intervention. As a result, Yahoo ensures continuous workload execution and prevents downstream processing delays across their massive data pipelines.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;"Managing high-scale data analytics at Yahoo requires resilient, automated infrastructure. Moving to flexible VMs in Managed Service for Apache Spark has transformed our approach; instead of stalling when a specific machine shape faces capacity constraints, our clusters now automatically pivot to our ranked fallback options. This has helped us reduce provisioning failures by 85%, providing the reliability we need to keep our global media platforms running smoothly."&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; - Akshay Jain, Senior Software Developer Engineer, Yahoo! &lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Strategic benefits of flexible infrastructure&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Adopting a flexible compute stack transforms your environment into a dynamic pool of resources that adapts to your operational needs. By moving away from rigid, single-machine type configurations, you ensure that your workloads reliably access the compute they need, regardless of supply fluctuations. This shift not only maximizes workload obtainability and reliability but also facilitates seamless hardware modernization by allowing you to prioritize newer VM generations while maintaining older types as reliable fallback options.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Build your resilient data pipeline&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Transitioning to a fluid compute strategy ensures your critical analytics remain operational despite regional resource shifts. Here is how you can begin optimizing your infrastructure today:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Audit your workloads: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Identify applications tightly coupled to specific VM families or zones and map out viable alternative hardware shapes.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Standardize resource policies: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Explore the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/managed-spark/docs/concepts/configuring-clusters/flexible-vms"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;documentation for Managed Spark flexible VMs&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to establish your preferred and fallback VM families.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Align financial strategy: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Utilize Flexible Committed Use Discounts (Flex CUDs) to maintain cost predictability when workloads dynamically pivot to alternative machine types.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Claim your credits: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;New customers may be eligible for &lt;/span&gt;&lt;a href="https://cloud.google.com/free"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;$300 in credits&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to try Managed Service for Apache Spark and other Google Cloud products at no cost.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;&lt;/div&gt;</description><pubDate>Fri, 04 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/how-yahoo-optimizes-apache-spark-with-flexible-vms/</guid><category>Streaming</category><category>Customers</category><category>Data Analytics</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>How Yahoo optimizes resources with flexible VMs in Managed Service for Apache Spark</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/how-yahoo-optimizes-apache-spark-with-flexible-vms/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Akshay Jain</name><title>Senior Software Engineer, Yahoo</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Surjit Singh</name><title>Data &amp; AI Engineer, Google Cloud</title><department></department><company></company></author></item><item><title>Simplify pipelines with new BigQuery identity columns</title><link>https://cloud.google.com/blog/products/data-analytics/bigquery-identity-columns-to-auto-generate-sequential-integers/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To further empower our customers in their data journey, we are excited to announce the launch of identity columns in BigQuery. This new feature allows users to define columns that automatically generate sequential 64-bit integer values, simplifying the way you manage unique identifiers within your tables.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Data engineers are always looking for ways to make data ingestion smoother and more reliable. BigQuery identity columns offer a powerful, built-in mechanism to automatically generate unique numerical values for your tables. By shifting the responsibility of ID generation to BigQuery, you can significantly reduce the complexity of your data pipelines and focus on delivering insights.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Key benefits for your data pipelines&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Implementing identity columns provides several advantages that help streamline the development and maintenance of your data architecture.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Streamlined ingestion&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: You can now ingest data without needing to pre-calculate unique keys in your application logic or ETL tools.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Reduced boilerplate&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: By using auto-generated sequences, your SQL code becomes cleaner and easier to maintain, as the database handles key management natively.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Integrated automation&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Identity columns work harmoniously with standard DML operations, ensuring that every new row receives a unique identifier automatically.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Flexible integration&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Whether you are using &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;INSERT&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; or &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;MERGE&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; statements, identity columns adapt to your existing workflow.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;How to implement identity columns&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Setting up an identity column is simple and can be done directly within your &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;CREATE TABLE&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; statement. You have two primary ways to define how these values are handled.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Definition options&lt;/strong&gt;&lt;/p&gt;
&lt;div align="left"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;&lt;table&gt;&lt;colgroup&gt;&lt;col/&gt;&lt;col/&gt;&lt;/colgroup&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th scope="col" style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Clause&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Description&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;code style="vertical-align: baseline;"&gt;GENERATED ALWAYS AS IDENTITY&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;BigQuery automatically manages and ensures the uniqueness of the values.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;code style="vertical-align: baseline;"&gt;GENERATED BY DEFAULT AS IDENTITY&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Provides an automatic value but still allows for manual overrides when necessary.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Example usage&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;The following SQL statement demonstrates how to create a table that automatically increments IDs, starting at 1 and increasing by one for each new entry.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;quot;CREATE TABLE my_project.my_dataset.orders (\r\n  order_id INT64 GENERATED ALWAYS AS IDENTITY (START WITH 1 INCREMENT BY 1),\r\n  customer_name STRING,\r\n  order_date DATE\r\n);\r\n\r\n-- Ingesting data is now simpler:\r\nINSERT INTO my_project.my_dataset.orders (customer_name, order_date)\r\nVALUES (&amp;#x27;Joe Doe&amp;#x27;, CURRENT_DATE());&amp;quot;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca34e00650&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Get started today&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Identity columns represent our ongoing commitment to providing a flexible, high-performance, and standards-compliant data platform. By automating the generation of surrogate keys, we are making it easier for you to build scalable and maintainable data architecture.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To learn more about how to implement this feature in your projects, please visit the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/identity-columns"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;BigQuery identity columns documentation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Wed, 02 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/bigquery-identity-columns-to-auto-generate-sequential-integers/</guid><category>BigQuery</category><category>Data Analytics</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Simplify pipelines with new BigQuery identity columns</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/bigquery-identity-columns-to-auto-generate-sequential-integers/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Wawrzek Hyska</name><title>Product Manager</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Aayush Bhatnagar</name><title>Software engineer</title><department></department><company></company></author></item><item><title>Introducing TabFM in BigQuery: Predictive analytics reimagined</title><link>https://cloud.google.com/blog/products/data-analytics/tabfm-adds-predictive-ml-to-bigquery/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Historically, enterprise predictive analytics tasks such as predicting churn, purchase intent, or fraud scoring have meant building custom models using libraries like XGBoost, Random Forest, or Deep Neural Networks (DNNs). While effective, the traditional train-tune-deploy-retrain cycle can be complex and time-consuming. Additionally, the overhead of manual feature engineering, hyperparameter tuning, lengthy and expensive training, and the need for specialized data science skills can lead businesses to underutilize predictive models in their decision-making. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Today, we are announcing the TabFM model in BigQuery. Developed by Google Research, TabFM is a state-of-the-art, pre-trained foundation model for regression and classification on tabular data. It leverages in-context learning (ICL) to deliver highly accurate predictions on your tabular datasets instantly via a single SQL statement, removing the separate training and deployment steps. TabFM on BigQuery is currently in preview. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Here is what TabFM brings to your BigQuery analytics:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Zero-shot predictions&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Skip model training, tuning, and artifact deployment. Simply pass your labeled historical data and new prediction tables into a single SQL function to get instant, high-quality predictions.&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Predictive ML for your agentic applications&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Building an agent for your business use? Add predictive powers to it with TabFM plus &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/use-bigquery-mcp"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;BigQuery MCP server&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. No runtimes or infrastructure to manage, just data in and predictions out.&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;State-of-the-art accuracy&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Outperforms custom-trained, out-of-the-box traditional models on complex datasets, achieving superior accuracy scores on industry benchmarks.&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Simple developer experience&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Runs natively in BigQuery and is accessible via simple SQL syntax. Automatically handles featurization tasks such as missing values, categorical encoding, etc., with no complex feature engineering pipelines to manage.&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Scalability:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Processes massive inference tables (up to millions of rows) in minutes using BigQuery’s distributed inference architecture.&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;The leading model for tabular predictions&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Google’s TabFM delivers industry-leading accuracy across a wide range of tabular data. In evaluations on the &lt;/span&gt;&lt;a href="https://huggingface.co/spaces/TabArena/leaderboard" rel="noopener" target="_blank"&gt;&lt;span style="vertical-align: baseline;"&gt;TabArena&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; benchmark, TabFM consistently outperforms both classic machine learning models and other tabular foundation mod&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;els.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/image1_vsRpJjZ.max-1000x1000.png"
        
          alt="image1"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="8522v"&gt;ELO ratings (↑) for the top 10 models across TabArena classification (upper) and regression (lower). (D) = default; (T+E) = tuned + ensemble. Higher scores denote superior performance.&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Learn more about the TabFM model &lt;/span&gt;&lt;a href="https://research.google/blog/introducing-tabfm-a-zero-shot-foundation-model-for-tabular-data/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Getting started with TabFM in BigQuery&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Using TabFM is straightforward. It is exposed directly through new, built-in SQL functions: AI.PREDICT and AI.EVALUATE.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;1. Get instant predictions with AI.PREDICT&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;To make predictions, you write a single query that passes your training  data and prediction data. The model automatically infers whether the task is a classification or regression problem based on the data type of your target label.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;quot;-- Classifying transactions as fraudulent or not\r\nSELECT *\r\nFROM AI.PREDICT(\r\n  TABLE `my_project.my_dataset.historical_transactions`, -- Training data (in-context examples)\r\n  TABLE `my_project.my_dataset.new_transactions`, -- Prediction data\r\n  label_col =&amp;gt; &amp;#x27;is_fraud&amp;#x27;-- Target column to predict\r\n);&amp;quot;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca35a45450&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In this example, the output contains all original columns from your prediction table plus predicted label and probability columns (e.g. predicted_is_fraud). No manual feature engineering or model creation was required.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;2. Evaluate models with AI.EVALUATE&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;You can quickly check prediction performance against a test set using the AI.EVALUATE function. This allows you to generate standard evaluation metrics in a single step.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;quot;-- Regression Evaluation for Customer Lifetime Value (LTV)\r\nSELECT *\r\nFROM AI.EVALUATE(\r\n  TABLE `my_project.my_dataset.historical_customer_ltv`,\r\n  TABLE `my_project.my_dataset.test_customer_ltv`,\r\n  label_col =&amp;gt; &amp;#x27;ltv&amp;#x27;\r\n);&amp;quot;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca2ff303d0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;AI.EVALUATE returns a robust set of metrics such as r2_score, mean_absolute_error etc. for regression problems and metrics such as precision, recall, and f1 for classification problems.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;TabFM in BigQuery under the hood&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Traditional machine learning requires fitting model parameters to a training dataset. TabFM, in contrast, uses in-context learning. Similar to how large language models (LLMs) learn a task from few-shot examples in a prompt, TabFM reads your training table as in-context examples and generates predictions for your target table in a single forward pass.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To handle the computational complexity and memory footprint of tabular foundation models, BigQuery performs distributed, parallelized inference on your data. Further, to optimize performance and resource utilization, it uses intelligent training-data sampling as well as distributed execution. This allows BigQuery to handle large input rows for training data while executing predictions quickly and efficiently across millions of rows of inference data.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Choosing the right tool for the job&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;TabFM introduces groundbreaking zero-shot capabilities to BigQuery, and complements existing offerings such as XGBoost models. Here’s how to choose between TabFM and other models:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Use TabFM when you need rapid, high-quality predictive insights without machine learning expertise, when historical datasets are small-to-medium sized, when data changes frequently, and when you need to retrain your models frequently to maintain accuracy. It is also a great fit for conversational or agentic workflows where you need predictive analysis on demand.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Use traditional models like XGBoost when you have very large historical datasets, require complete control over custom hyperparameter tuning, have a high number of features that exceed current limits of TabFM, or need feature-importance explainability, i.e., which of the input features contributed most to the prediction.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Predictive machine learning made easy&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With TabFM natively integrated into BigQuery, predictive ML is now as easy as running a standard SELECT query. By eliminating the manual overhead of model training, tuning, and management, TabFM lets developers, data scientists and analysts go from raw data to rich predictive insights in seconds.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To get started today, check out the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/reference/standard-sql/bigqueryml-syntax-ai-predict"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;public documentation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. For questions or feedback reach out to our team at &lt;/span&gt;&lt;a href="mailto:bqml_feedback@google.com"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;bqml_feedback@google.com&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.  We look forward to seeing what you build!&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Tue, 01 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/tabfm-adds-predictive-ml-to-bigquery/</guid><category>BigQuery</category><category>Data Analytics</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Introducing TabFM in BigQuery: Predictive analytics reimagined</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/tabfm-adds-predictive-ml-to-bigquery/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Vaibhav Sethi</name><title>Group Product Manager</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Xi Cheng</name><title>Engineering Manager</title><department></department><company></company></author></item><item><title>BigQuery Graph is now GA: the knowledge foundation for the agentic era</title><link>https://cloud.google.com/blog/products/data-analytics/bigquery-graph-connecting-data-and-ai-at-scale/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Many of the questions that matter in enterprise data aren't just about individual rows — they're about how things connect: how two accounts are linked, what path a payment took, what context grounds an AI agent's answer. That’s what a graph is built to solve. Historically, unlocking these insights meant extracting data into standalone graph databases, creating silos and operational overhead. To remove these barriers, we brought native graph capabilities directly to the data warehouse. Today, we are announcing the general availability of &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/graph-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;BigQuery Graph&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We introduced BigQuery Graph in &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/introducing-bigquery-graph?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;preview&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to unify graph and relational analytics. ISO-standard Graph Query Language (GQL) sits alongside SQL, traversals run natively, and there’s no ETL. And because it’s built on BigQuery, BigQuery Graph inherits and expands its capabilities: It reaches petabyte-scale without the memory bottlenecks of a scale-up database, runs under your existing row- and column-level security, and calls BigQuery ML and AI functions in the same query. One engine, two jobs — large-scale graph analytics, and connected context for AI agents.&lt;/span&gt;&lt;/p&gt;
&lt;p style="padding-left: 40px;"&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;"BigQuery Graph has been a game-changer for our threat detection pipeline, allowing us to move beyond simple, siloed alerts. By modeling our security signal data as a property graph, we can now perform complex, multi-hop traversals in seconds - something that was previously computationally prohibitive. This graph-centric approach automatically clusters anomalies into coherent attack stories, which, combined with the seamless integration of Gemini models, helps us generate actionable threat narratives. We look forward to integrating native BigQuery Graph algorithms to further streamline our workflows." - Pete Rubio, VP of Global engineering at Thales Cybersecurity Products&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Since preview, we saw data teams across industries adopt BigQuery Graph for both analytical and agentic workflows:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Threat and fraud detection:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;  Security and financial organizations correlate signals across event logs to uncover multi-hop attack paths, fraud networks, and suspicious transaction loops.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Supply chain digital twins&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Manufacturing and logistics organizations map dependencies across suppliers, parts, and distribution routes to simulate disruptions and optimize fulfillment.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Identity resolution and Customer 360&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Ad-tech and retail platforms stitch fragmented user identifiers and behavioral touchpoints into unified customer profiles across channels.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Knowledge graphs and AI agent grounding&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Enterprise AI teams build structured knowledge graphs from unstructured documents, providing domain context to ground Gemini models and GraphRAG workflows. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Network lineage and infrastructure management&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Telecommunications and enterprise IT teams track complex network topologies, service dependencies, and data lineage across multi-hop paths.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;What’s new in BigQuery Graph&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Reaching GA is more than a stability milestone. The work fell into two movements: we made the graph engine itself faster and broader, and we built an agentic ecosystem around it — so agents can build a graph, chat with it, and keep an auditable memory on it. Some of what follows is generally available today; some is in preview or rolling out over the coming weeks.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;A faster, broader graph engine&lt;/strong&gt;&lt;/h3&gt;
&lt;p style="padding-left: 40px;"&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;“Advertising has spent decades optimizing individual events; the agentic era will optimize the relationships between them. At Yahoo, BigQuery Graph gives our AI agents connected context - campaigns, audiences, exposures, and outcomes, traversable with standard GQL right where our monetization data already lives, with no separate graph engine and no data movement. Our agents don't just read the graph; they reason over it and write their conclusions back as new relationships. That's how monetization moves beyond automation, to autonomous systems we can trust to act.” - &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Mikul Bhatt, Director of Engineering, Monetization Platform at Yahoo&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Borderless graph Lakehouse&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Agents are only as good as the context they can reason over, and that context is rarely in one place. With &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/lakehouse/docs/about-borderless-lakehouse"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;borderless Lakehouse&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, a single BigQuery Graph can span native BigQuery tables and open Iceberg tables in other clouds — through Databricks Unity Catalog, AWS Glue, or Snowflake — traversed in place, without copying data or building ETL pipelines.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Say a support agent needs to answer, &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;"who supplies the product behind this customer's delayed order, and where are they based?"&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; The customer data sits in an Iceberg lakehouse on Google Cloud, the product and supplier records in a Databricks catalog on AWS. Instead of stitching the sources together per request, the agent traverses one virtual knowledge graph that already connects them — over data that never moved.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_Virtual_Graph.max-1000x1000.png"
        
          alt="1, Virtual Graph"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="5usnm"&gt;Figure 1: A diagram illustrating a virtual knowledge graph spanning across Google Cloud (blue nodes), AWS (yellow nodes), and other clouds (green nodes) without data movement.&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The following DDL statement shows how you can define this virtual graph, mapping your node and edge tables directly across both cloud environments:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;-- A virtual knowledge graph spanning two clouds - no data movement\r\nCREATE OR REPLACE PROPERTY GRAPH `my_project.retail.virtual_kg`\r\n  NODE TABLES (\r\n    -- Google Cloud\r\n    `my_project.gcs_lake.retail.customers` AS Customer KEY (customer_id),\r\n    -- AWS\r\n    `my_project.dbx_fed_catalog.retail.products` AS Product  KEY (product_id),\r\n    `my_project.dbx_fed_catalog.retail.suppliers` AS Supplier KEY (supplier_id)\r\n  )\r\n  EDGE TABLES (\r\n    `my_project.gcs_lake.retail.purchases` AS Bought KEY (purchase_id)\r\n      SOURCE KEY (customer_id) REFERENCES Customer (customer_id)\r\n      DESTINATION KEY (product_id) REFERENCES Product (product_id),\r\n    `my_project.dbx_fed_catalog.retail.products` AS Supplied_By KEY (product_id)\r\n      SOURCE KEY (product_id) REFERENCES Product (product_id)\r\n      DESTINATION KEY (supplier_id) REFERENCES Supplier (supplier_id)\r\n  );&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca34ed3610&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With that, the agent gets a grounded, multi-hop answer assembled across two clouds in a single traversal:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;quot;-- Agent grounding: trace a customer to the supplier behind their product, across clouds\r\nGRAPH `my_project.retail.virtual_kg`\r\nMATCH (c:Customer {customer_id: &amp;#x27;C1&amp;#x27;})-[:Bought]-&amp;gt;\r\n      (:Product)-[:Supplied_By]-&amp;gt;(s:Supplier)\r\nRETURN s.name AS supplier, s.country AS supplier_country&amp;quot;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca34ed0e90&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Faster and more expressive GQL&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;BigQuery Graph is built for questions about connection: how two accounts are linked, what path a payment took, which entities sit within a few hops of a flagged one. These are the questions SQL joins struggle to express, and they're where a graph engine earns its place. At GA, we've made them both faster to run and easier to write:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Faster execution.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; GA optimizes path-finding for acyclic and undirected traversals: against public benchmarks, GQL is 2x faster since preview and undirected traversal 100x, with faster, more resource-efficient cycle detection in &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;ACYCLIC&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;TRAIL&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; path modes. Lower query latency keeps the neighborhood and path lookups that ground an agent's answer responsive under frequent, interactive access.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;More expressive queries.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; With the new &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/reference/standard-sql/graph-query-statements#gql_call"&gt;&lt;code style="text-decoration: underline; vertical-align: baseline;"&gt;CALL&lt;/code&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt; statement &lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;and extended &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/reference/standard-sql/graph-subqueries"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;subquery&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; support, you can run a graph subquery for each entity in a result, or invoke a reusable named function, so a complex question breaks into parts instead of one sprawling pattern. The same functions an analyst writes become the building blocks an agent calls as a tool.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Built for the agentic era&lt;/strong&gt;&lt;/h3&gt;
&lt;p style="padding-left: 40px;"&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;“Companies have plenty of workforce data, but very little shared understanding of what their people can do or where they fit. BigQuery Graph lets us turn that scattered information into a reusable property graph and traverse connections across people, roles, capabilities, and evidence at scale, so the same connected workforce context can support thousands of decisions instead of being recreated one decision at a time. That gives AI a stronger foundation for much harder questions about how work should get done.”  -&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; Heiko Roth, Founder &amp;amp; CEO, Workerbee&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Chat with your graphs&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;You don't have to write GQL to explore a graph. BigQuery &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/conversational-analytics?content_ref=when%20you%20ask%20questions%20about%20your%20graph%20the%20agent%20constructs%20sql%20queries%20to%20answer%20them%20agents%20can%20use%20descriptions%20and%20synonyms%20that%20you%20define%20on%20your%20graph#graphs"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;conversational analytics&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; lets you &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/graph-chat"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;chat with your graph&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; directly in natural language: it reads the relationships in your schema to translate a question into SQL or GQL, and visualizes the traversal for path-based answers. The agent draws on graph metadata like descriptions and synonyms to keep results grounded — the relationships that make a graph a graph are exactly what cut the ambiguity and hallucination that plague free-form natural language querying. You can also connect &lt;/span&gt;&lt;a href="https://cloud.google.com/gemini-enterprise?utm_source=google&amp;amp;utm_medium=cpc&amp;amp;utm_campaign=1713762-Gemini_Enterprise-DR-NA-US-en-Google-BKWS-EXA-GEnterprise&amp;amp;utm_content=c-Hybrid+%7C+BKWS+-+MIX+%7C+Txt_Gemini+Enterprise-189528400785&amp;amp;utm_term=gemini+enterprise&amp;amp;gclsrc=aw.ds&amp;amp;gad_source=1&amp;amp;gad_campaignid=23370621055&amp;amp;gclid=Cj0KCQjw4orUBhCjARIsAIbF3qwrXsr1khkuSsBNMPTjrHynNaAJTcSyWSavMuVwERJmMqfKVkmO9LIaAjf7EALw_wcB&amp;amp;e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to BigQuery Graph through an &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/use-bigquery-mcp"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;MCP server&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, or &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/create-data-agents#publish-agent-gemini-enterprise"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;publish&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; the conversational data agent to it directly.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/2_Graph_CA_Blog_V1_2x_high_res.gif"
        
          alt="2, Graph CA Blog V1 2x high res"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Build a graph with an agent&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Standing up a graph — modeling tables into nodes and edges, then writing GQL against them — is work you can hand to the data agent you already use. We've packaged BigQuery Graph expertise into an agent skill that makes your agent fluent in graph: GQL pattern matching, blending graph and SQL, and schema design that follows our recommended practices. The capabilities are accessible out of the box in your preferred agentic coding tool, such as Antigravity, Visual Studio Code, Claude Code, and Codex, with the Google Cloud &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/data-agent-kit/overview"&gt;Data Agent Kit extension&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The skill is also learning to author, not just advise — a capability rolling out soon. Point it at a dataset, a model document, or an ER diagram and it proposes the nodes and edges, then verifies each relationship against your data before building, showing you the match rates: this one resolves at, say, 98%, that one 56%. You get a graph you can trust from day one.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/3_Graph_GA_Skill_Demo_V2.gif"
        
          alt="3, Graph GA Skill Demo V2"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Give your agents an auditable memory&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Grounding an agent is half the job; the other half is remembering what it did. As agents move from advising to acting, every decision has to be explainable after the fact — which option was chosen, which policy applied, &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;which&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; alternatives were rejected. With &lt;/span&gt;&lt;a href="https://adk.dev/integrations/bigquery-agent-analytics/#context-graph" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;context graph&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; in BigQuery Agent Analytics, each action an agent takes is captured and shaped into a context graph: a typed, queryable trace of the agent's reasoning, stored right in BigQuery Graph. Because the trace is itself a graph, "why did the agent do this?" is a single traversal — and the outcomes you join back to those decisions become the data that improves the next one.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Get started with BigQuery Graph today&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;BigQuery Graph runs graph analytics and grounds AI agents on your data, across clouds.&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;To get started, &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;check out the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/graph-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;overview and data model&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to see how GQL, node tables, and edge tables fit together, then put them to work on your team’s common patterns. Trace suspicious money movement and synthetic identities in the &lt;/span&gt;&lt;a href="https://codelabs.developers.google.com/codelabs/fraud-bigquery-graph#0" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;fraud detection codelab&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, stitch fragmented emails, devices, and cookies into one customer in the &lt;/span&gt;&lt;a href="https://codelabs.developers.google.com/codelabs/identity-resolution-bigquery-graph#0" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;identity resolution codelab&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, or model a supply chain as a &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/modeling-a-digital-twin-using-bigquery-graph?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;digital twin&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; you can query for hidden dependencies when disruption hits.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;From there, take it toward agents. The &lt;/span&gt;&lt;a href="https://codelabs.developers.google.com/bqaa-context-graph" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;agent context graph codelab&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; turns raw event logs into a graph that audits, explains, and traces what your autonomous agents actually did — the connected memory behind a system you can trust to act. If your workloads span both real-time operational transactions and massive-scale analytics, explore our &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/the-unified-graph-solution-with-spanner-graph-and-bigquery-graph?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;unified graph solution&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to see how Spanner Graph and BigQuery Graph work together. And when you are ready to go deeper — our &lt;/span&gt;&lt;a href="https://cloud.google.com/resources/graph-ebook"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;ebook&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; walks the journey end-to-end.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Mon, 31 Aug 2026 23:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/bigquery-graph-connecting-data-and-ai-at-scale/</guid><category>BigQuery</category><category>Data Analytics</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>BigQuery Graph is now GA: the knowledge foundation for the agentic era</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/bigquery-graph-connecting-data-and-ai-at-scale/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Bei Li</name><title>Sr. Staff Software Engineer</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Candice Chen</name><title>Product Manager</title><department></department><company></company></author></item><item><title>From weeks to minutes: The new agentic era of data pipelines</title><link>https://cloud.google.com/blog/products/data-analytics/build-data-pipelines-in-less-time-with-data-agent-kit/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Data pipelines are the backbone of the modern enterprise, yet a barrier to entry exists for orchestrating them, making this critical capability unavailable to many data professionals. Following our &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/managed-apache-airflow-scaling-data-and-ai-workloads"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;announcements at Google Cloud NEXT ’26&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, where we introduced the Orchestration Pipelines framework, we are fundamentally changing this dynamic.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To bring this powerful framework directly to practitioners, we offer the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/data-cloud-extension"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Data Agent Kit&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; — a unified, freely available, and open-source collection of data engineering and data science tools that integrate directly into your preferred IDE or CLI (such as VS Code, Claude Code, or Codex).&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The Data Agent Kit seamlessly embeds the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/orchestration-pipelines/overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Orchestration Pipelines&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; framework into your workflow in two distinct ways. First, it provides a dedicated &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/data-agent-kit/build-pipelines"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Data Engineering tab&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for comprehensive pipeline management. Second, it includes a specialized agentic skill designed to author, deploy, and troubleshoot production-grade &lt;/span&gt;&lt;a href="https://airflow.apache.org/docs/apache-airflow/stable/core-concepts/dags.html" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Apache Airflow® DAGs&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; using natural language.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By pairing these specialized agent skills with a declarative YAML DSL, all data personas — from analysts to ML engineers — can bypass complex Python Airflow boilerplate. This framework decouples high-level orchestration logic from underlying compute execution, democratizing access to powerful MLOps capabilities across your entire data organization.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In this post, we will walk through an exemplary MLOps use case to demonstrate how easily this can be achieved.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Setting up your environment&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Before authoring your first Orchestration Pipeline, you need to set up your local development environment. Getting started takes less than two minutes.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;1. Install and configure the extension&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To install the extension in your preferred IDE or CLI — such as VS Code, VS Code forks, Antigravity, Claude Code, Antigravity CLI, or Codex — and authenticate it with your Google Cloud account, follow the step-by-step setup guide in the official documentation:&lt;/span&gt; &lt;a href="https://docs.cloud.google.com/data-cloud-extension/vs-code/install"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Data Agent Kit installation guide&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;2. Verify orchestration pipeline skills&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Once installed, verify that the required agent skills are active:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Open the ‘&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Google Cloud Data Agent Kit’&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; panel on the VS Code activity bar.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Navigate to ‘&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Settings’&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; then ‘&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Skills’&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Ensure the ‘&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;gcp-pipelines-orchestration’&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; skill is enabled.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;a href="https://github.com/gemini-cli-extensions/data-agent-kit-starter-pack/tree/main/skills/gcp-pipeline-orchestration" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;This skill provides&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; the agent with deep contextual knowledge of pipeline syntax, variable substitution, secret management, and automated incident diagnosis for Airflow runs.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;3. Building your first pipeline&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To start authoring, building, and validating orchestration pipelines directly inside the any VS Code compatible IDE using natural language prompts, follow the official building guide: &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/data-cloud-extension/vs-code/build-pipelines"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Build pipelines guide&lt;/strong&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;An example business problem: Proactive supply chain management&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Let’s walk through an example business problem. &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;In the logistics and retail sector, customer satisfaction hinges on accurate delivery estimates. When an order is delayed without warning, customer churn can spike and support costs can escalate.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To address this, we are building an end-to-end MLOps architecture that predicts the exact transit time (in days) based on warehouse location, customer location, and order characteristics. By predicting these delays before shipping, operations teams can proactively notify customers or automatically upgrade shipping tiers before Service Level Agreements (SLAs) are breached.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To make this architecture fully reproducible, we use the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;bigquery-public-data.thelook_ecommerce&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/public-data"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;public dataset in BigQuery&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. For demo purposes, we split this static dataset into training and inference sets. In a real-life scenario, inference would be performed on new, incoming data. This dataset provides authentic operational complexity:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Geographical data:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Latitude and longitude for both customer addresses (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;users&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;) and distribution centers (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;distribution_centers&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Temporal data:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Granular order lifecycle timestamps (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;created_at&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;shipped_at&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;delivered_at&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Order attributes:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Product categories, pricing, and fulfillment status (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;orders&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;order_items&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By combining this dataset with BigQuery, &lt;/span&gt;&lt;a href="https://cloud.google.com/products/managed-service-for-apache-spark"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Managed Service for Apache Spark&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; serverless, &lt;/span&gt;&lt;a href="https://cloud.google.com/products/gemini-enterprise-agent-platform"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise Agent Platform&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and &lt;/span&gt;&lt;a href="https://www.getdbt.com/product/what-is-dbt" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;dbt&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, we will demonstrate how to build an automated, self-healing MLOps loop that handles training, daily batch inference, and model drift evaluation.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;The agentic workflow: From prompt to pipeline in minutes&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With the extension configured, we can bypass boilerplate Python for DAG authoring entirely. Inside VS Code, we opened the Data Agent Kit chat and provided a single natural language prompt to define our continuous MLOps feedback loop:&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="font-style: italic; vertical-align: baseline;"&gt;Note:&lt;/strong&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt; The detailed prompt was crafted with repeatability in mind specifically for this blog post. In real-life scenarios, you can achieve the same result in a more conversational way, pipeline by pipeline. The complete prompt and all generated files are available in the &lt;/span&gt;&lt;a href="https://github.com/GoogleCloudPlatform/orchestration-pipelines/tree/main/examples/blogpost-2026" rel="noopener" target="_blank"&gt;&lt;span style="font-style: italic; text-decoration: underline; vertical-align: baseline;"&gt;Orchestration-pipelines GitHub repository&lt;/span&gt;&lt;/a&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Note:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; While frontier models equipped with the Orchestration Pipelines skill can often scaffold complete workflows in a single step, LLM responses naturally vary based on model versions, workspace context, and token depth. If a specific parameter, dataset path, or dependency is omitted in the initial pass, simply provide a short follow-up prompt.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Within minutes, the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/data-engineering-agent-pipelines"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Data Agent Kit&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; generated the underlying PySpark scripts, dbt configurations, and the three declarative YAML pipelines.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Please find below the generated YAML pipelines and a visual diagram of them. This pipeline is a simplified example designed to showcase Orchestration Pipelines capabilities. In practice, recommended production MLOps setups will vary depending on your specific use cases and operational needs.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_XMQ1O65.max-1000x1000.jpg"
        
          alt="1"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Pipeline 1: The training engine&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;This pipeline serves as our heavy-compute engine. The agent generated a YAML definition that first queries BigQuery to extract historical completed orders. It then dynamically provisions a Managed Spark serverless cluster to calculate geographical distances and train a model for production use. Finally, it pushes the trained model to Gemini Enterprise Agent Platform Model Registry.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;modelVersion: &amp;quot;1.0&amp;quot;\r\npipelineId: &amp;quot;training-pipeline&amp;quot;\r\nrunner: airflow\r\nowner: &amp;quot;mlops&amp;quot;\r\ntags:\r\n  - &amp;quot;job:datacloud:antigravity&amp;quot;\r\ndefaults:\r\n  projectId: &amp;quot;your-project-id&amp;quot;\r\n  location: &amp;quot;us-central1&amp;quot;\r\n  executionConfig:\r\n    retries: 0\r\n\r\nactions:\r\n  - sql:\r\n      name: &amp;quot;extract_training_data&amp;quot;\r\n      engine:\r\n        bigquery:\r\n          location: &amp;quot;US&amp;quot;\r\n          destinationTable: &amp;quot;your-project-id.mlops.training_dataset&amp;quot;\r\n      query:\r\n        path: &amp;quot;blogpostdemo/training_query.sql&amp;quot;\r\n\r\n  - pyspark:\r\n      name: &amp;quot;train_model_dataproc&amp;quot;\r\n      dependsOn:\r\n        - &amp;quot;extract_training_data&amp;quot;\r\n      engine:\r\n        dataprocServerless:\r\n          location: &amp;quot;us-central1&amp;quot;\r\n          resourceProfile:\r\n            inline:\r\n              runtimeConfig:\r\n                version: &amp;quot;2.3&amp;quot;\r\n                properties:\r\n                  &amp;quot;spark.dataproc.driverEnv.PYTHONPATH&amp;quot;: &amp;quot;./libs/lib/python3.11/site-packages&amp;quot;\r\n                  &amp;quot;spark.executorEnv.PYTHONPATH&amp;quot;: &amp;quot;./libs/lib/python3.11/site-packages&amp;quot;\r\n      mainFilePath: &amp;quot;blogpostdemo/train_model.py&amp;quot;\r\n      environment:\r\n        requirements:\r\n          inline:\r\n            list:\r\n              - &amp;quot;tensorflow==2.14.1&amp;quot;\r\n              - &amp;quot;numpy&amp;lt;2.0.0&amp;quot;\r\n              - &amp;quot;protobuf&amp;lt;5.0.0dev&amp;quot;\r\n              - &amp;quot;google-cloud-storage&amp;quot;\r\n\r\n  - ai:\r\n      name: &amp;quot;upload_model_vertex&amp;quot;\r\n      dependsOn:\r\n        - &amp;quot;train_model_dataproc&amp;quot;\r\n      agentPlatform:\r\n        projectId: &amp;quot;your-project-id&amp;quot;\r\n        location: &amp;quot;us-central1&amp;quot;\r\n        modelUpload:\r\n          modelName: &amp;quot;transit_days_predictor&amp;quot;\r\n          modelArtifactUri: &amp;quot;gs://your-bucket-name/models/tf_transit_days_model&amp;quot;\r\n          servingContainerImageUri: &amp;quot;us-docker.pkg.dev/vertex-ai/prediction/tf2-cpu.2-14:latest&amp;quot;&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca34f5f850&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Pipeline 2: Daily inference&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;For our daily operational workflow, this lightweight pipeline applies the trained model to all currently in-transit orders. It queries the dataset via BigQuery job, executes inference job via Gemini Enterprise Agent Platform, and writes the results back to a BigQuery table to flag potential SLA breaches for the customer support team.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;modelVersion: &amp;quot;1.0&amp;quot;\r\npipelineId: &amp;quot;inference-pipeline&amp;quot;\r\nrunner: airflow\r\nowner: &amp;quot;mlops&amp;quot;\r\ntags:\r\n  - &amp;quot;job:datacloud:antigravity&amp;quot;\r\ndefaults:\r\n  projectId: &amp;quot;your-project-id&amp;quot;\r\n  location: &amp;quot;us-central1&amp;quot;\r\n  executionConfig:\r\n    retries: 0\r\n\r\nactions:\r\n  - sql:\r\n      name: &amp;quot;extract_inference_data&amp;quot;\r\n      engine:\r\n        bigquery:\r\n          location: &amp;quot;US&amp;quot;\r\n          destinationTable: &amp;quot;your-project-id.mlops.inference_dataset&amp;quot;\r\n      query:\r\n        path: &amp;quot;blogpostdemo/inference_query.sql&amp;quot;\r\n\r\n  - ai:\r\n      name: &amp;quot;run_vertex_batch_prediction&amp;quot;\r\n      dependsOn:\r\n        - &amp;quot;extract_inference_data&amp;quot;\r\n      agentPlatform:\r\n        projectId: &amp;quot;your-project-id&amp;quot;\r\n        location: &amp;quot;us-central1&amp;quot;\r\n        batchInference:\r\n          jobDisplayName: &amp;quot;inference_job&amp;quot;\r\n          modelName: &amp;quot;projects/your-project-id/locations/us-central1/models/your-model-id&amp;quot;\r\n          bigquerySource: &amp;quot;bq://your-project-id.mlops.inference_dataset&amp;quot;\r\n          bigqueryDestinationPrefix: &amp;quot;bq://your-project-id.mlops&amp;quot;&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca34ed3e10&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Pipeline 3: Automated evaluation and branching&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;The daily evaluation pipeline acts as our automated quality gate. It triggers dbt models to join our predictions with actual delivery timestamps, calculating absolute errors and SLA breaches.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Using built-in logic, the pipeline automatically evaluates these metrics. If the model’s error rate exceeds our acceptable threshold, it conditionally triggers the ‘training-pipeline’ to generate a fresh model.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;modelVersion: &amp;quot;1.0&amp;quot;\r\npipelineId: &amp;quot;evaluation-pipeline&amp;quot;\r\nrunner: airflow\r\nowner: &amp;quot;mlops&amp;quot;\r\ntags:\r\n  - &amp;quot;job:datacloud:antigravity&amp;quot;\r\ndefaults:\r\n  projectId: &amp;quot;your-project-id&amp;quot;\r\n  location: &amp;quot;us-central1&amp;quot;\r\n  executionConfig:\r\n    retries: 0\r\n\r\nactions:\r\n  - pipeline:\r\n      name: &amp;quot;run_dbt_models&amp;quot;\r\n      framework:\r\n        dbt:\r\n          airflowWorker:\r\n            projectDirectoryPath: &amp;quot;blogpostdemo/dbt_project&amp;quot;\r\n\r\n  - python:\r\n      name: &amp;quot;check_retraining_condition&amp;quot;\r\n      dependsOn:\r\n        - &amp;quot;run_dbt_models&amp;quot;\r\n      mainFilePath: &amp;quot;blogpostdemo/evaluate_drift.py&amp;quot;\r\n      pythonCallable: &amp;quot;check_drift&amp;quot;\r\n      engine:\r\n        local: {}\r\n\r\n  - orchestrationPipeline:\r\n      name: &amp;quot;trigger_retraining_pipeline&amp;quot;\r\n      dependsOn:\r\n        - &amp;quot;check_retraining_condition&amp;quot;\r\n      pipelineId: &amp;quot;training-pipeline&amp;quot;\r\n      bundleId: &amp;quot;my-first-bundle&amp;quot;\r\n      waitForCompletion: false&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca34ed1690&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Automated deployment to Managed Service for Apache Airflow&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Authoring pipeline logic is only half the battle; deploying it securely and reliably to production is where data teams historically lose valuable time.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/orchestration-pipelines"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Orchestration Pipelines&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, deployment is streamlined through standard CI/CD practices. Rather than manually writing deployment scripts or configuring complex environment boundaries, the Data Agent Kit automatically generates the necessary continuous integration workflows (such as GitHub Actions) for your workspace.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This means you can simply click commit, and the framework will seamlessly package and deploy your Orchestration Pipeline bundle directly to your Managed Airflow environment.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For a comprehensive guide on integrating these automated workflows into your existing CI/CD pipelines, review the official guide:&lt;/span&gt; &lt;a href="https://docs.cloud.google.com/orchestration-pipelines/deploy-orchestration-pipelines#deploy-run"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Deploying Orchestration Pipelines&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Day-two operations: Monitoring and agentic troubleshooting&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Maintaining these pipelines is just as intuitive as building them. By bringing the orchestration control plane directly into your IDE, the Data Agent Kit provides real-time monitoring of your Managed Airflow runs without requiring you to constantly context-switch between browser tabs.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_6wcKwIA.max-1000x1000.png"
        
          alt="2"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="1tm1o"&gt;The Data Agent Kit provides real-time monitoring of your Managed Airflow runs directly within your IDE.&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/3_qAVdnpZ.max-1000x1000.png"
        
          alt="3"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="1tm1o"&gt;The Data Agent Kit visualises the created pipeline.&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Inevitably, infrastructure or data issues occur—perhaps a Managed Spark cluster hits an out-of-memory exception due to a seasonal data spike, or a BigQuery quota is reached. Resolving these issues no longer requires digging through thousands of lines of raw execution logs.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If a pipeline fails, the Data Agent Kit provides out-of-the-box agentic troubleshooting. With the click of a "Troubleshoot" button in your IDE, the Data Engineering Agent analyzes the failure context. It can accurately distinguish between infrastructure quota limits and code-level bugs, instantly providing a root-cause summary and suggesting an inline fix (such as scaling up the compute template).&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/4_tCCNsgc.max-1000x1000.png"
        
          alt="4"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="1tm1o"&gt;Agentic troubleshooting instantly diagnoses pipeline failures, identifies infrastructure bottlenecks, and suggests inline fixes.&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Summary: Accelerating time to value&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Building a resilient MLOps architecture — extracting historical data, executing dbt transformations, provisioning Managed Spark ML compute, integrating Gemini Enterprise Agent Platform for model registry and inference, and configuring cross-DAG conditional triggers — traditionally takes platform engineering teams weeks of writing complex Python Operator logic.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With Orchestration Pipelines and the Data Agent Kit, this entire lifecycle was authored, deployed, and easily maintained in a matter of minutes. By replacing boilerplate infrastructure code with a declarative, agent-ready standard, we are ensuring your data organization spends less time orchestrating pipelines and more time delivering tangible business value.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Get Started Today:&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Review the&lt;/span&gt;&lt;a href="https://docs.cloud.google.com/orchestration-pipelines/overview"&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Orchestration Pipelines documentation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Install the&lt;/span&gt; &lt;a href="https://docs.cloud.google.com/data-agent-kit"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Data Agent Kit&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; in your preferred IDE or CLI and configure your workspace.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Learn more about the broader ecosystem in our recent blog post: &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/data-agent-kit-brings-data-skills-and-tools-to-your-ide-or-cli"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Data Agent Kit brings data skills and tools to your IDE or CLI&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Explore reference architectures in the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/data-cloud-extension/vs-code/train-models"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Data Agent Kit documentation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Mon, 31 Aug 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/build-data-pipelines-in-less-time-with-data-agent-kit/</guid><category>Data Analytics</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>From weeks to minutes: The new agentic era of data pipelines</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/build-data-pipelines-in-less-time-with-data-agent-kit/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Rafal Biegacz</name><title>Senior Software Engineering Manager</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Alexandre Moueddene</name><title>Software Engineer</title><department></department><company></company></author></item><item><title>Using OKF with Knowledge Catalog to serve context for agents</title><link>https://cloud.google.com/blog/products/data-analytics/scale-okf-bundles-across-an-organization-with-knowledge-catalog/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We continue to iterate on the &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/how-the-open-knowledge-format-can-improve-data-sharing"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Open Knowledge Format&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (OKF), an open specification that formalizes the &lt;/span&gt;&lt;a href="https://gist.github.com/karpathy/442a6bf555914893e9891c11519de94f" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;LLM-wiki pattern&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; into a portable, interoperable format. But a big question remains: How can you share and govern access to an OKF bundle across an organization?&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;OKF v0.1 established a portable format for the context agents need: markdown files with YAML frontmatter, one required field, and five conventions. Then, &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/okf-v0-2-adds-trust-signals"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;OKF v0.2&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; added the trust signals (provenance, verification, freshness, attestation) that a machine-authored bundle requires to be relied on, allowing a team to publish a trustworthy bundle for its own agents. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;However, what OKF does not answer is how teams share their bundles across an organization. A git repo per bundle is portable, but it is not searchable alongside the data it describes, it cannot be secured and governed using the same organizational identity and compliance policies, and it does not sit next to the technical metadata (schemas, lineage, ownership) that data teams already work in. Every downstream agent must know where each bundle resides, and that does not scale beyond a small number of bundles.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To scale an OKF bundle across an organization, you can use &lt;/span&gt;&lt;a href="https://cloud.google.com/products/knowledge-catalog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Knowledge Catalog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, Google Cloud's context engine for agents. By mapping the bundle onto &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/dataplex/docs/catalog-overview#terminology"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Knowledge Catalog's existing types&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, every concept becomes discoverable, governed, and reachable by any agent already reading from the catalog.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Knowledge Catalog is the context engine for agents&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Every agent that queries Knowledge Catalog reads from one governed index over what the organization already has in BigQuery, Cloud Storage, operational databases, and applications. Each entry carries schema, lineage, ownership, and tags, and can be extended with typed aspects that add domain-specific fields. The same catalog exposes search and cross-project lookup to retrieve optimized context for each agentic query. The context retrieval is secure and governed by IAM controls, so agents can only see the entries they have access to based on IAM identity. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Publishing an OKF bundle into Knowledge Catalog takes a one-time setup and a single push. Both use the OKF &lt;/span&gt;&lt;a href="https://github.com/GoogleCloudPlatform/knowledge-catalog/tree/main/toolbox/mdcode/demo/okf" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;sample code&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; in the Knowledge Catalog repository, whose wrappers call &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;gcloud dataplex&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; for setup and delegate push to &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;kcmd&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; (the Metadata-as-Code CLI in the same repository).&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The setup registers three Knowledge Catalog resources: an EntryGroup to hold the bundle, an EntryType named &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;okf-bundle&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; for its concepts, and an AspectType named &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;okf&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; that carries the OKF signal fields (from the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;okf-aspect.json&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; schema in the &lt;/span&gt;&lt;a href="https://github.com/GoogleCloudPlatform/knowledge-catalog/tree/main/toolbox/mdcode/demo/okf" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;sample code&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;). The push then creates one &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;okf-bundle&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; Entry per concept, each with two Aspects: an &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;overview&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; Aspect for the markdown body, and an &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;okf&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; Aspect for the structured signals. Display name, description, and tags live on the Entry itself. The bundle's &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;index.md&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; navigation files and its root &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;log.md&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; are also published as Entries: index files carry only the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;overview&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; Aspect (no OKF frontmatter), and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;log.md&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; carries both Aspects with &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;okf_type: Log&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Everything Knowledge Catalog already does for technical metadata (search, IAM, lineage, cross-project discovery) applies equally to OKF bundles, alongside the data they describe.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;The &lt;/strong&gt;&lt;strong style="vertical-align: baseline;"&gt;okf&lt;/strong&gt;&lt;strong style="vertical-align: baseline;"&gt; AspectType&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The &lt;/span&gt;&lt;a href="https://github.com/GoogleCloudPlatform/knowledge-catalog/blob/main/toolbox/mdcode/demo/okf/okf-aspect.json" rel="noopener" target="_blank"&gt;&lt;code style="text-decoration: underline; vertical-align: baseline;"&gt;okf-aspect.json&lt;/code&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; schema in the sample code defines the AspectType. It carries 13 fields covering the full &lt;/span&gt;&lt;a href="https://github.com/GoogleCloudPlatform/open-knowledge-format/blob/main/SPEC.md" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;OKF v0.2 spec&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;/p&gt;
&lt;div align="center"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;&lt;table&gt;&lt;colgroup&gt;&lt;col/&gt;&lt;col/&gt;&lt;col/&gt;&lt;col/&gt;&lt;/colgroup&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th scope="col" style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;#&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Field&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Type&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Purpose&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;1&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;code style="vertical-align: baseline;"&gt;okf_type&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;string&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The OKF document type (freeform, e.g. &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;BigQuery Table&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;Metric&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;Attested Computation&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;).&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;2&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;code style="vertical-align: baseline;"&gt;generated&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;record &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;{by, at}&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Actor and timestamp for the last meaningful change.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;3&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;code style="vertical-align: baseline;"&gt;sources&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;array of &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;{id, resource, title, author, usage_count, last_modified}&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Materials the concept derives from, with credibility signals.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;4&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;code style="vertical-align: baseline;"&gt;verified&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;array of &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;{by, at}&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Verification events. A &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;human:&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; actor marks the highest trust tier.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;5&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;code style="vertical-align: baseline;"&gt;status&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;string&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Lifecycle state: &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;draft&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;stable&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, or &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;deprecated&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;6&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;code style="vertical-align: baseline;"&gt;stale_after&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;datetime&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Absolute point in time (RFC3339 with an explicit offset) on or after which the content is stale.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;7&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;code style="vertical-align: baseline;"&gt;usage_window&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;record &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;{from, to}&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Period the source usage counts were measured over.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;8&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;code style="vertical-align: baseline;"&gt;runtime&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;string&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;How an Attested Computation runs (e.g., &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;bigquery&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;).&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;9&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;code style="vertical-align: baseline;"&gt;parameters&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;array of &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;{name, type, required}&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Typed named holes a caller may fill. The only surface a caller may vary.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;10&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;code style="vertical-align: baseline;"&gt;computation&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;string&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Path to a file holding the computation body.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;11&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;code style="vertical-align: baseline;"&gt;executor&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;record &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;{resource, receipt[]}&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;How the computation runs and what evidence it must return.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;12&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;code style="vertical-align: baseline;"&gt;attester&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;record &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;{resource}&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Deterministic code that takes a receipt and returns a verdict.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;13&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;code style="vertical-align: baseline;"&gt;extra&lt;/code&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;string&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Producer-defined frontmatter the template does not model, as JSON &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;[path, value]&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; pairs. Keeps the round-trip lossless.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Every field is annotated with a display name, a description, and a mandatory index. Any top-level scalar field in the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;okf&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; Aspect (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;okf_type&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;status&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;stale_after&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;runtime&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;computation&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;extra&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;) can drive Knowledge Catalog search predicates directly, so &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;aspect:acme-analytics.us-central1.okf.okf_type=Metric&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; returns every OKF Metric in scope. Scalar subfields of record fields (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;generated.by&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;usage_window.from&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;executor.resource&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;attester.resource&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;) also drive predicates. The array fields (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;sources&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;verified&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;parameters&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;) are not server-side searchable on their subfields; agents narrow on them client-side after &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;entries.get&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; with &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;view=ALL&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;. One caveat for search predicates on &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;datetime&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;-typed fields (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;stale_after&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;generated.at&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;usage_window.from&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;/&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;.to&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;), use a bare date (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;stale_after=2026-12-31&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;) or a range comparison (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;stale_after&amp;gt;2026-01-01&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;), not the full RFC3339 timestamp.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Pushing a bundle&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;code style="vertical-align: baseline;"&gt;kcmd push&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; reads an OKF bundle from git and writes each concept as an Entry in the target Knowledge Catalog EntryGroup. &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;index.md&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; files become Entries too, and each concept is parented to the index above it, so the bundle's directory structure survives as a browsable hierarchy.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;code style="vertical-align: baseline;"&gt;kcmd&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; expects a bundle in the Documents Layout: markdown files under a &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;catalog/&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; subdirectory, and a &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;catalog.yaml&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; at the bundle root that lists the snapshot's entry and aspect types. The &lt;/span&gt;&lt;a href="https://github.com/GoogleCloudPlatform/knowledge-catalog/tree/main/toolbox/mdcode/demo/okf" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;sample code&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;'s &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;setup.ts&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; generates &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;catalog.yaml&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; from its &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;--entry-group&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; flag (default &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;okf_demo&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;), so a reader wiring the sample to a new bundle passes the flag rather than editing &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;catalog.yaml&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; by hand.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Here is an end-to-end workflow for the &lt;/span&gt;&lt;a href="https://github.com/GoogleCloudPlatform/open-knowledge-format/tree/main/bundles/acme_retail" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Acme Retail bundle&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; that we introduced in the &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/okf-v0-2-adds-trust-signals?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;OKF v0.2 blog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;# One-time setup (if required): install bun, clone the repo, build kcmd, configure gcloud\r\ncurl -fsSL https://bun.sh/install | bash\r\nexport BUN_INSTALL=&amp;quot;$HOME/.bun&amp;quot; &amp;amp;&amp;amp; export PATH=&amp;quot;$BUN_INSTALL/bin:$PATH&amp;quot;\r\ngit clone https://github.com/GoogleCloudPlatform/knowledge-catalog\r\ncd knowledge-catalog/toolbox/mdcode &amp;amp;&amp;amp; npm install &amp;amp;&amp;amp; npm run build\r\n\r\n# Authenticate, set project and enable dataplex apis\r\ngcloud auth login\r\ngcloud config set project &amp;lt;your-project&amp;gt;\r\ngcloud config set compute/region &amp;lt;your-location&amp;gt;\r\ngcloud services enable dataplex.googleapis.com\r\ngcloud auth application-default login\r\n\r\n# Push the Acme Retail bundle\r\ncd demo/okf\r\nbun run setup.ts   # creates the EG (default \&amp;#x27;okf_demo\&amp;#x27;)\r\nbun run push.ts    # pushes okf/bundles/acme_retail into the EG setup created&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca34ed2010&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To pick a different EntryGroup name or push a different bundle, pass &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;--entry-group your-name&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; to &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;setup.ts&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;--bundle path/to/your/bundle&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; to &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;push.ts&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;. For example: &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;bun run setup.ts --entry-group acme-bundle&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; followed by &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;bun run push.ts&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;. This regenerates the manifest, so subsequent &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;push&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;pull&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;cleanup&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; all target the new EG; delete earlier EGs manually with &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;gcloud dataplex entry-groups delete &amp;lt;name&amp;gt; --project &amp;lt;your-project&amp;gt; --location &amp;lt;your-location&amp;gt;&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The &lt;/span&gt;&lt;a href="https://github.com/GoogleCloudPlatform/open-knowledge-format/tree/main/bundles/acme_retail" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Acme Retail bundle&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is a synthetic OKF bundle for a US retailer's BigQuery estate. It contains nine leaf concepts across six directories (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;attesters&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;tables&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;metrics&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;computations&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;policies&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;skills&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;), each with its own &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;index.md&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, plus a bundle root with its own &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;index.md&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;log.md&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;. That's 17 pushed Entries in total; Dataplex auto-creates one &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;&amp;lt;eg&amp;gt;_entry&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; alongside, so &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;gcloud dataplex entries list&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; returns 18 rows.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;After the push completes:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Every concept markdown file is a Knowledge Catalog Entry, discoverable by search across the whole project or organization, depending on IAM configuration.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;The &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;revenue-ytd&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; Attested Computation appears in the console with its sanctioned SQL, its executor, its attester, its verification history, and the full concept body.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;An analyst searching Knowledge Catalog for "revenue" finds Acme Retail's business definition alongside the BigQuery table it computes from, both under one permission model.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;A downstream agent that already calls LookupContext for BigQuery table Entries retrieves the bundle's context by adding the OKF entry names to its &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;resources&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; list.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Further, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;metrics/revenue.md&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; becomes an Entry with two Aspects. The full &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;entries.get&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; response (with &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;view=ALL&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;) looks like:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;{\r\n  &amp;quot;name&amp;quot;: &amp;quot;projects/acme-analytics/locations/us-central1/entryGroups/acme-retail/entries/metrics/revenue&amp;quot;,\r\n  &amp;quot;entryType&amp;quot;: &amp;quot;projects/acme-analytics/locations/us-central1/entryTypes/okf-bundle&amp;quot;,\r\n  &amp;quot;createTime&amp;quot;: &amp;quot;2026-08-15T00:48:39.123456Z&amp;quot;,\r\n  &amp;quot;updateTime&amp;quot;: &amp;quot;2026-08-15T00:48:57.234567Z&amp;quot;,\r\n  &amp;quot;parentEntry&amp;quot;: &amp;quot;projects/acme-analytics/locations/us-central1/entryGroups/acme-retail/entries/metrics/index&amp;quot;,\r\n  &amp;quot;entrySource&amp;quot;: {\r\n    &amp;quot;displayName&amp;quot;: &amp;quot;Revenue&amp;quot;,\r\n    &amp;quot;description&amp;quot;: &amp;quot;Recognized revenue for a period, per Acme\&amp;#x27;s FY2026 revenue-recognition policy. Backed by an Attested Computation.&amp;quot;,\r\n    &amp;quot;labels&amp;quot;: {\r\n      &amp;quot;finance&amp;quot;: &amp;quot;true&amp;quot;,\r\n      &amp;quot;revenue&amp;quot;: &amp;quot;true&amp;quot;,\r\n      &amp;quot;headline-metric&amp;quot;: &amp;quot;true&amp;quot;\r\n    },\r\n    &amp;quot;location&amp;quot;: &amp;quot;us-central1&amp;quot;\r\n  },\r\n  &amp;quot;aspects&amp;quot;: {\r\n    &amp;quot;dataplex-types.global.overview&amp;quot;: {\r\n      &amp;quot;aspectType&amp;quot;: &amp;quot;projects/dataplex-types/locations/global/aspectTypes/overview&amp;quot;,\r\n      &amp;quot;createTime&amp;quot;: &amp;quot;2026-08-15T00:48:57.111111Z&amp;quot;,\r\n      &amp;quot;updateTime&amp;quot;: &amp;quot;2026-08-15T00:48:57.111111Z&amp;quot;,\r\n      &amp;quot;aspectSource&amp;quot;: {},\r\n      &amp;quot;data&amp;quot;: {\r\n        &amp;quot;content&amp;quot;: &amp;quot;# Definition\\n\\nRevenue for a fiscal year is the sum of `net_amount` over orders that (a) reached `order_status = \&amp;#x27;delivered\&amp;#x27;`, (b) completed the 30-day return window, and (c) fall in the fiscal year by `order_ts`. Multi-currency orders are converted to USD at the `order_ts` daily reference rate. [^revenue-policy]\\n\\nThe sanctioned computation is [`computations/revenue-ytd.md`](../computations/revenue-ytd.md). Consumers MUST run and attest that computation rather than composing their own SUM. The attester rejects any receipt whose executed SQL does not match the sanctioned form.\\n\\n# Reporting cuts\\n\\n- **By fiscal year:** the sanctioned computation takes `year` as its sole parameter.\\n- **By channel or category:** these are approved narrations, not new metrics. Join the receipt\&amp;#x27;s row-level result to `orders.channel` or to `order_lines` × `products.category` client-side. Do NOT rewrite the sanctioned SQL.\\n\\n# Trust and freshness\\n\\n- **Verified:** VP Finance sign-off on 2026-07-01, against the FY2026 policy.\\n- **Stale after 2026-12-31:** Finance re-issues the revenue recognition policy each January. Consumers of this concept after 2027-01-01 MUST re-verify the definition against the new policy before serving.\\n\\n[^revenue-policy]: Revenue Recognition Policy (FY2026)&amp;quot;,\r\n        &amp;quot;contentType&amp;quot;: &amp;quot;MARKDOWN&amp;quot;\r\n      }\r\n    },\r\n    &amp;quot;acme-analytics.us-central1.okf&amp;quot;: {\r\n      &amp;quot;aspectType&amp;quot;: &amp;quot;projects/acme-analytics/locations/us-central1/aspectTypes/okf&amp;quot;,\r\n      &amp;quot;createTime&amp;quot;: &amp;quot;2026-08-15T00:48:57.222222Z&amp;quot;,\r\n      &amp;quot;updateTime&amp;quot;: &amp;quot;2026-08-15T00:48:57.222222Z&amp;quot;,\r\n      &amp;quot;aspectSource&amp;quot;: {},\r\n      &amp;quot;data&amp;quot;: {\r\n        &amp;quot;okf_type&amp;quot;: &amp;quot;Metric&amp;quot;,\r\n        &amp;quot;generated&amp;quot;: { &amp;quot;by&amp;quot;: &amp;quot;reference_agent/gemini-2.5-pro&amp;quot;, &amp;quot;at&amp;quot;: &amp;quot;2026-06-30T14:00:00Z&amp;quot; },\r\n        &amp;quot;verified&amp;quot;: [ { &amp;quot;by&amp;quot;: &amp;quot;human:jsmith@acme&amp;quot;, &amp;quot;at&amp;quot;: &amp;quot;2026-07-01T09:00:00Z&amp;quot; } ],\r\n        &amp;quot;status&amp;quot;: &amp;quot;stable&amp;quot;,\r\n        &amp;quot;stale_after&amp;quot;: &amp;quot;2026-12-31T00:00:00Z&amp;quot;,\r\n        &amp;quot;sources&amp;quot;: [\r\n          {\r\n            &amp;quot;id&amp;quot;: &amp;quot;revenue-policy&amp;quot;,\r\n            &amp;quot;resource&amp;quot;: &amp;quot;policies/revenue-recognition.md&amp;quot;,\r\n            &amp;quot;title&amp;quot;: &amp;quot;Revenue Recognition Policy (FY2026)&amp;quot;,\r\n            &amp;quot;author&amp;quot;: &amp;quot;human:jsmith@acme&amp;quot;,\r\n            &amp;quot;last_modified&amp;quot;: &amp;quot;2026-06-15T00:00:00Z&amp;quot;\r\n          }\r\n        ]\r\n      }\r\n    }\r\n  }\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca3439dc10&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;overview&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; Aspect holds the full body of &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;revenue.md&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;. The &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;okf&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; Aspect carries the structured signal fields, so agents get provenance, source, and OKF type in a form they can filter on directly instead of parsing markdown. Server-side searchEntries filters on the top-level scalar fields and on the scalar subfields of record fields; agents narrow further on the array-element subfields client-side after entries.get. (Aspects and EntryTypes are keyed by project number in real API responses and search predicates; the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;acme-analytics&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; project ID is shown throughout for readability.)&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;What pushing your OKF to Knowledge Catalog enables&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Once the bundle is in Knowledge Catalog, it provides two capabilities to any agent that reads from the catalog:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Discoverability across the organization.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Agents find bundle concepts through the same searchEntries and LookupContext APIs they already use for cataloged data, so an OKF bundle appears alongside BigQuery tables and other resources in every query it matches.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Governance.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Bundle Entries inherit IAM from the EntryGroup, so a single agent call returns exactly what the caller is permitted to read, with no parallel permission model to maintain.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Discoverability across the organization&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;OKF bundle Entries appear in searchEntries results alongside BigQuery tables and other cataloged resources, so an agent already querying the catalog picks up new bundles automatically. To retrieve a concept's body, trust signals, or linked concepts from a match, the agent moves to LookupContext and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;entries.get&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;A LookupContext call looks like this:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;POST https://dataplex.googleapis.com/v1/projects/acme-analytics/locations/us-central1:lookupContext\r\n{\r\n  &amp;quot;resources&amp;quot;: [\r\n    &amp;quot;projects/acme-analytics/locations/us-central1/entryGroups/acme-retail/entries/metrics/revenue&amp;quot;\r\n  ],\r\n  &amp;quot;options&amp;quot;: { &amp;quot;format&amp;quot;: &amp;quot;yaml&amp;quot;, &amp;quot;context_budget&amp;quot;: &amp;quot;8000&amp;quot; }\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca3439da90&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The response is a single &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;context&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; field containing a pre-formatted YAML block. The block carries the entry's &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;catalogEntry&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, its type, its description, its tags as labels, and its &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;overview&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;: the full markdown body of the concept, including its trust and freshness section. LookupContext does not render custom Aspects, so an agent that needs the structured OKF signal fields (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;okf_type&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;generated&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;sources&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, and the other ten) reads them with &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;entries.get&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;view=ALL&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; alongside the LookupContext call.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;There is no repository clone, no manual Aspect merging, and no re-parse of frontmatter. The agent uses the same API call any Knowledge Catalog client already makes.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;An agent traversing an OKF bundle typically follows a three-step flow. An agent that already knows the specific Entry names it needs skips step 1. An agent that already knows the target EntryGroup and wants to enumerate the bundle exhaustively substitutes &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;entryGroups.entries.list&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; for step 1.&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;searchEntries returns candidate Entry names and descriptions. Its &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;scope&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; accepts a project or organization; narrowing within that scope happens through query terms, including aspect predicates like &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;aspect:acme-analytics.us-central1.okf.okf_type=Metric&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;LookupContext on the top few Entry names (up to ten per call) returns the full concept body as pre-formatted YAML; &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;context_budget&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; caps the response size.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;code style="vertical-align: baseline;"&gt;entries.get&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; with &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;view=ALL&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; on any Entry returns its structured OKF signals (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;okf_type&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;generated&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;sources&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, and the other ten) directly, which the agent can then filter or attest on.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When a concept's &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;sources[]&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; references another concept by path, the agent calls LookupContext on that Entry name to walk the reference.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The full response for the Revenue Entry:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;quot;resources:\r\n -\r\n  catalogEntry: projects/acme-analytics/locations/us-central1/entryGroups/acme-retail/entries/metrics/revenue\r\n  type: OKF Document\r\n  description: Recognized revenue for a period, per Acme&amp;#x27;s FY2026 revenue-recognition\r\n    policy. Backed by an Attested Computation.\r\n  overview: |-\r\n    # Definition\r\n\r\n    Revenue for a fiscal year is the sum of `net_amount` over orders that (a) reached `order_status = &amp;#x27;delivered&amp;#x27;`, (b) completed the 30-day return window, and (c) fall in the fiscal year by `order_ts`. Multi-currency orders are converted to USD at the `order_ts` daily reference rate. [^revenue-policy]\r\n\r\n    The sanctioned computation is [`computations/revenue-ytd.md`](../computations/revenue-ytd.md). Consumers MUST run and attest that computation rather than composing their own SUM. The attester rejects any receipt whose executed SQL does not match the sanctioned form.\r\n\r\n    # Reporting cuts\r\n\r\n    - **By fiscal year:** the sanctioned computation takes `year` as its sole parameter.\r\n    - **By channel or category:** these are approved narrations, not new metrics. Join the receipt&amp;#x27;s row-level result to `orders.channel` or to `order_lines` × `products.category` client-side. Do NOT rewrite the sanctioned SQL.\r\n\r\n    # Trust and freshness\r\n\r\n    - **Verified:** VP Finance sign-off on 2026-07-01, against the FY2026 policy.\r\n    - **Stale after 2026-12-31:** Finance re-issues the revenue recognition policy each January. Consumers of this concept after 2027-01-01 MUST re-verify the definition against the new policy before serving.\r\n\r\n    [^revenue-policy]: Revenue Recognition Policy (FY2026)\r\n  labels:\r\n    finance: &amp;#x27;true&amp;#x27;\r\n    revenue: &amp;#x27;true&amp;#x27;\r\n    headline-metric: &amp;#x27;true&amp;#x27;&amp;quot;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7fca3439fa10&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Governance&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Permissions on the EntryGroup use standard Knowledge Catalog IAM. An agent that names both a bundle concept and the BigQuery table it grounds against in one call receives both, each subject to its own existing access control list (ACL), so the response carries only what the caller is already permitted to read. There is no parallel permission model to maintain.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Reading agents use &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;roles/dataplex.catalogViewer&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, which grants the read paths: &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;entries.get&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, LookupContext, and searchEntries. The identity that runs &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;kcmd push&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; uses &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;roles/dataplex.catalogEditor&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, which grants the write paths: &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;entries.create&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;entries.patch&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;. One EntryGroup per bundle-owning team is the multi-team pattern, and IAM on the EntryGroup cascades to its Entries.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;LookupContext resolves the entry names it is given, up to ten per call, within a single location. It does not follow links out of a concept's body, so an agent that wants a referenced concept must name it explicitly. Place the bundle's EntryGroup in the same location as the data it describes to fetch both in one call.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Lifecycle&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;code style="vertical-align: baseline;"&gt;kcmd push&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; is an idempotent upsert. Re-running is safe (no duplicates, no error), but every push writes every Entry. Concept deletes require an explicit &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;kcmd delete&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; on the Entry, or &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;cleanup.ts&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; to remove the whole EntryGroup at once; &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;cleanup.ts&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; deletes only the EntryGroup and its Entries, so the shared &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;okf&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; AspectType and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;okf-bundle&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; EntryType stay in place for other bundles that reference them. For continuous ingestion in production, wire a CI job to &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;kcmd push&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; on every commit to the bundle repository, using a service-account credential with &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;roles/dataplex.catalogEditor&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; on the target EntryGroup.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Getting started&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;OKF defines what a trustworthy bundle looks like. Knowledge Catalog makes it reachable across the organization. To get started, check out the following resources:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Read the &lt;/span&gt;&lt;a href="https://github.com/GoogleCloudPlatform/open-knowledge-format/blob/main/SPEC.md" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;OKF v0.2 spec&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and browse the &lt;/span&gt;&lt;a href="https://github.com/GoogleCloudPlatform/open-knowledge-format/tree/main/bundles/acme_retail" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Acme Retail bundle&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Author a small bundle for one domain your team owns.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Sync it into your Knowledge Catalog project using the &lt;/span&gt;&lt;a href="https://github.com/GoogleCloudPlatform/knowledge-catalog/tree/main/toolbox/mdcode/demo/okf" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;sample code&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;'s &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;setup.ts&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; (which registers the resources) and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;push.ts&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; (which delegates to &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;kcmd&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Point your existing agents at Knowledge Catalog. New context becomes reachable through the same LookupContext and searchEntries calls they already use.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;&lt;/div&gt;</description><pubDate>Wed, 26 Aug 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/scale-okf-bundles-across-an-organization-with-knowledge-catalog/</guid><category>Data Analytics</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Using OKF with Knowledge Catalog to serve context for agents</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/scale-okf-bundles-across-an-organization-with-knowledge-catalog/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Firat Elbey</name><title>Group Product Manager, Data Analytics</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Sam McVeety</name><title>Tech Lead, Data Analytics</title><department></department><company></company></author></item><item><title>Serverless Apache Spark on Google Cloud: Architecture Choices &amp; AI Troubleshooting</title><link>https://cloud.google.com/blog/products/data-analytics/serverless-apache-spark-on-google-cloud-architecture-ai-troubleshooting/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In modern enterprise data engineering, Apache Spark remains a cornerstone framework for processing massive datasets at scale. However, managing infrastructure such as provisioning clusters, tuning YARN configurations, and avoiding costs for idle hardware often detracts from what matters most: building resilient data pipelines. Google Cloud addresses this operational overhead via its &lt;/span&gt;&lt;a href="https://cloud.google.com/products/managed-service-for-apache-spark"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Managed Service for Apache Spark&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, offering flexible deployment modes of serverless and managed clusters tailored to specific operational needs.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This technical guide walks through the architectural decision matrix for deploying Spark on Google Cloud, details resource and cost optimization techniques, and demonstrates how to apply built-in &lt;/span&gt;&lt;a href="https://cloud.google.com/products/gemini/cloud-assist"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Cloud Assist&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to rapidly troubleshoot and resolve serverless batch pipeline failures. While there is benefit to reading these three parts in a sequence, each one can be read independently and add value to how you approach Spark development on Google Cloud. &lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Part 1: Choosing your Apache Spark deployment model&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When launching Spark workloads on Managed Service for Apache Spark, the first major decision point is evaluating whether to construct traditional managed clusters or transition to a zero-management, serverless infrastructure footprint.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Decision #1: Managed clusters vs. serverless&lt;/strong&gt;&lt;/h3&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_uYxUREr.max-1000x1000.png"
        
          alt="1"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="m5sk9"&gt;*Created using Nano Banana 2 in Gemini Enterprise Agent Platform&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Choosing between traditional Managed Spark clusters and serverless depends on ecosystem requirements, infrastructure control needs, and financial utilization patterns:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Workload frequency, latency sensitive workloads &amp;amp; financial fit:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; For continuous, highly predictable, 24/7 streaming or batch processing pipelines where cluster nodes maintain constant high utilization baselines (80%+) or when the workflow’s accumulated startup time risk meeting SLA target, a permanently running, finely tuned traditional cluster, with custom YARN autoscaling rules, can sometimes be more cost-predictable. Conversely, for intermittent, bursty, ad-hoc, or orchestrator-triggered pipelines, Managed Spark serverless is highly optimal, eliminating operational management, requiring less planning time and ensuring you don’t pay for idle compute time.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Ecosystem &amp;amp; component requirements:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Managed Spark serverless is strictly optimized for Apache Spark 3.x+ codebases. If your processing pipeline relies on other ecosystem components such as Apache Flink, Presto/Trino, Hive LLAP, or Apache HBase, or if you are locked into a legacy Spark 2.x codebase, you must use Managed Spark clusters.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Infrastructure customization needs:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Managed Spark serverless abstracts away the underlying virtual machine (VM) layer. If your workload mandates deep OS-level hardware tuning, custom OS initialization actions, root SSH access to instances, specific local SSD configurations, or custom machine shapes, a traditional cluster is required. Note that serverless does support custom Docker container images for bundling specific application-level libraries.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Decision #2: Serverless interactive sessions vs. serverless batches&lt;/strong&gt;&lt;/h3&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_3XNMzOd.max-1000x1000.png"
        
          alt="2"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="m5sk9"&gt;*Created using Nano Banana 2 in Gemini Enterprise Agent Platform&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Once you select the serverless deployment mode, you must choose the appropriate execution model based on your development stage and operational requirements. Managed Service for Apache Spark provides two options for running serverless workloads:&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Serverless interactive sessions&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Interactive sessions are great for iterative and exploratory use cases. You write blocks of code, inspect intermediate DataFrames, modify variables, and generate visualizations with your dataset held warm in-memory.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Primary interface&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Designed for human-in-the-loop interaction. Developers execute code cell-by-cell using their IDE of choice, such as Colab, Gemini Enterprise Agent Platform Workbench, Antigravity, Jupyter notebooks, etc.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Idle cost profile&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Compute resources remain active to support immediate execution during developer thinking time, which can incur some idle compute charges if sessions are left inactive.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Serverless batches&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Batches are useful when you know what you want to run, and need automated, non-interactive execution. The engine runs fully completed, packaged PySpark scripts (.py) or Java/Scala application files (.jar) from start to finish without manual human intervention.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Primary interface&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Managed by automated orchestrators, such as &lt;/span&gt;&lt;a href="https://cloud.google.com/products/managed-service-for-apache-airflow"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Managed Service for Apache Airflow&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/scheduler/docs"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Scheduler&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, or &lt;/span&gt;&lt;a href="https://www.skills.google/course_templates/691" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;CI/CD pipelines&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Idle cost profile&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Billed strictly for the duration of the run. Compute resources are provisioned on-demand, run the script, and immediately shut down upon completion to prevent idle costs.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;The development-to-production lifecycle&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;These execution options are designed to work together as a natural pipeline lifecycle. During the initial development phase, you open a serverless interactive session within your notebook interface to explore datasets, clean schemas, and prototype transformations. Once your logic is validated and the transformations are finalized, you package the code into a Python script and schedule it as a serverless batch job orchestrated by Managed Service for Apache Airflow for production execution. This transition minimizes ongoing development costs while maintaining operational reliability.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Part 2: Advanced performance tuning and DCU cost optimization&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;While serverless Managed Spark eliminates the operational overhead of cluster maintenance, running production enterprise-grade pipelines on default settings can result in performance bottlenecks or budget waste. Resource allocation must be explicitly declared during submission using runtime configuration properties to maintain an efficient Data Compute Unit (DCU) burn rate.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Google recently introduced history-based &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/managed-spark/docs/concepts/autotuning"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;autotuning&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. In the context of serverless, this capability automatically applies optimizations based on best practices and historical execution. It does this by grouping recurring batch workloads into what Google calls cohorts. The autotuner analyzes the telemetry and statistics from previous runs under that same cohort name to figure out where the bottlenecks are.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Customizing driver and executor shapes&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By default, serverless batches allocate generic specifications (4 cores and 16,000MB RAM). This can cause critical efficiency issues depending on the nature of the application:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;The Memory-Bound job:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Pipelines processing highly uncompressed data volumes may hit Out-Of-Memory (OOM) errors and crash. To counter this, increase heap sizing independently using &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;spark.driver.memory&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;spark.executor.memory&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;The Compute-Bound Job:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Processing-intensive jobs running mathematical modeling or heavy tokenization might saturate CPUs while leaving expensive RAM sitting idle. Fine-tune processing concurrency per instance by explicitly adjusting &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;spark.driver.cores&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;spark.executor.cores&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Remember that by default increasing cores, automatically provisions a proportionate baseline   &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;of memory to match the vCPU-to-RAM ratio. This is why overriding the values for both cores and memory is critical  &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Controlling autoscaling boundaries&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Managed Spark serverless dynamically scales up and down the number of active executors based on backlogged tasks. However, unconstrained scaling can lead to budget overruns if a rogue code loop or unoptimized cartesian join is introduced.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;As a defensive guardrail, always declare an explicit upper limit using &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;spark.dynamicAllocation.maxExecutors&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;. This acts as your budget deadman-switch. By capping this at a reasonable ceiling, you guarantee that even if the code behaves sub-optimally, the job will never scale past a fixed infrastructure footprint.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;High priority (SLA-driven):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Set &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;maxExecutors&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; to a higher ceiling to allow resource bursting and minimize overall runtime duration.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Low priority (nightly batch):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Set &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;maxExecutors&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; to a low, tight ceiling. The workload will run longer but will consume a predictable, flat, cost-efficient stream of DCUs.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Managing shuffle storage efficiency&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When execution involves wide transformations like &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;groupBy()&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;join()&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;, or &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;distinct()&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;, data must be redistributed across the network, generating intermediate disk writes known as shuffle storage&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;.&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Spark defaults to a static setting of 200 partitions (&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;spark.sql.shuffle.partitions&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;). If you are processing a massive, multi-gigabyte dataset, 200 partitions means each individual chunk will be too large. When a partition's size exceeds available executor RAM (e.g., a 1GB partition trying to process inside 0.5GB of assigned heap space), data spills onto disk. This slows execution and incurs additional billing fees for premium or standard shuffle storage blocks. A helpful rule of thumb: Dynamically scale your partition parameters based on total data size so that each partition handles roughly&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;100MB to 200MB&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;of data in memory. This may require a few iterations before the optimal results are achieved.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The above properties are the main tunable properties. Additional Serverless runtime configuration properties can be found in this &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/managed-spark/docs/concepts/spark-properties-serverless"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;link&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Part 3: Operational diagnosis with Gemini Cloud Assist&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When automated data pipelines fail in production, data engineers are traditionally forced to spend hours sifting through verbose, disjointed log files across drivers and executors. Managed Service for Apache Spark addresses this friction by natively integrating &lt;/span&gt;&lt;a href="https://cloud.google.com/products/gemini/cloud-assist"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Cloud Assist&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; into the Google Cloud console, allowing engineers to diagnose and resolve failures using natural language.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To illustrate this operational shift, we examine the typical troubleshooting lifecycle for a failed PySpark ETL pipeline that reads customer transaction data from a &lt;/span&gt;&lt;a href="https://cloud.google.com/storage"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Storage&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (GCS) bucket, applies transformations, and encounters unexpected runtime errors.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Stage 1: Diagnosing missing execution parameters&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;During the initial execution attempt of a new pipeline, the batch job status switches from pending to running, and ultimately ends in a failed state with a generic exit message:&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Application failed with exit code 1&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Rather than manually querying Cloud Logging or navigating through multiple sections of the console, the engineer can locate the error log and select the ‘Investigate log’ option. This action opens a native conversation pane where Gemini Cloud Assist automatically analyzes the driver telemetry and system logs.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/3_B6AqsPB.max-1000x1000.png"
        
          alt="3"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In this scenario, the assistant explains in plain English that the PySpark script failed because required runtime arguments (such as the source GCS bucket path) were omitted during submission. It instantly identifies the exact lines in the script expecting these arguments, eliminating the need to read through the stack trace.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--medium
      
      
        h-c-grid__col
        
        h-c-grid__col--4 h-c-grid__col--offset-4
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/4_aZwid0v.max-1000x1000.png"
        
          alt="4"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Stage 2: Resolving schema and data type anomalies&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Once the missing arguments are resolved and the job is re-submitted, the pipeline runs but encounters a secondary data anomaly. In high-volume ingest pipelines, upstream source files frequently contain corrupted records or formatting inconsistencies.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Upon the second failure, the engineer again prompts Gemini Cloud Assist to investigate the logs. The assistant identifies a &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;TypeError&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; and pinpoints the exact DataFrame transformation causing the crash: a division operation (&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;df['amount'] / df['transaction_id']&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;) that failed because the schema auto-inferred the columns as strings.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/5_KItNZsH.max-1000x1000.png"
        
          alt="5"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Additionally, the assistant scans the underlying GCS file data to identify the root cause: non-numeric anomalies (such as text strings within numerical cells) in the source dataset.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--medium
      
      
        h-c-grid__col
        
        h-c-grid__col--4 h-c-grid__col--offset-4
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/6_DMwKNVq.max-1000x1000.png"
        
          alt="6"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Stage 3: Generating and deploying verified code fixes&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Rather than manually rewriting the PySpark logic to cast schema types and catch null values, the engineer can prompt Gemini Cloud Assist directly to generate a resilient solution:&lt;/span&gt;&lt;/p&gt;
&lt;p style="padding-left: 40px;"&gt;&lt;strong style="vertical-align: baseline;"&gt;User Prompt&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;"Suggest how to rewrite the code to divide the amount by quantity instead of transaction_id. In addition, add logic to skip invalid records without failing the process."&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The assistant generates the corrected PySpark code block, using resilient casting and null-handling functions (such as &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;coalesce&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;try_cast&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;).&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By implementing this corrected script, the orchestration pipeline can filter out bad source records smoothly without crashing the entire batch run. The subsequent execution completes successfully, preserving data freshness SLAs.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Unlock serverless Apache Spark: Benefits and next steps&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Managing data processing pipelines should not require a deep specialization in infrastructure configuration. By pairing the hands-off scale of serverless batches with explicit resource tuning — such as dynamic allocation caps and calculated shuffle sizing — data teams can maintain strict control over performance and cost profiles. When failures do occur, integrating Gemini Cloud Assist directly into your logging workflows transforms complex troubleshooting from a manual log-sifting exercise into a rapid, automated cycle.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To start putting these architectures into practice, you can explore the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/dataproc/docs"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Managed Service for Apache Spark documentation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and execute a serverless batch directly in the &lt;/span&gt;&lt;a href="https://console.cloud.google.com/dataproc"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud console&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For a deep architectural analysis of these concepts, get instant access to &lt;/span&gt;&lt;a href="https://services.google.com/fh/files/misc/google_cloud_apache_spark_whitepaper.pdf" rel="noopener" target="_blank"&gt;&lt;span style="font-style: italic; text-decoration: underline; vertical-align: baseline;"&gt;A practitioner’s guide to Apache Spark® in the agentic era&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. This guide includes step-by-step workflows, Codelabs, and runnable PySpark and Terraform templates directly from our GitHub repository. If you are new to Google Cloud, you can test these blueprints on serverless and managed clusters at zero cost by signing up for a &lt;/span&gt;&lt;a href="https://cloud.google.com/free"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;free trial with $300 in credits&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Wed, 19 Aug 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/serverless-apache-spark-on-google-cloud-architecture-ai-troubleshooting/</guid><category>Streaming</category><category>Data Analytics</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/image7_5rgoVhK.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Serverless Apache Spark on Google Cloud: Architecture Choices &amp; AI Troubleshooting</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/image7_5rgoVhK.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/serverless-apache-spark-on-google-cloud-architecture-ai-troubleshooting/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Lior (Leo) Ginzberg</name><title>Data &amp; Analytics Customer Engineer, Google Cloud</title><department></department><company></company></author></item><item><title>How to modernize Apache Hive using Google Cloud’s Lakehouse runtime catalog</title><link>https://cloud.google.com/blog/products/data-analytics/lakehouse-runtime-catalog-helps-modernize-apache-hive/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For over a decade, the Apache Hive Metastore (HMS) has served as the de facto metadata authority for big data analytics. Whether it was deployed on Hadoop clusters, self-managed Compute Engine VMs backed by MySQL or PostgreSQL, HMS provided the central schema registry that let Apache Spark, Presto, and Hive query raw &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;.parquet&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;.orc&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; files.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;However, as enterprise data architectures scale to petabytes and span multiple query engines (such as Google Cloud Managed Service for Apache Spark, BigQuery, and Trino), legacy Hive Metastores often become critical operational bottlenecks.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In this blog, we explore why legacy metastores struggle in modern cloud environments at agent scale, and show you how the serverless Google Cloud Lakehouse runtime catalog that we &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/biglake-metastore-now-supports-iceberg-rest-catalog?e=a"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;introduced last year&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; can help: Built on the open Apache Iceberg REST catalog specification, it is a runnable, zero-data-copy migration solution to help you transition your production Hive tables in minutes.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;The challenges of legacy Hive Metastores&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When speaking with data engineers and infrastructure leads running production analytics at scale, three core pain points consistently emerge with standalone Hive Metastores:&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Architectural and scaling bottlenecks&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Standalone HMS deployments rely on relational database backends (such as MySQL or Postgres) to track table schemas, partitions, and storage locations. As data lakes grow to hundreds of thousands of partitioned tables, partition pruning and bulk listing operations lead to key performance bottlenecks on the relational database. A complex Spark job requesting partition metadata can spike metastore CPU to 100%, causing cluster-wide query delays or out-of-memory (OOM) failures.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Siloed identity and security governance&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Legacy metastores were designed around perimeter-based Hadoop security models. Enforcing modern granular data governance — such as table-level access control lists (ACLs) — across both Apache Spark compute jobs and enterprise SQL engines like BigQuery requires maintaining fragmented, duplicated security policies across two distinct control planes.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Operational overhead and  total cost of ownership (TCO)&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Managing high-availability MySQL/Postgres instances, patching HMS daemons, tuning JDBC connection pools, and paying for idle instance-based metastore servers creates unnecessary operational toil for data platform teams, whose time is better spent building high-leverage data products for agents.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;The solution: Lakehouse runtime catalog&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To solve these architectural bottlenecks without forcing data engineers to rewrite petabytes of existing storage payloads, we built the Lakehouse runtime catalog with support for Iceberg Rest Catalog and Hive Catalog.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The Lakehouse runtime catalog is a fully serverless, highly available, and unified metadata registry designed from the ground up to support both legacy Hive/Parquet tables and modern open table formats like&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Apache Iceberg. By natively implementing the Apache Iceberg REST Catalog specification&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;,&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; the Lakehouse runtime catalog decouples metadata discovery from compute engines. This decoupling of the catalog and compute engines ensures multiple Iceberg compatible engines can access the same data in a zero copy fashion thereby reducing the need for customers to maintain multiple copies of the data and enables them to take their workloads to production sooner.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_i5lkwYb.max-1000x1000.png"
        
          alt="1"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This approach offers a number of architectural benefits:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Multi-engine interoperability&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Once registered, tables are immediately discoverable and queryable across Google Cloud Managed Spark, BigQuery, and open-source engines via standard REST interfaces.&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Open APIs&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Supports Iceberg Rest Catalog and Hive Catalog which enables different teams to use their preferred analytics tools on a single, unified dataset.&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Zero-data copy&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Table definitions point directly to your existing data in Google Cloud Storage. You do not move, rewrite, or duplicate your underlying data.&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;AI-powered governance, security and trusted context&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The Lakehouse runtime catalog integrates directly with &lt;/span&gt;&lt;a href="https://cloud.google.com/products/knowledge-catalog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Knowledge Catalog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://www.cloud-iam.com/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud IAM&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, allowing you to define trusted context for your agents and table-level security that apply consistently across all compute engines. Further it supports &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;key authorization mechanisms, such as credential vending&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;. This means you can access your tables without needing direct access to the files in the underlying Cloud Storage bucket.&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Enterprise-readiness, scale and reduced TCO: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Backed by Google’s planet-scale infrastructure and Spanner, enabling your metadata to scale with your data. Support for Cloud Storage dual-region and multi-region buckets enables failover use cases. It also provides reduced TCO due to serverless and no-ops environments, and scalability for any workload size.&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Zero-copy migration from legacy Hive Metastore in action&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To demonstrate how smooth cutover is in practice, we have provided a capability that lets you modernize your legacy self-managed Hive Metastore to the Google Cloud Lakehouse. This capability connects directly to your legacy Hive Metastore, extracts external table definitions and partition maps, and registers them cleanly into the serverless Lakehouse catalog and then start using the data in Google Managed Spark, BigQuery and Conversational Analytics agents with Gemini. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Modernize to the Lakehouse and immediately tap your data in key agentic journeys&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/2_ykGR7QL.gif"
        
          alt="2"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Ready to modernize your data architecture?&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Modernizing from a legacy Hive Metastore to Google Cloud’s &lt;/span&gt;&lt;a href="https://cloud.google.com/products/lakehouse?e=48754805&amp;amp;hl=en"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Lakehouse&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; minimizes data silos across analytics engines and agents, unifies multi-engine governance, provides trusted context to your agents and slashes operational TCO. In other words, it helps prepare your modern cloud environments to operate at agent scale.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Get started and &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/hdfs-data-lake-transfer"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;migrate your Apache Hive Metastore tables to Google Cloud&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; today, and get ready for the agentic era. Learn more about Google Cloud Lakehouse &lt;/span&gt;&lt;a href="https://cloud.google.com/products/lakehouse?e=a"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Wed, 19 Aug 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/data-analytics/lakehouse-runtime-catalog-helps-modernize-apache-hive/</guid><category>Data Analytics</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>How to modernize Apache Hive using Google Cloud’s Lakehouse runtime catalog</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/data-analytics/lakehouse-runtime-catalog-helps-modernize-apache-hive/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Vinod Ramachandran</name><title>Product Manager</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Pratibha Suryadevara</name><title>Vice President</title><department></department><company></company></author></item></channel></rss>