<?xml version="1.0" encoding="utf-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:media="http://search.yahoo.com/mrss/"><channel><title>Developers &amp; Practitioners</title><link>https://cloud.google.com/blog/topics/developers-practitioners/</link><description>Developers &amp; Practitioners</description><atom:link href="https://cloudblog.withgoogle.com/blog/topics/developers-practitioners/rss/" rel="self"></atom:link><language>en</language><lastBuildDate>Fri, 02 Oct 2026 15:57:01 +0000</lastBuildDate><image><url>https://cloud.google.com/blog/topics/developers-practitioners/static/blog/images/google.a51985becaa6.png</url><title>Developers &amp; Practitioners</title><link>https://cloud.google.com/blog/topics/developers-practitioners/</link></image><item><title>GKE CPU startup boost: Accelerate app starts without over-provisioning</title><link>https://cloud.google.com/blog/products/containers-kubernetes/gke-cpu-startup-boost-faster-pod-starts-lower-costs/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Whether you’re launching microservices in response to sudden traffic spikes, deploying new software releases, or scaling up application replicas, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;pod startup time&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; is critical to maintaining a fast, responsive user experience for applications running on Google Kubernetes Engine (GKE).&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-paragraph_with_image"&gt;&lt;div class="article-module h-c-page"&gt;
  &lt;div class="h-c-grid uni-paragraph-wrap"&gt;
    &lt;div class="uni-paragraph
      h-c-grid__col h-c-grid__col--8 h-c-grid__col-m--6 h-c-grid__col-l--6
      h-c-grid__col--offset-2 h-c-grid__col-m--offset-3 h-c-grid__col-l--offset-3"&gt;

      






  

    &lt;figure class="article-image--wrap-small
      
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_mhC0eeP.max-1000x1000.png"
        
          alt="1"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  





      &lt;p data-block-key="0udcy"&gt;Yet, platform engineers and developers face a persistent dilemma: Applications often demand significantly more CPU power during startup than they do during steady-state operations. Sizing CPU requests for normal, steady-state usage leads to CPU throttling during launch, which can result in sluggish cold starts and readiness probe timeouts. On the flip side, over-provisioning baseline CPU requests to satisfy short-lived startup bursts wastes valuable compute resources, inflating infrastructure bills.&lt;/p&gt;&lt;p data-block-key="ciu5o"&gt;Today, we are excited to announce &lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/how-to/boost-application-startup"&gt;&lt;b&gt;CPU startup boost&lt;/b&gt;&lt;/a&gt; for GKE in preview. Integrated directly into GKE's Vertical Pod Autoscaler (VPA), CPU startup boost dynamically elevates a container's CPU allocation during initialization and seamlessly scales it back to baseline steady-state levels once the application is ready - &lt;b&gt;all without restarting your containers.&lt;/b&gt;&lt;/p&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Why modern applications need extra CPU at boot time&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When a new container launches, it may perform intensive initialization tasks before it begins serving user requests. Depending on your tech stack, the following startup workloads require substantial CPU cycles:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Java JVM applications&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Frameworks like Spring Boot require high CPU burst capacity for class loading, classpath scanning, instantiating dependency injection containers, and running Just-in-Time (JIT) compilation.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Node.js servers&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Apps parse JavaScript files, build complex module dependency trees (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;require&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;/&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;import&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;), and execute V8 engine optimization and JIT compilation passes during initial execution.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Python and AI/ML microservices&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: These services spend initial cycles importing heavy libraries (such as PyTorch, NumPy, or LangChain), compiling &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;.pyc&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; bytecode, establishing ORM database schemas, and pre-loading cache structures.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If you size CPU requests strictly for steady-state performance, these initialization workloads experience CPU throttling on launch, delaying readiness probes. To prevent slow cold starts, teams frequently overprovision CPU requests. However, once the application stabilizes, those extra CPU resources sit idle, increasing your cloud spend without adding value.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-paragraph_with_image"&gt;&lt;div class="article-module h-c-page"&gt;
  &lt;div class="h-c-grid uni-paragraph-wrap"&gt;
    &lt;div class="uni-paragraph
      h-c-grid__col h-c-grid__col--8 h-c-grid__col-m--6 h-c-grid__col-l--6
      h-c-grid__col--offset-2 h-c-grid__col-m--offset-3 h-c-grid__col-l--offset-3"&gt;

      






  

    &lt;figure class="article-image--wrap-small
      
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_zZExLw4.max-1000x1000.jpg"
        
          alt="2"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  





      &lt;h3 data-block-key="0udcy"&gt;&lt;b&gt;How CPU startup boost can help&lt;/b&gt;&lt;/h3&gt;&lt;p data-block-key="8k1a3"&gt;CPU Startup Boost solves this by giving your workloads temporary vCPU "boosts" during launch, and automatically returning them to baseline once initialization completes.&lt;/p&gt;&lt;p data-block-key="c9jmj"&gt;Key benefits:&lt;/p&gt;&lt;ul&gt;&lt;li data-block-key="96uhr"&gt;&lt;b&gt;Faster cold starts&lt;/b&gt;: Reduce application initialization times by up to &lt;b&gt;2x&lt;/b&gt;, accelerating auto-scaling responsiveness during unexpected traffic surges.&lt;/li&gt;&lt;li data-block-key="f852p"&gt;&lt;b&gt;Optimized cloud spend&lt;/b&gt;: Right-size steady-state CPU requests to fit actual runtime needs rather than paying for idle startup headroom.&lt;/li&gt;&lt;li data-block-key="5opq6"&gt;&lt;b&gt;Zero pod restarts&lt;/b&gt;: Dynamic resource resizing happens live inside the running container.&lt;/li&gt;&lt;li data-block-key="3r67d"&gt;&lt;b&gt;Flexible policy controls&lt;/b&gt;: Apply simple pod-level multiplier factors (e.g., 2x CPU during startup) or define granular, container-specific rules for complex multi-container pods.&lt;/li&gt;&lt;/ul&gt;
    &lt;/div&gt;
  &lt;/div&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Under the hood: Kubernetes In-place Pod Resize&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Historically, changing a pod's resource requests or limits required deleting and recreating the pod. This disruptive process triggered container restarts, cache invalidation, and node rescheduling overhead.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To fix that, CPU startup boost builds on Kubernetes &lt;/span&gt;&lt;a href="https://kubernetes.io/docs/tasks/configure-pod-container/resize-container-resources/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;In-place Pod Resize (IPPR)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Tracked under &lt;/span&gt;&lt;a href="https://github.com/kubernetes/enhancements/tree/master/keps/sig-node/1287-in-place-pod-resize" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;KEP-1287&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, IPPR introduced dynamic, in-place resource mutation. Introduced as Alpha in Kubernetes 1.27, promoted to Beta in v1.33, and graduating to General Availability (GA) in v1.35, IPPR allows the Kubernetes control plane and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;kubelet&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; to update container CPU and memory requests on running pods &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;without restarting the container process&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;GKE leverages IPPR within the VPA  to apply startup CPU boosts at pod admission and smoothly step them down post-readiness.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;How CPU startup boost works (pod lifecycle overview)&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;CPU startup boost operates across three distinct phases:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Admission phase&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: When you deploy a pod, the VPA admission webhook intercepts the creation request. It calculates the elevated CPU request based on your policy (e.g., 2x multiplier or +2 vCPUs) and injects the boosted CPU request along with tracking annotations into the pod spec.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Startup phase&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The pod is scheduled and initialized with the boosted CPU allocation. Your application completes class loading, JIT compilation, or module parsing at top speed without experiencing CPU throttling.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Unboosting phase&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: As soon as the pod's &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;readinessProbe&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; passes (plus any configured &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;durationSeconds&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; cooldown delay), the VPA updater issues an in-place resize request. The CPU request steps back down to your baseline level while the container continues running uninterrupted.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Prerequisites and availability&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;CPU startup boost is available today in preview on GKE:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;GKE version&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Version &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;1.36.0-gke.4447000&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; or later on Standard and Autopilot clusters.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Cluster Modes&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Enabled natively on GKE Autopilot (VPA is active by default). On GKE Standard, simply ensure Vertical Pod Autoscaling (VPA) is enabled.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Workload Support&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Works with standard Kubernetes controllers, including Deployments and StatefulSets.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Getting started: Configuring CPU startup boost&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Configuring CPU startup boost is as simple as adding a &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;startupBoost&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; section to your &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;VerticalPodAutoscaler&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; manifest.&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Example 1: Pod-level boost with fixed steady-state (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;updateMode: "Off"&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;)&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If you want to use VPA purely for startup boost while keeping steady-state CPU requests locked to your manifest definitions, set &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;updateMode&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; to &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;"Off"&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;apiVersion: &amp;quot;autoscaling.k8s.io/v1&amp;quot;\r\nkind: VerticalPodAutoscaler\r\nmetadata:\r\n  name: java-app-startup-boost\r\n  namespace: default\r\nspec:\r\n  targetRef:\r\n    apiVersion: &amp;quot;apps/v1&amp;quot;\r\n    kind: Deployment\r\n    name: java-app\r\n  updatePolicy:\r\n    updateMode: &amp;quot;Off&amp;quot;\r\n  startupBoost:\r\n    cpu:\r\n      type: &amp;quot;Factor&amp;quot;\r\n      factor: 2\r\n      durationSeconds: 10&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e67a9af10&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;In this example, GKE doubles the container's CPU request during launch and holds the boosted allocation for 10 seconds after readiness probes pass before scaling back to baseline.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Example 2: Combining startup boost with continuous VPA auto-scaling&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If you want GKE to boost CPU during launch &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;and&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; continuously optimize steady-state resources post-startup, set &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;updateMode&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; to &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;"InPlaceOrRecreate"&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;apiVersion: &amp;quot;autoscaling.k8s.io/v1&amp;quot;\r\nkind: VerticalPodAutoscaler\r\nmetadata:\r\n  name: nodejs-app-vpa\r\n  namespace: default\r\nspec:\r\n  targetRef:\r\n    apiVersion: &amp;quot;apps/v1&amp;quot;\r\n    kind: Deployment\r\n    name: nodejs-service\r\n  updatePolicy:\r\n    updateMode: &amp;quot;InPlaceOrRecreate&amp;quot;\r\n  startupBoost:\r\n    cpu:\r\n      type: &amp;quot;Factor&amp;quot;\r\n      factor: 3\r\n      durationSeconds: 15&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e67a9a450&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Example 3: Granular container-level boost&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For pods running sidecars (such as logging agents or service mesh proxies) that do not require extra CPU on boot, target specific app containers:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;apiVersion: &amp;quot;autoscaling.k8s.io/v1&amp;quot;\r\nkind: VerticalPodAutoscaler\r\nmetadata:\r\n  name: app-container-boost\r\nspec:\r\n  targetRef:\r\n    apiVersion: &amp;quot;apps/v1&amp;quot;\r\n    kind: Deployment\r\n    name: API-gateway\r\n  updatePolicy:\r\n    updateMode: &amp;quot;Off&amp;quot;\r\n  resourcePolicy:\r\n    containerPolicies:\r\n    - containerName: &amp;quot;web-app&amp;quot;\r\n      mode: &amp;quot;Off&amp;quot;\r\n      startupBoost:\r\n        cpu:\r\n          type: &amp;quot;Quantity&amp;quot;\r\n          quantity: &amp;quot;2&amp;quot;\r\n          durationSeconds: 5&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e67a995d0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Verifying startup boost in your cluster&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;You can verify that GKE applied and downscaled the startup boost using &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;kubectl&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;1. &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Inspect pod annotations&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Check for the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;vpaCpuStartupBoost&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; tracking annotation:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;kubectl get pod &amp;lt;POD_NAME&amp;gt; -o yaml&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e67a99190&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;Look for annotations indicating the original baseline and boosted CPU requests.&lt;/span&gt;&lt;/p&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;2. &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Monitor in-place resize events&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Confirm that GKE downscaled the CPU request back to baseline after readiness:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;kubectl get events --field-selector reason=InPlaceResizedByVPA&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e67a9b390&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;An event with &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;reason=InPlaceResizedByVPA&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; confirms successful in-place downscaling post-startup.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Best practices for production workloads&lt;/span&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Pairing with Horizontal Pod Autoscaler (HPA)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: When using HPA alongside CPU startup boost, ensure a robust &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;readinessProbe&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; is defined and keep &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;durationSeconds&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; short (e.g., &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;0s&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;–&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;10s&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;). This prevents HPA from falsely interpreting initialization CPU spikes as high steady-state load.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Handle traffic spikes with &lt;/strong&gt;&lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/how-to/configure-capacity-buffer"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;GKE capacity buffers API&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;: By reducing the startup tax at scale, you can achieve higher workload density on fewer nodes to improve overall utilization. Consider adopting the GKE capacity buffers API to absorb sudden traffic surges with minimal operational overhead while maintaining strict SLOs. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;GKE Autopilot resource ratios&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: On GKE Autopilot, remember that pods must maintain valid CPU-to-memory ratios. Ensure baseline memory allocations accommodate the boosted CPU ratio during startup.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Get started today&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;CPU startup boost gives GKE users the best of both worlds: lightning-fast cold starts for CPU-intensive workloads like Java, Node.js, and Python, paired with maximum resource efficiency and lower cloud costs.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Ready to accelerate your GKE workloads?&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Explore the &lt;/span&gt;&lt;a href="https://cloud.google.com/kubernetes-engine/docs/how-to/cpu-startup-boost"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GKE CPU startup boost documentation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Learn more about &lt;/span&gt;&lt;a href="https://cloud.google.com/kubernetes-engine/docs/how-to/vertical-pod-autoscaling"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Vertical Pod Autoscaling on GKE&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Try CPU startup boost on your GKE Standard or Autopilot clusters running GKE &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;1.36.0-gke.4447000&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; or later!&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Fri, 02 Oct 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/containers-kubernetes/gke-cpu-startup-boost-faster-pod-starts-lower-costs/</guid><category>Compute</category><category>GKE</category><category>Developers &amp; Practitioners</category><category>Containers &amp; Kubernetes</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>GKE CPU startup boost: Accelerate app starts without over-provisioning</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/containers-kubernetes/gke-cpu-startup-boost-faster-pod-starts-lower-costs/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Abdel Sghiouar</name><title>Cloud Developer Advocate</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Jakub Pawełczak</name><title>Cloud Product Manager</title><department></department><company></company></author></item><item><title>Democratizing Managed Lustre with lower cost and frictionless development</title><link>https://cloud.google.com/blog/topics/developers-practitioners/democratizing-managed-lustre-with-lower-cost-and-frictionless-development/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This is the first of a two-part series exploring how Google Cloud is bringing the foundational values of a high-performance parallel filesystem–TB/s throughput, sub-ms latency at high client scale, and POSIX support–to a broader set of use cases and users.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Historically, due to the cost and special purpose nature of parallel filesystems, colder data had to be stored outside of the filesystem and AI developers have had to maintain separate, slower environments for writing code, compiling libraries, and managing repositories. This fragmentation increases the toil of manual data staging, dataset copying, and managing disjointed namespaces.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Google Cloud Managed Lustre is solving these problems through our &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;6 cents/GB*month Dynamic Tier&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;and by optimizing Managed Lustre performance for a range of development tasks and workloads – making Managed Lustre a “One-Stop Shop” for high-performance AI and HPC workloads.&lt;/strong&gt;&lt;/p&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Lower Cost: More Lustre for Less with the Dynamic Tier&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The Managed Lustre Dynamic Tier provides &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;sub-ms latency for hot data, which allows you to store all of your data in a single namespace, and costs only 6 cents/GB*month&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Throughput, capacity scale and client scale:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Throughput scales linearly with capacity up to 80 PB, while sub-ms latency for hot data remains stable as you scale to tens of thousands of clients.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Single-flat fee:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Predictable pricing. No independent charges for disk media types, data movement within the namespace, or metadata IOPS.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Read Latencies:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Sub-ms latencies for High-Performance Cache (SSD).  The Capacity Pool (“HDD”) is built on Google Cloud Hyperdisk throughput, which has an &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/compute/docs/disks/hd-types/hyperdisk-throughput"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;average read latency of 10 to 30 ms&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;strong style="vertical-align: baseline;"&gt;Recommended workloads for Dynamic Tier&lt;/strong&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Multi-Epoch Training and/or Training with Optimized Fetch Sizes: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Hot data is promoted to the High Performance Cache (SSD) after the first run. Larger data prefetch will allow you to take advantage of the Dynamic Tier cost structure and gain from low-latency SSD.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Write-Heavy Checkpointing:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Bursty checkpoint writes land directly in the High Performance Cache. Older checkpoints are transparently demoted to the Capacity Pool (HDD).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Rapid Checkpoint Restore:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;  New checkpoints are written to the High Performance Cache, enabling low-latency checkpoint restores.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Interactive Snappiness for Developers:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Low-latency tasks like git cloning, compiling libraries, or running notebooks benefit from a local-disk feel (~300µs average read latencies) on the same shared workspace hosting large training sets.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Frictionless development: Lustre as a one-stop shop for developer’s workloads&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In addition to Managed Lustre’s scalability for large AI and HPC workloads (checkpoint/restart/data-loading), it also meets the demands for interactive work, meaning developers can start on Managed Lustre and stay on Managed Lustre throughout the entire workload lifecycle:&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Unified Foundation &amp;amp; Interactive Performance&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Consolidates the AI and HPC lifecycle into a single namespace, providing a "local disk" feel for interactive work (Read more about the &lt;/span&gt;&lt;a href="https://cloud.google.com/products/managed-lustre"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;latency benefits of Managed Lustre experienced by Salesforce and others&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;).&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Latency:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; ~300µs average read latency—delivering up to 4x better responsiveness than alternative distributed file systems.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Accelerated Setup:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Untar the Linux kernel in ~2 minutes (4.7x faster than alternative file solutions), run a 20-worker parallel git clone of Python in ~40 seconds, compile Python in ~200s.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/democratizing-lustre-with-lower-cost-and-f.max-1000x1000.jpg"
        
          alt="democratizing-lustre-with-lower-cost-and-frictionless-development-02-managed-lustre-performance-values"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/democratizing-lustre-with-lower-cost-and-f.max-1000x1000_xLOfL6w.jpg"
        
          alt="democratizing-lustre-with-lower-cost-and-frictionless-development-01-performance-comparision"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;High-Concurrency Broadcast &amp;amp; Cluster Startup&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Managed Lustre maximizes GPU ROI by preventing storage bottlenecks during cluster initialization. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When thousands of worker nodes attempt to read the exact same file simultaneously (such as a shared model checkpoint, base weights, or container layer), traditional distributed file systems can choke on localized hotspotting, leaving high-cost GPU clusters idle for minutes.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Improves Aggregate Throughput for a large number of clients reading the same file:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Demonstrates a 67% improvement over alternative file solutions.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Parallel Loading:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Imports libraries like PyTorch across 4,000+ processes in under 60 seconds.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/democratizing-lustre-with-lower-cost-and-f.max-1000x1000_9kShfM1.jpg"
        
          alt="democratizing-lustre-with-lower-cost-and-frictionless-development-03-performance-advantage"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Run One-Stop Shop Workflows for Yourself&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Here is the code for the tests we’ve run, so that you can perform your own testing.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Low latency for interactive access&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We used &lt;/span&gt;&lt;a href="https://github.com/axboe/fio" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;fio&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to emulate small, low-concurrency reads and writes:&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;sup&gt;&lt;span style="vertical-align: baseline;"&gt;1 &lt;/span&gt;&lt;/sup&gt;&lt;span style="color: #5f6368; font-size: 16px; font-style: italic; font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;Storage system specs: 500 MBps per TiB tier of Managed Lustre, 108,000 GiB capacity. Zonal Filestore at 102,400 GiB capacity. Average throughput of 36.7 GB/s to 2,048 client VMs reading the same 40 GiB file.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;# Read workload\r\nfio --ioengine=libaio --filesize=100M --ramp_time=2s \\\r\n    --runtime=2m --time_based --numjobs=1 --direct=1 --verify=0 --randrepeat=0 \\\r\n    --group_reporting --directory=~/LUSTRE_MOUNT \\\r\n    --name=randread --blocksize=4k --iodepth=1 --readwrite=randread \\\r\n    --buffer_compress_percentage=50\r\n\r\n# Write workload\r\nfio --ioengine=libaio --filesize=100M --ramp_time=2s \\\r\n    --runtime=2m --time_based --numjobs=1 --direct=1 --verify=0 --randrepeat=0 \\\r\n    --group_reporting --directory=~/LUSTRE_MOUNT \\\r\n    --name=randwrite --blocksize=4k --iodepth=1 --readwrite=randwrite \\\r\n    --buffer_compress_percentage=50&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e7414c5d0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Accelerated setup&lt;/strong&gt;&lt;/h3&gt;
&lt;h4&gt;&lt;strong style="vertical-align: baseline;"&gt;How to run Linux untar&lt;/strong&gt;&lt;/h4&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;# Download a kernel tarball\r\nwget -P /tmp https://cdn.kernel.org/pub/linux/kernel/v5.x/linux-5.18.9.tar.xz\r\n\r\n# Extract the archive to the Lustre mount\r\nmkdir ~/LUSTRE_MOUNT/kernel\r\ntar -C ~/LUSTRE_MOUNT/kernel -xf /tmp/linux-5.18.9.tar.xz&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e7414ef50&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In the above use case, you will want to take care to avoid the metadata performance tax that can come from running as root (Namely, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;tar&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; issues&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt; chown&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; and&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt; chmod&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; calls to make extracted files’ owner+permissions match the ones recorded in the archive.).  If you still wish to run as root (and have verified that this approach is compatible with your setup), you may &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;specify&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;`--no-same-owner --no-same-permissions`&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;in order to ensure that extracted files maintain&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; root&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; as owner and have root's default file permissions. In other words, it makes extraction as root behave like extraction as non-root (by ignoring the owner+permissions in the archive).&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;strong style="vertical-align: baseline;"&gt;How to run Python gitclone&lt;/strong&gt;&lt;/h4&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;git config --global checkout.workers 20\r\nmkdir ~/LUSTRE_MOUNT/python\r\ngit clone https://github.com/python/cpython.git ~/LUSTRE_MOUNT/python&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e6767ca90&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;How to run Python compile&lt;/strong&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;pushd ~/LUSTRE_MOUNT/python\r\n./configure &amp;gt; /dev/null\r\nmake &amp;gt; /dev/null\r\npopd&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e6767f490&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;High scale distribution&lt;/strong&gt;&lt;/h3&gt;
&lt;h4&gt;&lt;strong style="vertical-align: baseline;"&gt;Aggregate throughput for distributing one large file to many nodes&lt;/strong&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Run the below on each client VM:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;# Start fio in server mode\r\nfio --server&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e6767c510&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;p data-block-key="iadne"&gt;Run the below on a selected client VM:&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;quot;# Create a 40 GiB file\r\nfio --name=job1 \\\r\n    --ioengine=libaio \\\r\n    --direct=1 \\\r\n    --buffer_compress_percentage=50 \\\r\n    --blocksize=4m \\\r\n    --iodepth=32 \\\r\n    --filesize=40g \\\r\n    --readwrite=write \\\r\n    --filename ~/LUSTRE_MOUNT/40gb_test\r\n\r\n# Create an fio job file for the read workload\r\ncat &amp;lt;&amp;lt;&amp;#x27;EOF&amp;#x27; &amp;gt; /tmp/read.fio\r\n[job1]\r\nfilename=${HOME}/LUSTRE_MOUNT/40gb_test\r\nrw=read\r\nbs=4m\r\nexitall_on_error=1\r\nEOF\r\n\r\n# Run the read workload using all client VMs in ~/hostfile\r\nfio --client ~/hostfile /tmp/read.fio&amp;quot;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e6767e210&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h4&gt;&lt;strong style="vertical-align: baseline;"&gt;Parallel loading of libraries across many processes&lt;/strong&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Run the below on a selected client VM:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;# Install PyTorch in a virtual env\r\npython3 -m venv ~/LUSTRE_MOUNT/env\r\nsource ~/LUSTRE_MOUNT/env/bin/activate\r\npip3 install --upgrade pip\r\npip3 install torch torchvision torchaudio \r\ndeactivate\r\n\r\n# Import PyTorch on all client VMs in ~/hostfile, 4 processes per host\r\nmpirun --allow-run-as-root --oversubscribe --hostfile ~/hostfile -N 4 \\\r\n  bash -c \&amp;#x27;source ~/LUSTRE_MOUNT/env/bin/activate &amp;amp;&amp;amp; python3 -c &amp;quot;import torch&amp;quot;\&amp;#x27;&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e6767d0d0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h1&gt;&lt;strong style="vertical-align: baseline;"&gt;Looking ahead and next steps&lt;/strong&gt;&lt;/h1&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By eliminating the manual data staging tax and lowering entry costs with the Dynamic Tier, Google Cloud Managed Lustre is evolving from an elite, single-purpose engine into a highly versatile, unified storage fabric for the entire AI lifecycle.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In the second part of this series, we will focus on &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;upcoming object integration features&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;. Stay tuned!&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Next steps&lt;/strong&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Run the benchmarks yourself (if you haven’t already): Deploy a Google Cloud Managed Lustre instance using the &lt;/span&gt;&lt;a href="https://console.cloud.google.com/"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud console&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and run tests provided above to benchmark your own workloads.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Explore the Dynamic Tier: Read the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/managed-lustre/docs/performance-tiers"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Managed Lustre Documentation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to learn more.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Stay tuned for Part 2: In the next installment of this series, we will dive deep into upcoming object integration features and how they further simplify AI and HPC storage.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Get started with centralizing your development-to-training lifecycle on Google Cloud Managed Lustre!&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Thu, 01 Oct 2026 13:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/developers-practitioners/democratizing-managed-lustre-with-lower-cost-and-frictionless-development/</guid><category>Developers &amp; Practitioners</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/democratizing-lustre-with-lower-cost-and-fri.max-600x600_y4nYtDt.jpg" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Democratizing Managed Lustre with lower cost and frictionless development</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/democratizing-lustre-with-lower-cost-and-fri.max-600x600_y4nYtDt.jpg</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/developers-practitioners/democratizing-managed-lustre-with-lower-cost-and-frictionless-development/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Barak Epstein </name><title>Senior Product Manager, Google Cloud Managed Lustre</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Yuval Ehrental</name><title>Software Engineer, Google Cloud Managed Lustre</title><department></department><company></company></author></item><item><title>Introducing the Server Side Cloud Swift SDK</title><link>https://cloud.google.com/blog/topics/developers-practitioners/introducing-the-server-side-cloud-swift-sdk/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For years, &lt;a href="https://swift.org/" rel="noopener nofollow noreferrer" target="_blank"&gt;Swift&lt;/a&gt; was perceived mainly as a UI language tied to Apple client devices. With Swift 6 and strict concurrency checking, it has matured into a viable systems and cloud language, pairing Rust-like data-race safety with predictable, reference-counted performance.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To support this ecosystem, Google engineering has launched the official &lt;/span&gt;&lt;a href="https://github.com/googleapis/google-cloud-swift" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud API Client Libraries for Swift&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. Built from the ground up for Swift 6.2+, this new SDK uses the latest non-blocking &lt;/span&gt;&lt;a href="https://github.com/apple/swift-nio" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Swift NIO&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; event loops, HTTP/2 multiplexing, &lt;/span&gt;&lt;a href="https://grpc.io" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;gRPC&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; transport, and zero-cost compile-time data race safety.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In this article, we'll walk you through all you need to know to get started, and to understand how the Server Side Cloud Swift SDK, or &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;google-cloud-swift&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, works.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;The rise of server-side Swift and cloud-native concurrency&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;"&gt;Traditional backend development often forces a trade-off between developer ergonomics and resource utilization. While managed runtimes offer rapid development, lower-level systems languages provide finer control over memory and CPU footprint, often at the expense of feature velocity.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;Server-side Swift aims to strike a practical balance. Swift pairs a lightweight runtime and Automatic Reference Counting (ARC) with expressive syntax. More importantly, Swift 6 introduces compile-time concurrency checking. &lt;span style="vertical-align: baseline;"&gt;When you share state between async tasks across a cloud microservice, the compiler enforces that types conform to &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;Sendable&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;. Data races are caught in your editor before a binary ever compiles or reaches production.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;At the network layer, every request to &lt;/span&gt;&lt;a href="https://cloud.google.com"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; APIs runs over event-driven, non-blocking sockets that scale across multicore &lt;/span&gt;&lt;a href="https://www.kernel.org" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Linux&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; server environments without spawning system threads per connection.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/sdk-architecture.max-1000x1000.png"
        
          alt="sdk-architecture"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Where to use the Swift SDK&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The Server Side Cloud Swift SDK is engineered for server, container, and automated DevOps environments.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When you build high-throughput microservices with Swift web frameworks like &lt;/span&gt;&lt;a href="https://hummingbird.codes" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Hummingbird&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; or &lt;/span&gt;&lt;a href="https://vapor.codes" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Vapor&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;google-cloud-swift&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; provides native access to Cloud Storage, AI, Identity and Access Management (IAM), and over one hundred other Google Cloud services. You can containerize your executable on Linux and deploy directly to &lt;/span&gt;&lt;a href="https://cloud.google.com/run"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Run&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/kubernetes-engine"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Kubernetes Engine (GKE)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, or &lt;/span&gt;&lt;a href="https://cloud.google.com/compute"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Compute Engine&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; VMs.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Because the SDK compiles on macOS, and Linux, you can develop the backend in your preferred development environment, and then seamlessly deploy to production. And using Swift on both the frontend and backend allows you to share application-specific types across both.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The SDK also excels at platform engineering and DevOps automation. You can author cross-platform CLI utilities and data rotation scripts that run on your developer laptop or inside CI/CD pipelines. These tools authenticate automatically against Google Cloud using &lt;/span&gt;&lt;a href="https://cloud.google.com/docs/authentication/application-default-credentials"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Application Default Credentials (ADC)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; or &lt;/span&gt;&lt;a href="https://cloud.google.com/iam/docs/workload-identity-federation"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Workload Identity Federation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If you're building an &lt;/span&gt;&lt;a href="https://developer.apple.com/ios/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;iOS&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://www.apple.com/ipados/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;iPadOS&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, or &lt;/span&gt;&lt;a href="https://www.apple.com/apple-vision-pro/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;visionOS&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; app for the &lt;/span&gt;&lt;a href="https://www.apple.com/app-store/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Apple App Store&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, you should not embed &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;google-cloud-swift&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; directly into your client bundle. Shipping Google Cloud service account keys or administrative credentials inside a client binary creates security risks. For direct client-side features, use the &lt;/span&gt;&lt;a href="https://github.com/firebase/firebase-ios-sdk" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Firebase SDK for Apple Platforms&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to handle user authentication, real-time &lt;/span&gt;&lt;a href="https://cloud.google.com/firestore"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Firestore&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; sync, and client-side security rules, or route requests through your own Cloud Run backend API.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Getting started with your IDE and packages&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Because &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;google-cloud-swift&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; treats Linux and macOS as first-class citizens, you can develop on Apple hardware with &lt;/span&gt;&lt;a href="https://developer.apple.com/xcode/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Xcode&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; or on Linux workstations with &lt;/span&gt;&lt;a href="https://code.visualstudio.com/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Visual Studio Code&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://swift.org/install/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;swiftly&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To install the official Swift compiler on Linux workstations using the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;swiftly&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; CLI installer, run:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;curl -O https://download.swift.org/swiftly/linux/swiftly-$(uname -m).tar.gz &amp;amp;&amp;amp;\r\ntar zxf swiftly-$(uname -m).tar.gz &amp;amp;&amp;amp;\r\n./swiftly init --quiet-shell-followup &amp;amp;&amp;amp;\r\n. &amp;quot;${SWIFTLY_HOME_DIR:-$HOME/.local/share/swiftly}/env.sh&amp;quot; &amp;amp;&amp;amp;\r\nhash -r&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e743ef3d0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Alternatively, you can download prebuilt toolchain tarballs directly from official &lt;/span&gt;&lt;a href="https://www.swift.org/download/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Swift Downloads&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for Ubuntu, Debian, Fedora, or Amazon Linux. Note that &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;google-cloud-swift&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; requires &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Swift 6.2 or later&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, so verify your compiler version with &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;swift --version&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; after installation.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To resolve &lt;/span&gt;&lt;a href="https://swift.org/package-manager/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Swift Package Manager&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; bare repository trust warnings when cloning across Linux filesystems, configure &lt;/span&gt;&lt;a href="https://git-scm.com" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Git&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; before building with &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;git config --global safe.bareRepository all&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Add the required packages to your &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;Package.swift&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; manifest:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;swift package add-dependency https://github.com/googleapis/swift-google-cloud-language-v2.git --from 0.4.0\r\nswift package add-target-dependency GoogleCloudLanguageV2 CloudBackendService --package swift-google-cloud-language-v2&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e743ef8d0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;On macOS you need to change the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;platforms&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; directive:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;// swift-tools-version: 6.2\r\nimport PackageDescription\r\n\r\nlet package = Package(\r\n  name: &amp;quot;CloudBackendService&amp;quot;,\r\n  // Applied when compiling on Darwin/macOS; ignored by SPM on Linux targets\r\n  platforms: [.macOS(.v15)],\r\n ... ...&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e743ed990&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;p data-block-key="bngnr"&gt;In most environments a default-initialized client can make requests:&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;import Foundation\r\nimport GoogleCloudLanguageV2\r\n\r\nfunc analyzeTextSentiment(text: String) async throws {\r\n  // Initialize explicit API key credentials\r\n  let client = try LanguageServiceClient()\r\n\r\n  // Configure request using structured builder closure\r\n  let document = Document().with {\r\n    $0.type = .plainText\r\n    $0.source = .content(text)\r\n  }\r\n\r\n  let response = try await client.analyzeSentiment(\r\n    request: AnalyzeSentimentRequest().with { $0.document = document }\r\n  )\r\n\r\n  if let sentiment = response.documentSentiment {\r\n    print(&amp;quot;Document sentiment score: \\(sentiment.score)&amp;quot;)\r\n  }\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e743efc10&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Notice how &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;Document().with { ... }&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; avoids verbose temporary variables or mutating setters by providing a clean, thread-safe configuration closure.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Networking, transport, and authentication&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The repository splits infrastructure primitives into modular packages under &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;packages/&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;code style="vertical-align: baseline;"&gt;swift-google-cloud-auth&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;: Implements &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/docs/authentication/application-default-credentials"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Application Default Credentials&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; discovery, service account JWT signing, external account exchange for Workload Identity Federation, and API keys.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;code style="vertical-align: baseline;"&gt;swift-google-cloud-wkt&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;: Provides idiomatic Swift types for Google Protocol Buffer well-known types, including nanosecond-precision &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;Timestamp&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; representations that bridge cleanly to Swift's &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;Date&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;code style="vertical-align: baseline;"&gt;swift-google-cloud-gax&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;: Handles Google API Extensions such as automated retry loops, exponential backoff, and pagination state machines.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When you initialize any client library without arguments, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;Credentials.default()&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; automatically scans your environment (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;GOOGLE_APPLICATION_CREDENTIALS&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, quota project variables, or the local Google Cloud CLI configuration) and authenticates connections over gRPC or HTTP/2.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If you need to programmatically override credentials with an API key or attach custom access headers, you can pass explicit configuration options. For example, you could modify the previous example to use the following:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;import Foundation\r\nimport GoogleCloudAuth\r\nimport GoogleCloudGax\r\nimport GoogleCloudLanguageV2\r\n\r\nfunc analyzeTextSentiment(apiKey: String, text: String) async throws {\r\n  // Initialize explicit API key credentials\r\n  let credentials = try Credentials(configuration: .apiKey(apiKey))\r\n  let client = try LanguageServiceClient(\r\n    ClientOptions().with { $0.credentials = credentials }\r\n  )\r\n\r\n  // Configure request using structured builder closure\r\n  let document = Document().with {\r\n    $0.type = .plainText\r\n    $0.source = .content(text)\r\n  }\r\n\r\n  let response = try await client.analyzeSentiment(\r\n    request: AnalyzeSentimentRequest().with { $0.document = document }\r\n  )\r\n\r\n  if let sentiment = response.documentSentiment {\r\n    print(&amp;quot;Document sentiment score: \\(sentiment.score)&amp;quot;)\r\n  }\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e743eff10&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;The autogenerated client ecosystem&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Google Cloud operates a vast ecosystem of APIs whose schemas update regularly. The teams supporting &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;google-cloud-swift&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; use code generators to automatically update the client libraries with the latest features and with new APIs. Using code generators produces stable APIs, without disruptive breaking changes. While the releases are on a fixed cadence, please contact Cloud Customer Care if you need a particular feature or API urgently.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Whether you need to rotate keys in &lt;/span&gt;&lt;a href="https://cloud.google.com/secret-manager"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Secret Manager&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; or invoke multimodal inference models via the &lt;/span&gt;&lt;a href="https://cloud.google.com/vertex-ai/generative-ai/docs/gemini-v1"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini API&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; on &lt;/span&gt;&lt;a href="https://cloud.google.com/vertex-ai"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise Agent Platform&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, the generated SDKs follow consistent naming and async/await signatures.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;These generated clients offer more than plain unary RPC wrappers. They also offer wrappers that simplify application development. For example, iterating over long results involves fetching pages of results with one RPC, iterating over the page of results, and then preparing a new request to retrieve the following page. Using the generated clients this becomes an asynchronous iterator. This example shows how to query project secrets using &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;GoogleCloudSecretManagerV1&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;import Foundation\r\nimport GoogleCloudSecretManagerV1\r\n\r\n@main\r\nstruct SecretManagerQuickstart {\r\n  static func main() async throws {\r\n    guard let projectId = CommandLine.arguments.dropFirst().first else {\r\n      print(&amp;quot;Usage: SecretManagerQuickstart &amp;lt;projectId&amp;gt;&amp;quot;)\r\n      exit(1)\r\n    }\r\n\r\n    // Connects using Application Default Credentials automatically\r\n    let client = try SecretManagerServiceClient()\r\n\r\n    let request = ListSecretsRequest().with {\r\n      $0.parent = &amp;quot;projects/\\(projectId)&amp;quot;\r\n    }\r\n\r\n    // Async sequence streams pages of secrets automatically\r\n    print(&amp;quot;Secrets in project \\(projectId):&amp;quot;)\r\n    for try await item in try client.listSecretsByItem(request: request) {\r\n      print(&amp;quot; - \\(item.name)&amp;quot;)\r\n    }\r\n  }\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e743ef410&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The pagination response returns an asynchronous sequence (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;AsyncSequence&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;). You can iterate over items with &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;for try await&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; while the client library fetches subsequent pages in the background over non-blocking NIO channels.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Where to go next&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;google-cloud-swift&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, server-side Swift developers can write end-to-end cloud infrastructure with compile-time race safety, native async/await ergonomic APIs, and zero OS thread congestion.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To inspect the source code, open issues, or contribute new veneers, visit the official repository at &lt;/span&gt;&lt;a href="https://github.com/googleapis/google-cloud-swift" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;googleapis/google-cloud-swift&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Want to discuss server-side Swift architectures or Cloud Run containerization? Join the &lt;/span&gt;&lt;a href="https://developers.google.com/program" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Developer Program&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to continue the conversation.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Thu, 01 Oct 2026 04:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/developers-practitioners/introducing-the-server-side-cloud-swift-sdk/</guid><category>Developers &amp; Practitioners</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/introducing-the-server-side-cloud-swift-sdk-.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Introducing the Server Side Cloud Swift SDK</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/introducing-the-server-side-cloud-swift-sdk-.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/developers-practitioners/introducing-the-server-side-cloud-swift-sdk/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Karl Weinmeister</name><title>Director, Developer Relations</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Carlos O'Ryan</name><title>Software Engineer</title><department></department><company></company></author></item><item><title>Empower your agents with the Google Cloud CLI remote MCP server</title><link>https://cloud.google.com/blog/products/ai-machine-learning/google-cloud-cli-remote-mcp-server-in-preview/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Today, we’re expanding our ecosystem of managed remote MCP servers by introducing the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/sdk/use-gcloud-mcp"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud CLI remote MCP server&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; in preview.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Powered by the popular &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/sdk/gcloud"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;gcloud&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/reference/bq-cli-reference"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;bq (BigQuery)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; command-line tools, this new server gives AI agents immediate, broad access to command-line operations for managing Google Cloud infrastructure and working with advanced BigQuery workflows securely and seamlessly. &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Why CLI matters for AI agents&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Agents are increasingly performing complex cloud operations, but standardizing how they interact with backend systems remains a challenge. The Google Cloud CLI remote MCP server bridges this gap by packaging the versatility of hundreds of gcloud and bq commands into one single MCP server. This results in two strong benefits for the agent:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Higher-level abstractions:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; CLI commands package complex multi-step workflows, validation checks, and high-level operations into unified commands rather than requiring multi-step API orchestration.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Leverages model training:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; LLMs are heavily pre-trained on public command-line documentation, syntaxes, and usage examples, making CLI invocation intuitive and highly accurate for models.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;The benefits of putting CLI behind remote MCP&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Managing cloud infrastructure with AI agents traditionally requires installing and maintaining Google Cloud CLI binaries inside agent execution environments. The Cloud CLI remote MCP server bridges CLI capabilities with MCP benefits by providing an isolated execution sandbox on Google Cloud infrastructure. This solves key infrastructure challenges:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Simplified dependency and runtime management:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; For teams building custom agents, maintaining local CLI versions and dependencies across dev, test, and production environments creates operational overhead. Remote MCP eliminates local installations and runtime maintenance.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Access for web-based agent endpoints:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Web-hosted agent platforms and web interfaces (such as Gemini Enterprise and other hosted enterprise agent platforms) run in environments where users cannot control or install local packages. Remote MCP enables secure, managed access to Google Cloud CLI operations directly from these surfaces.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Enterprise-grade security and governance&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Connecting an AI agent to your infrastructure requires strict, enterprise-ready safeguards. This remote server leverages Google Cloud's standard identity and governance frameworks to keep your environments secure:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Zero ambient credentials:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The server isolates execution in a network-restricted proxy boundary with no ambient credentials. Authentication and authorization are handled through &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/iam/docs/agent-identity-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Agent Identity&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://developers.google.com/identity/protocols/oauth2" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;OAuth 2.0&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/iam/docs"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Identity and Access Management (IAM)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Strict policy enforcement:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Every command executed through the remote MCP server is run with the permissions of the authenticated caller identity. Both standard IAM permissions and organization policy service constraints are strictly enforced against downstream target resources.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Advanced protection with Model Armor:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; To minimize the risks associated with AI tool calling, the Cloud CLI remote MCP server integrates with &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/model-armor/model-armor-mcp-google-cloud-integration"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Model Armor&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. You can proactively screen LLM prompts and responses to protect against risks like prompt injection and malicious inputs.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Cloud audit logging:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The Cloud CLI remote MCP server can be configured to log every tool invocation to Audit Logs (Data Access logs under cloudcli.googleapis.com/mcp). Security teams can gain full visibility into caller identities, OAuth clients, and IAM authorization decisions (mcp.googleapis.com/tools.call) without exposing sensitive command payloads or personally identifiable information (PII).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Connecting to the Google Cloud CLI Remote MCP Server&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Integrating cloud management into your agents no longer requires packaging Google Cloud CLI binaries, managing local execution runtimes, or maintaining dependencies inside agent container images. Because the Google Cloud CLI remote MCP server implements the standard Model Context Protocol, any MCP-compatible agent platform or orchestration runtime can connect immediately via standard configuration:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;{\r\n  &amp;quot;mcpServers&amp;quot;: {\r\n    &amp;quot;google-cloud-cli&amp;quot;: {\r\n      &amp;quot;uri&amp;quot;: &amp;quot;https://cloudcli.googleapis.com/mcp&amp;quot;,\r\n      ...\r\n    }\r\n  }\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e743ee850&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Your agent immediately gains access to execute gcloud and bq commands in a secure, network-isolated cloud sandbox.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Authentication is handled via keyless Agent Identity for hosted Google Cloud platforms, or standard OAuth 2.0 for external runtimes. For authentication options, see the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/mcp/set-up-authentication-mcp-servers"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;MCP Authentication Guide&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Bringing infrastructure management to agents&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;a href="https://docs.cloud.google.com/sdk/use-gcloud-mcp"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;The Cloud CLI remote MCP server&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; exposes two powerful tools, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;run_gcloud_command&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;run_bq_command&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, giving your AI agents broad, immediate access to Google Cloud operations through natural language.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Managing cloud infrastructure with &lt;/strong&gt;&lt;code&gt;&lt;strong style="vertical-align: baseline;"&gt;run_gcloud_command&lt;/strong&gt;&lt;/code&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;run_gcloud_command&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, agents can execute the full breadth of gcloud operations to manage, diagnose, and secure your Google Cloud environment. An example follows:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Observability and incident diagnostics&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: An agent streamlines incident diagnostics by automating command execution and reducing context-switching across tools.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/1_IPA9ALa.gif"
        
          alt="1"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Extending BigQuery operations with the run_bq_command tool&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;While the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/bigquery/docs/use-bigquery-mcp"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;BigQuery MCP server&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; already helps organizations analyze and explore data using AI agents, with the introduction of &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;run_bq_command&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, agents can now tackle advanced BigQuery tasks such as resource allocation, job monitoring, and task scheduling by unlocking the full scope of &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/sdk/reference/mcp#mcp-tools"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;bq CLI&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; functionality. Key capabilities include:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Automating scheduled queries&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: An agent utilizes BigQuery Data Transfer Service configurations to schedule queries automatically. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Job and resource management&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Gain deep insight into query execution details, including processed data volume, slot usage, and execution plans, as well as managing reservations.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Access and permissions control:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Data administrators and owners can inspect and update table permissions directly through the agent.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/2_JA5HWbH.gif"
        
          alt="2"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Pricing and availability&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The Google Cloud CLI MCP server is available today in public preview. There is no additional charge to use the MCP server itself. You pay only for the GCP resources you create and any applicable data transfer costs.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://docs.cloud.google.com/sdk/use-gcloud-mcp"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;To get started&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, enable the Cloud CLI Execution API (`&lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;cloudcli.googleapis.com&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;`) in your Google Cloud project, grant the required &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;MCP Tool User&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; (`&lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;roles/mcp.toolUser&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;`) IAM role to your agent or user identity, and configure your MCP client to connect to `cloudcli.googleapis.com/mcp`.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/sdk/use-gcloud-mcp"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Use the Google Cloud CLI Remote MCP Server Guide&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/sdk/reference/mcp#mcp-tools"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud CLI MCP Reference&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/mcp/overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Remote MCP Servers Overview&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/mcp/supported-products"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Full List of Google OneMCP Servers&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/mcp/set-up-authentication-mcp-servers"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Set up authentication to Google and Google Cloud MCP servers&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/gemini-enterprise-agent-platform/govern/agent-identity-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise Agent Platform &amp;amp; Agent Identity Overview&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://www.youtube.com/watch?v=-fb0ycu4kiU" rel="noopener" target="_blank"&gt;&lt;span data-rich-links='{"fple-t":"Automate Google Cloud with Cloud CLI Remote MCP Server","fple-u":"https://www.youtube.com/watch?v=-fb0ycu4kiU","fple-mt":null,"type":"first-party-link"}' style="text-decoration: underline; vertical-align: baseline;"&gt;Automate Google Cloud with Cloud CLI Remote MCP Server&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Wed, 30 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/ai-machine-learning/google-cloud-cli-remote-mcp-server-in-preview/</guid><category>Data Analytics</category><category>Developers &amp; Practitioners</category><category>AI &amp; Machine Learning</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Empower your agents with the Google Cloud CLI remote MCP server</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/ai-machine-learning/google-cloud-cli-remote-mcp-server-in-preview/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Prosper Nwankpa</name><title>Senior Engineering Manager, Google Cloud</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Adam Hwang</name><title>Software Engineering Manager, Google Cloud</title><department></department><company></company></author></item><item><title>Data Agent Kit is now GA: Bring Google Data Cloud to any coding agent</title><link>https://cloud.google.com/blog/topics/developers-practitioners/data-agent-kit-is-now-ga-bring-google-data-cloud-to-any-coding-agent/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Today, &lt;/span&gt;&lt;a href="https://cloud.google.com/products/data-agent-kit?utm_campaign=CDR_0xaea1deef_default_b566338695&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Data Agent Kit&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; is generally available. Data Agent Kit is a free set of Model Context Protocol (MCP) tools and agent skills that lets the coding agent you already use work directly with your Google Cloud data products, whether you're using Antigravity, Claude Code, Codex, or other popular tools.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With GA, we are adding support for &lt;/span&gt;&lt;a href="https://cloud.google.com/bigquery/docs/graph-overview?utm_campaign=CDR_0xaea1deef_default_b566338695&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;BigQuery Graph&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/bigtable?utm_campaign=CDR_0xaea1deef_default_b566338695&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Bigtable&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and &lt;/span&gt;&lt;a href="https://cloud.google.com/dataproc-serverless/docs/overview?utm_campaign=CDR_0xaea1deef_default_b566338695&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Managed Service for Apache Spark&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; access to your open Lakehouse, along with dozens of quality-of-life improvements that make everyday work faster and smoother.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/00-data-agent-kit-is-now-ga-bring-google-data-cloud-to-any-coding-agent-da.gif"
        
          alt="00-data-agent-kit-is-now-ga-bring-google-data-cloud-to-any-coding-agent-dak-promo"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;What is Data Agent Kit?&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Coding agents have become remarkably good at writing SQL, PySpark, and pipeline code. What they don't have by default is context about your environment: which tables exist, how they're partitioned, which ones your team trusts, or why last night's job failed. Without that, even a strong agent has to work from assumptions, and you end up pasting schemas and error logs into the chat to fill in the gaps.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Data Agent Kit fills that gap with two things:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;MCP tools:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Connections to more than 15 Google Data Cloud services, so your agent can inspect schemas, run queries, read job logs, and manage resources in your live environment.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Google-authored skills:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://github.com/GoogleCloudPlatform/data-agent-kit-plugin" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Open-source instructions&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; from Google Cloud engineers that teach your agent data best practices, like optimizing BigQuery SQL, designing Bigtable row keys, and building dbt (&lt;/span&gt;&lt;a href="https://www.getdbt.com/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;data build tool&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;) pipelines.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;You can use Data Agent Kit wherever you already work: as an IDE extension for VS Code, Antigravity IDE, Cursor, and other VS Code-compatible editors; as a plugin for Antigravity 2.0, Antigravity CLI, Claude Code, and Codex; or in &lt;/span&gt;&lt;a href="https://cloud.google.com/shell?utm_campaign=CDR_0xaea1deef_default_b566338695&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Shell&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://cloud.google.com/workstations?utm_campaign=CDR_0xaea1deef_default_b566338695&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Workstations&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, where it comes pre-installed. The IDE extension also brings a lightweight version of the Google Cloud console into your editor, so you can browse data, run queries, and review your agent's work without switching windows.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;How it works&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Say you ask your agent, "Forecast next month's demand for our top-selling products and check whether we have enough inventory to meet it." Data Agent Kit loads the relevant skills, so the agent follows Google's best practices for the task. It searches &lt;/span&gt;&lt;a href="https://cloud.google.com/dataplex/docs/introduction?utm_campaign=CDR_0xaea1deef_default_b566338695&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Knowledge Catalog&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to find the sales and inventory tables your team trusts. It then uses MCP tools to run a forecast in &lt;/span&gt;&lt;a href="https://cloud.google.com/bigquery?utm_campaign=CDR_0xaea1deef_default_b566338695&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;BigQuery&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, check current stock levels in AlloyDB for PostgreSQL, and bring the combined answer back to your editor or terminal. Every step runs with your own IAM permissions.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/01-data-agent-kit-is-now-ga-bring-google-d.max-1000x1000.png"
        
          alt="01-data-agent-kit-is-now-ga-bring-google-data-cloud-to-any-coding-agent_how_it_works"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="qngqb"&gt;How Data Agent Kit connects to data.&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;What you can build&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Data Agent Kit covers analytics, operational databases, the Lakehouse, and pipelines. Here's what that looks like in practice, starting with what's new at GA.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Analytics and graph: BigQuery and BigQuery Graph (New in GA)&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Graphs are a natural way to explore relationships, like which products people buy together or how suppliers connect to your inventory. Building one usually means hand-writing &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;CREATE PROPERTY GRAPH&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; DDL, learning GQL, and working out which keys actually form edges.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Instead, you describe the graph you want and your agent builds it. The &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;bigquery-graph-author&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; skill maps your tables to nodes and edges, checks each proposed relationship against the actual data, and shows you a plan to approve before creating anything. It can even start from an ER diagram or data model you already have. The &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;bigquery-graph-query&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; skill then writes the GQL, and the graph visualizer in the IDE lets you click through the results.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/02-data-agent-kit-is-now-ga-bring-google-data-cloud-to-any-coding-agent_bq.gif"
        
          alt="02-data-agent-kit-is-now-ga-bring-google-data-cloud-to-any-coding-agent_bqgraph"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="qngqb"&gt;Building and visualizing a BigQuery property graph.&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Operational and real-time databases: Spanner, AlloyDB, Cloud SQL, and Bigtable (New in GA)&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Some features have to load instantly, like a personalized feed, a live counter, or a "recently viewed" rail on your storefront. Bigtable is built for exactly that, and it rewards a well-designed row key.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With GA, Bigtable joins &lt;/span&gt;&lt;a href="https://cloud.google.com/spanner?utm_campaign=CDR_0xaea1deef_default_b566338695&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Spanner&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://cloud.google.com/alloydb?utm_campaign=CDR_0xaea1deef_default_b566338695&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;AlloyDB&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and &lt;/span&gt;&lt;a href="https://cloud.google.com/sql?utm_campaign=CDR_0xaea1deef_default_b566338695&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud SQL&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; as a fully supported database in Data Agent Kit. The new &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;bigtable-basics&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; skill designs your schema around how the data will be read and flags hotspots and full table scans before you create anything. Your agent can then create the table and query it with GoogleSQL, with column families flattened into readable columns. In the IDE, you can browse Bigtable instances and tables in the catalog explorer and run queries from the SQL editor.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/03-data-agent-kit-is-now-ga-bring-google-data-cloud-to-any-coding-agent_bi.gif"
        
          alt="03-data-agent-kit-is-now-ga-bring-google-data-cloud-to-any-coding-agent_bigtable"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="qngqb"&gt;Agent-guided Bigtable schema design and querying.&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Lakehouse and Spark: Managed Service for Apache Spark and Lakehouse for Apache Iceberg (New in GA)&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Your agent could already query the Apache Iceberg tables in your Lakehouse through BigQuery. Now it can also work with those same tables using serverless Spark on &lt;/span&gt;&lt;a href="https://cloud.google.com/dataproc-serverless/docs/overview?utm_campaign=CDR_0xaea1deef_default_b566338695&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Managed Service for Apache Spark&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, with no cluster to manage. Each session keeps its state, so temporary views carry across statements, and you get Iceberg's full feature set, including branching, time travel, and schema evolution.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Your tables don't all have to live on Google Cloud, either. The &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;federate-lakehouse-catalog&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; skill connects your Lakehouse to AWS Glue and Databricks Unity Catalog, so your agent can query that data in place without building an ingestion pipeline first.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/04-data-agent-kit-is-now-ga-bring-google-d.max-1000x1000.png"
        
          alt="04-data-agent-kit-is-now-ga-bring-google-data-cloud-to-any-coding-agent_lakehouse"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="qngqb"&gt;Branching and querying Iceberg tables with Managed Service for Apache Spark.&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Pipelines and orchestration: dbt, Dataform, and Managed Service for Apache Airflow&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Once your logic works, your agent can turn it into a pipeline that runs on its own. It writes dbt or &lt;/span&gt;&lt;a href="https://cloud.google.com/dataform?utm_campaign=CDR_0xaea1deef_default_b566338695&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Dataform&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; models, then the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;gcp-pipeline-orchestration&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; skill schedules them together with your notebooks as an Orchestration Pipeline on &lt;/span&gt;&lt;a href="https://cloud.google.com/composer?utm_campaign=CDR_0xaea1deef_default_b566338695&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Managed Service for Apache Airflow&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. Pipelines can also include Gemini Enterprise Agent Platform steps, like uploading a model or running batch inference.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In the IDE, you can follow each run on a visual pipeline canvas. If a task needs attention, click &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Diagnose&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; to hand its logs to your agent. Troubleshooting skills for Airflow and Spark trace the root cause and propose a fix for you to approve.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/05-data-agent-kit-is-now-ga-bring-google-data-cloud-to-any-coding-agent_pi.gif"
        
          alt="05-data-agent-kit-is-now-ga-bring-google-data-cloud-to-any-coding-agent_pipelines"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="qngqb"&gt;Troubleshooting a failed Airflow DAG with an agent.&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Improving the developer experience&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;GA also streamlines setup and day-to-day workflows. When you get started, you simply sign in once and select the Google Cloud services you use. Data Agent Kit automatically enables the required APIs, installs the matching skills, and configures your MCP servers with no manual setup files. Inside the IDE, the SQL editor and notebooks now support inline code generation, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;@&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; references to tables, and diff views for suggested changes. The extension also shares your active project, open file, and the error from the query you just ran with your agent, so asking it to "fix this query" just works.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We also made the core tools faster and more responsive. New Spark notebooks automatically create and select a Spark Connect runtime, and Spark SQL queries in the editor run in isolated sessions with built-in execution metrics. &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;The catalog explorer now loads faster and includes BigQuery public datasets in the sidebar. Tuned notebook skills help your agent finish notebook tasks more quickly while using fewer tokens. &lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Ready for the enterprise&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Data Agent Kit is included at no additional cost; you pay standard pricing only for the Google Cloud services your agent uses. Skills also steer the agent toward cost-aware query patterns, like checking partition keys and running a dry run before executing a BigQuery query to avoid accidental full-table scans. And because the skills are open source on GitHub, your team can audit them, fork them, or write custom skills for your own internal standards.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For access control, the agent connects as you or as a service account you impersonate, so row- and column-level security policies apply automatically. Admins can also govern MCP access with Identity and Access Management (IAM), screen MCP traffic with &lt;/span&gt;&lt;a href="https://cloud.google.com/security-command-center/docs/model-armor-overview?utm_campaign=CDR_0xaea1deef_default_b566338695&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Model Armor&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, and scope agents with &lt;/span&gt;&lt;a href="https://cloud.google.com/vpc-service-controls?utm_campaign=CDR_0xaea1deef_default_b566338695&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;VPC Service Controls&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and Principal Access Boundary policies.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Install and get started&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;You can set up Data Agent Kit in under a minute in either your IDE or your terminal.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Option 1: Install the IDE Extension&lt;/span&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;IDE extension:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Search for "Google Cloud Data Agent Kit" in the Extensions panel of VS Code, Antigravity IDE, Cursor, or any VS Code-compatible editor. You can also install it from the &lt;/span&gt;&lt;a href="https://marketplace.visualstudio.com/items?itemName=googlecloudtools.datacloud" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;VS Code Marketplace&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; or &lt;/span&gt;&lt;a href="https://open-vsx.org/extension/googlecloudtools/datacloud" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Open VSX&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Antigravity 2.0:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Go to &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Settings &amp;gt; Customizations &amp;gt; Build with Google Plugins&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, then download Data Agent Kit.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Cloud Shell and Cloud Workstations:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Already installed by default; just open the editor and sign in.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Option 2: Install the CLI Plugin&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Run the command for your preferred coding agent using the official &lt;/span&gt;&lt;a href="https://github.com/GoogleCloudPlatform/data-agent-kit-plugin" rel="noopener" target="_blank"&gt;&lt;code style="text-decoration: underline; vertical-align: baseline;"&gt;GoogleCloudPlatform/data-agent-kit-plugin&lt;/code&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; repository:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;# Antigravity CLI\r\nagy plugin install https://github.com/GoogleCloudPlatform/data-agent-kit-plugin\r\n\r\n# Claude Code\r\nclaude plugin install data-agent-kit-starter-pack@claude-plugins-official\r\n\r\n# Codex CLI\r\ncodex plugin marketplace add GoogleCloudPlatform/data-agent-kit-plugin\r\ncodex plugin add dak@dak-marketplace&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e6767e7d0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Then try a first prompt:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;quot;What are this week&amp;#x27;s fastest rising search terms in the US that weren&amp;#x27;t in the last week&amp;#x27;s top 10? Use the BigQuery public Google Trends dataset.&amp;quot;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e67a98c90&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/06-data-agent-kit-is-now-ga-bring-google-data-cloud-to-any-coding-agent_cl.gif"
        
          alt="06-data-agent-kit-is-now-ga-bring-google-data-cloud-to-any-coding-agent_cli"&gt;
        
        &lt;/a&gt;
      
        &lt;figcaption class="article-image__caption "&gt;&lt;p data-block-key="qngqb"&gt;Data Agent Kit plugin running in Claude Code.&lt;/p&gt;&lt;/figcaption&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Next steps &amp;amp; hands-on resources&lt;/span&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Read the docs:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Explore the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/data-agent-kit?utm_campaign=CDR_0xaea1deef_default_b566338695&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Data Agent Kit documentation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://cloud.google.com/products/data-agent-kit?utm_campaign=CDR_0xaea1deef_default_b566338695&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;product overview page&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Explore the skills:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Browse, star, and contribute on &lt;/span&gt;&lt;a href="https://github.com/GoogleCloudPlatform/data-agent-kit-plugin" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GitHub&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Build an analytics workflow:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Try the &lt;/span&gt;&lt;a href="https://codelabs.developers.google.com/dak-analytics-eng-antigravity-ide?utm_campaign=CDR_0xaea1deef_default_b566338695&amp;amp;utm_medium=external&amp;amp;utm_source=blog" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Analytics with Data Agent Kit and Antigravity IDE&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; codelab, and read the companion blog, &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/data-analytics/agentic-analytics-with-the-data-agent-kit?utm_campaign=CDR_0xaea1deef_default_b566338695&amp;amp;utm_medium=external&amp;amp;utm_source=blog"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Agentic analytics with the Data Agent Kit&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Build a data science pipeline:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Try the &lt;/span&gt;&lt;a href="https://codelabs.developers.google.com/dak-data-science-antigravity-ide?utm_campaign=CDR_0xaea1deef_default_b566338695&amp;amp;utm_medium=external&amp;amp;utm_source=blog" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Fraud detection pipeline with Data Agent Kit and Antigravity IDE&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; codelab.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Wed, 30 Sep 2026 13:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/developers-practitioners/data-agent-kit-is-now-ga-bring-google-data-cloud-to-any-coding-agent/</guid><category>Developers &amp; Practitioners</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/dak_blog_banner.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Data Agent Kit is now GA: Bring Google Data Cloud to any coding agent</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/dak_blog_banner.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/developers-practitioners/data-agent-kit-is-now-ga-bring-google-data-cloud-to-any-coding-agent/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Arun Nair</name><title>Product Manager, Google</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Jeff Nelson</name><title>Developer Advocate, Google</title><department></department><company></company></author></item><item><title>Graph Workflows in ADK: Everything You Need to Know</title><link>https://cloud.google.com/blog/topics/developers-practitioners/graph-workflows-in-adk-everything-you-need-to-know/</link><description>&lt;div class="block-paragraph"&gt;&lt;p data-block-key="z0o0k"&gt;Graph engineering is the design work: breaking a task into nodes, connecting them with edges, and deciding where code, models, or people control the next step. The &lt;a href="https://adk.dev/" target="_blank"&gt;Agent Development Kit (ADK)&lt;/a&gt;'s Workflow turns that design into an executable process, with functions and agents doing the work. Through a refund example, this post shows how to run steps in parallel, route decisions, pause for human review, and process a list of cases. It also explains when to declare the paths in a static graph and when to let Python schedule further work as results arrive.&lt;/p&gt;&lt;p data-block-key="2irer"&gt;&lt;b&gt;TL;DR:&lt;/b&gt; Using a refund workflow in ADK, we'll cover fan-out and fan-in, deterministic and agent routers, human-in-the-loop pauses, parallel workers, and dynamic orchestration— along with when to use a static graph or let Python decide what runs next.&lt;/p&gt;&lt;h2 data-block-key="611o0"&gt;&lt;b&gt;Start with a single agent&lt;/b&gt;&lt;/h2&gt;&lt;p data-block-key="d27bu"&gt;We can give one agent the tools and instructions to handle the refund request from start to finish:&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;refund_agent = Agent(\r\n    name=&amp;quot;refund_agent&amp;quot;, model=MODEL,\r\n    tools=[fetch_order, fetch_payment, fetch_history],\r\n    instruction=&amp;quot;&amp;quot;&amp;quot;You handle refund requests.\r\n    1. Look up the order.\r\n    2. Check the payment record.\r\n    3. Check the customer\&amp;#x27;s refund history.\r\n    4. Deny if there\&amp;#x27;s an open chargeback or it\&amp;#x27;s past 30 days.\r\n       Approve if it\&amp;#x27;s under $50 and they\&amp;#x27;ve had fewer than three\r\n       refunds this year.\r\n    5. Write the customer an email explaining the decision.&amp;quot;&amp;quot;&amp;quot;,\r\n)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;lang-py&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e67cdf790&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;p data-block-key="z0o0k"&gt;This puts the model in charge of choosing the tools, applying the policy, and writing the reply. But the prompt already describes distinct pieces of work: three lookups, a decision, and a response. Making those pieces separate nodes lets us decide how each should run.&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/graph-workflows-in-adk-everything-you-need.max-1000x1000_wEZYObA.png"
        
          alt="graph-workflows-in-adk-everything-you-need-to-know-01-five-nodes"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;p data-block-key="z0o0k"&gt;Assume the customer has selected an order and clicked “Request refund.” The app knows the order ID, so the workflow can begin with the lookups.&lt;/p&gt;&lt;h2 data-block-key="2v81e"&gt;&lt;b&gt;Let independent steps run together&lt;/b&gt;&lt;/h2&gt;&lt;p data-block-key="1ca48"&gt;The order, payment, and refund-history lookups all need the order ID, but none needs another lookup's result. Although the prompt lists them one after another, there is no reason for them to wait for each other. We can run all three in parallel.&lt;/p&gt;&lt;p data-block-key="6rgi5"&gt;The policy decision is different: it needs all three records. So the workflow splits into three paths, then brings their results together before continuing. These two moves are called &lt;b&gt;fan-out&lt;/b&gt; and &lt;b&gt;fan-in&lt;/b&gt;.&lt;/p&gt;&lt;p data-block-key="4am57"&gt;In ADK, the lookup functions can become nodes directly. A nested tuple starts them together, and a &lt;code&gt;JoinNode&lt;/code&gt; waits for their results. First, the imports and a lookup signature:&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;import asyncio\r\nfrom pydantic import BaseModel, Field\r\n\r\nfrom google.adk import Agent, Context, Event, Workflow\r\nfrom google.adk.events import RequestInput\r\nfrom google.adk.workflow import JoinNode, START, node\r\n\r\nasync def fetch_order(node_input: str) -&amp;gt; dict:\r\n    ...&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;lang-py&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e6710a310&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;p data-block-key="z0o0k"&gt;Each lookup receives the order ID through &lt;code&gt;node_input&lt;/code&gt; and returns a dictionary. We can connect them like this:&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;join_case = JoinNode(name=&amp;quot;join_case&amp;quot;)\r\n\r\nedges=[\r\n    (START, (fetch_order, fetch_payment, fetch_history), join_case),\r\n]&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;lang-py&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e6710b310&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;p data-block-key="z0o0k"&gt;Read this from left to right: start all three lookups, then continue through &lt;code&gt;join_case&lt;/code&gt; once they finish.&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/graph-workflows-in-adk-everything-you-need.max-1000x1000_59cATcj.png"
        
          alt="graph-workflows-in-adk-everything-you-need-to-know-02-fanout-join"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;p data-block-key="z0o0k"&gt;The join returns a dictionary keyed by node name. The next node can read &lt;code&gt;node_input["fetch_order"]&lt;/code&gt;, &lt;code&gt;node_input["fetch_payment"]&lt;/code&gt;, and &lt;code&gt;node_input["fetch_history"]&lt;/code&gt; without a model call to collect the results. The &lt;a href="https://github.com/google/adk-python/blob/main/docs/guides/workflow/join_node/index.md" target="_blank"&gt;JoinNode guide&lt;/a&gt; explains the details.&lt;/p&gt;&lt;p data-block-key="a65ee"&gt;If the payment lookup needed a transaction ID from the order lookup, those two would run in sequence. &lt;b&gt;Dependencies determine the edges&lt;/b&gt;, even when the prompt lists every step in order.&lt;/p&gt;&lt;p data-block-key="5cu7f"&gt;You can learn more about fan out and fan in in this video:&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-video"&gt;



&lt;div class="article-module article-video "&gt;
  &lt;figure&gt;
    &lt;a class="h-c-video h-c-video--marquee"
      href="https://youtube.com/watch?v=Mzr7byMFy_4"
      data-glue-modal-trigger="uni-modal-Mzr7byMFy_4-"
      data-glue-modal-disabled-on-mobile="true"&gt;

      
        &lt;img src="//img.youtube.com/vi/Mzr7byMFy_4/maxresdefault.jpg"
             alt="Video about how to build a production-ready, multi-agent AI system from scratch using Graph Engineering and Google&amp;#x27;s Agent Development Kit (ADK)."/&gt;
      
      &lt;svg role="img" class="h-c-video__play h-c-icon h-c-icon--color-white"&gt;
        &lt;use xlink:href="#mi-youtube-icon"&gt;&lt;/use&gt;
      &lt;/svg&gt;
    &lt;/a&gt;

    
      &lt;figcaption class="article-video__caption h-c-page"&gt;
        
          &lt;h4 class="h-c-headline h-c-headline--four h-u-font-weight-medium h-u-mt-std"&gt;Graph Engineering with ADK&lt;/h4&gt;
        
        
      &lt;/figcaption&gt;
    
  &lt;/figure&gt;
&lt;/div&gt;

&lt;div class="h-c-modal--video"
     data-glue-modal="uni-modal-Mzr7byMFy_4-"
     data-glue-modal-close-label="Close Dialog"&gt;
   &lt;a class="glue-yt-video"
      data-glue-yt-video-autoplay="true"
      data-glue-yt-video-height="99%"
      data-glue-yt-video-vid="Mzr7byMFy_4"
      data-glue-yt-video-width="100%"
      href="https://youtube.com/watch?v=Mzr7byMFy_4"
      ng-cloak&gt;
   &lt;/a&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;p data-block-key="z0o0k"&gt;One detail matters when turning tools into nodes: the parameter named &lt;code&gt;node_input&lt;/code&gt; receives the previous node's output. Other names bind to &lt;code&gt;ctx.state&lt;/code&gt; by default, so &lt;code&gt;order_id&lt;/code&gt; would look for &lt;code&gt;ctx.state["order_id"]&lt;/code&gt; and raise a &lt;code&gt;ValueError&lt;/code&gt; if it is missing. Here, the str annotation also converts &lt;code&gt;START&lt;/code&gt;'s &lt;code&gt;types.Content&lt;/code&gt; input to a string.&lt;/p&gt;&lt;h2 data-block-key="1og89"&gt;&lt;b&gt;Route each request to the right workflow&lt;/b&gt;&lt;/h2&gt;&lt;p data-block-key="1mldp"&gt;So far, the customer has explicitly requested a refund. In a broader support conversation, we first need to identify what they want and send the request to the right process. “The shoes are the wrong size. Could you send me a different pair?” should go to an exchange workflow, while a request for money back should enter our refund workflow.&lt;/p&gt;&lt;p data-block-key="82plm"&gt;That introduces a &lt;b&gt;router&lt;/b&gt;, a node that chooses which branch runs next.&lt;/p&gt;&lt;ul&gt;&lt;li data-block-key="duqvc"&gt;A &lt;b&gt;deterministic router&lt;/b&gt; follows explicit rules, the same inputs produce the same route.&lt;/li&gt;&lt;li data-block-key="27pqb"&gt;A &lt;b&gt;nondeterministic router&lt;/b&gt; can choose different branches for the same input.&lt;/li&gt;&lt;li data-block-key="fe33q"&gt;An &lt;b&gt;agent router&lt;/b&gt; uses a model to interpret the request, so its choice can vary.&lt;/li&gt;&lt;/ul&gt;&lt;p data-block-key="cn1a2"&gt;Routers choose among the paths defined by the workflow. In our support example, we can use an agent to identify the customer's intent, then fixed rules to apply the refund policy.&lt;/p&gt;&lt;h3 data-block-key="3tth6"&gt;&lt;b&gt;Identify the intent with an agent router&lt;/b&gt;&lt;/h3&gt;&lt;p data-block-key="40a6e"&gt;An agent can interpret the customer's message and classify the request. In one ADK pattern, it returns a structured category, then a small function emits the corresponding &lt;code&gt;Event(route=...)&lt;/code&gt; to send the request to the chosen workflow:&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;Customer message → classification agent → route function\r\n                                           ├─ REFUND → refund workflow\r\n                                           ├─ EXCHANGE → exchange workflow\r\n                                           └─ CLARIFY → ask a follow-up question&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e6710bd90&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;p data-block-key="z0o0k"&gt;The model identifies the intent; the graph defines the available destinations. If the request is unclear, the workflow can ask a follow-up question. ADK's &lt;a href="https://github.com/google/adk-python/blob/main/contributing/samples/workflows/route/agent.py" target="_blank"&gt;routing sample&lt;/a&gt; shows this pattern.&lt;/p&gt;&lt;p data-block-key="d6j2d"&gt;This intent router would sit before our refund workflow. For the selected-order example, the “Request refund” button has already established the intent, so we can enter that workflow directly.&lt;/p&gt;&lt;h3 data-block-key="4h73u"&gt;&lt;b&gt;Apply the refund policy with a deterministic router&lt;/b&gt;&lt;/h3&gt;&lt;p data-block-key="49cfk"&gt;Inside the refund workflow, the three lookups give us the facts for another decision: approve, deny, or ask a person to review. This time, the policy gives us explicit thresholds, so a function can choose the path:&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;def refund_policy(amount, days_ago, chargeback_open, priors):\r\n    if chargeback_open or days_ago &amp;gt; 30:\r\n        return &amp;quot;DENY&amp;quot;\r\n    if amount &amp;lt; 50 and priors &amp;lt; 3:\r\n        return &amp;quot;AUTO_APPROVE&amp;quot;\r\n    return &amp;quot;MANUAL_REVIEW&amp;quot;\r\n\r\ndef route_refund(node_input):\r\n    case = {**node_input[&amp;quot;fetch_order&amp;quot;], **node_input[&amp;quot;fetch_payment&amp;quot;],\r\n            **node_input[&amp;quot;fetch_history&amp;quot;]}\r\n    route = refund_policy(case[&amp;quot;amount_usd&amp;quot;], case[&amp;quot;placed_days_ago&amp;quot;],\r\n                          case[&amp;quot;chargeback_open&amp;quot;], case[&amp;quot;prior_refunds_12mo&amp;quot;])\r\n    return Event(output=case, route=route)        # the function names the path&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;lang-py&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e6710be90&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;code&gt;route_refund&lt;/code&gt; combines the records and returns the case with one of three route names: &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;AUTO_APPROVE&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;DENY&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, or &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;MANUAL_REVIEW&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;. Manual review handles cases that meet neither automatic rule.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This is a &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;deterministic router&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: the same case data produces the same decision, and we can test the policy without calling a model.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt; &lt;/p&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;&lt;table border="1" style="border-collapse: collapse; width: 100%;"&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;td style="width: 31.4907%;"&gt; &lt;/td&gt;
&lt;td style="width: 31.4907%;"&gt;&lt;strong&gt;Deterministic router&lt;/strong&gt;&lt;/td&gt;
&lt;td style="width: 31.4907%;"&gt;&lt;strong&gt;&lt;span style="vertical-align: baseline;"&gt;Agent router&lt;/span&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="width: 31.4907%;"&gt;&lt;strong&gt;Decisions come best from&lt;/strong&gt;&lt;/td&gt;
&lt;td style="width: 31.4907%;"&gt;Rules in code&lt;/td&gt;
&lt;td style="width: 31.4907%;"&gt;A model interpreting the input&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="width: 31.4907%;"&gt;&lt;strong&gt;Best fit&lt;/strong&gt;&lt;/td&gt;
&lt;td style="width: 31.4907%;"&gt;Known facts and explicit policy&lt;/td&gt;
&lt;td style="width: 31.4907%;"&gt;Meaning that is hard to capture in rules&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="width: 31.4907%;"&gt;&lt;strong&gt;Example&lt;/strong&gt;&lt;/td&gt;
&lt;td style="width: 31.4907%;"&gt;Deny an order older than 30 days&lt;/td&gt;
&lt;td style="width: 31.4907%;"&gt;Recognize an exchange request&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="width: 31.4907%;"&gt;&lt;strong&gt;Model call for routing&lt;/strong&gt;&lt;/td&gt;
&lt;td style="width: 31.4907%;"&gt;None&lt;/td&gt;
&lt;td style="width: 31.4907%;"&gt;Required&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;div align="left"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt; &lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Both routers choose among defined paths. &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;An agent router can sit inside a static graph&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: the model's choice varies, while the possible connections stay the same.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For an automatic approval or denial, the next step is to explain the decision to the customer. We give each route a notice agent that writes the reply from the case data. ADK passes the dictionary in &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;Event(output=case, ...)&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; to that agent as a JSON user message:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;approve_notice = Agent(\r\n    name=&amp;quot;approve_notice&amp;quot;, model=MODEL,\r\n    instruction=&amp;quot;Tell the customer their refund is approved and when to expect the &amp;quot;\r\n                &amp;quot;money, using the case JSON you receive. Short email, warm, no fluff.&amp;quot;,\r\n)\r\ndenial_notice = Agent(\r\n    name=&amp;quot;denial_notice&amp;quot;, model=MODEL,\r\n    instruction=&amp;quot;Tell the customer their refund was declined and exactly why, based &amp;quot;\r\n                &amp;quot;on the case JSON you receive. If it carries a reviewer_note, that is &amp;quot;\r\n                &amp;quot;the reason. Short email, direct and kind. Do not invent policy.&amp;quot;,\r\n)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;lang-py&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e67177cd0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;p data-block-key="z0o0k"&gt;The refund workflow now has a path from the initial lookups to a decision and a reply, with a third branch for cases that need a person:&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;workflow = Workflow(\r\n    name=&amp;quot;refund_decision&amp;quot;,\r\n    edges=[\r\n        (START, (fetch_order, fetch_payment, fetch_history),\r\n         join_case, route_refund),\r\n        (route_refund, {&amp;quot;AUTO_APPROVE&amp;quot;:  approve_notice,\r\n                        &amp;quot;MANUAL_REVIEW&amp;quot;: escalate_to_human,\r\n                        &amp;quot;DENY&amp;quot;:          denial_notice}),\r\n    ],\r\n)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;lang-py&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e671763d0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The first chain fetches the records, joins them, and applies the policy. The second maps the router's decision to a destination. Automatic decisions go straight to a notice agent; &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;MANUAL_REVIEW&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; goes to the human-review node we'll define next.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/graph-workflows-in-adk-everything-you-need.max-1000x1000_9lNpFLs.png"
        
          alt="graph-workflows-in-adk-everything-you-need-to-know-03-refund-graph"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;That third branch needs more than another function call. A reviewer may take minutes or days to answer, so the workflow must pause with the case pending and continue once the person decides.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The review node yields RequestInput, which records the pending request and pauses the run. On resume, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;rerun_on_resume=True&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; runs the node again, with the answer available in &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;ctx.resume_inputs&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;REVIEW = &amp;quot;refund:review&amp;quot;\r\n\r\nclass ReviewDecision(BaseModel):\r\n    approve: bool = Field(description=&amp;quot;True to refund, False to decline.&amp;quot;)\r\n    note: str = Field(&amp;quot;&amp;quot;, description=&amp;quot;Why, in the reviewer\&amp;#x27;s words.&amp;quot;)\r\n\r\n@node(rerun_on_resume=True)\r\nasync def escalate_to_human(ctx: Context, node_input: dict):\r\n    answer = ctx.resume_inputs.get(REVIEW)\r\n    if answer is None:                       # first pass: ask, then stop\r\n        yield RequestInput(\r\n            interrupt_id=REVIEW,\r\n            message=f&amp;quot;Refund ${node_input[\&amp;#x27;amount_usd\&amp;#x27;]} on order &amp;quot;\r\n                    f&amp;quot;{node_input[\&amp;#x27;order_id\&amp;#x27;]}?&amp;quot;,\r\n            payload=node_input,              # what the reviewer is shown\r\n            response_schema=ReviewDecision,\r\n        )\r\n        return\r\n    # second pass: the answer is here, so route on it\r\n    yield Event(output={**node_input, &amp;quot;reviewer_note&amp;quot;: answer.get(&amp;quot;note&amp;quot;, &amp;quot;&amp;quot;)},\r\n                route=&amp;quot;AUTO_APPROVE&amp;quot; if answer[&amp;quot;approve&amp;quot;] else &amp;quot;DENY&amp;quot;)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;lang-py&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e67177090&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The response schema gives the node an &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;approve&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; value to route on and a note to carry forward. If the reviewer declines, the notice agent receives their reason with the case. One more edge connects the review decision to the reply:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;edges=[\r\n    ...,\r\n    (escalate_to_human, {&amp;quot;AUTO_APPROVE&amp;quot;: approve_notice,\r\n                         &amp;quot;DENY&amp;quot;:         denial_notice}),\r\n]&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;lang-py&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e67175b50&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;ADK ships two variants of this. The one above is the single-node pattern: the node reruns and reads &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;ctx.resume_inputs&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, as in the &lt;/span&gt;&lt;a href="https://github.com/google/adk-python/blob/main/contributing/samples/workflows/request_input_rerun/agent.py" rel="noopener" target="_blank"&gt;&lt;code style="text-decoration: underline; vertical-align: baseline;"&gt;request_input_rerun&lt;/code&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt; sample&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. The &lt;/span&gt;&lt;a href="https://github.com/google/adk-python/blob/main/contributing/samples/workflows/request_input/agent.py" rel="noopener" target="_blank"&gt;&lt;code style="text-decoration: underline; vertical-align: baseline;"&gt;request_input&lt;/code&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt; sample&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; shows the two-node variant, where one node yields &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;RequestInput&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; and the reviewer's answer arrives as the next node's &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;node_input&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; — so there is no &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;ctx.resume_inputs&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; to find in that file. The companion scripts linked below include the code that sends the reviewer's answer back to the run.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;A &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;Workflow&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; is also a node. This whole refund process can become one step in a larger customer-service workflow.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Apply the same step to a batch of cases&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We now have a process for one refund. Suppose a batch of cases arrives with the records already collected. Each needs the same policy check, and the number of cases changes from batch to batch. We can apply one node to every item using a &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;parallel worker&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In ADK, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;parallel_worker=True&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; runs a node once per item in an input list and collects the results in the original order. We can reuse our policy function:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;@node(parallel_worker=True)\r\ndef review_case(node_input):\r\n    # Each worker receives one case from the input list.\r\n    case = node_input\r\n    decision = refund_policy(\r\n        case[&amp;quot;amount_usd&amp;quot;], case[&amp;quot;placed_days_ago&amp;quot;],\r\n        case[&amp;quot;chargeback_open&amp;quot;], case[&amp;quot;prior_refunds_12mo&amp;quot;],\r\n    )\r\n    return {&amp;quot;order_id&amp;quot;: case[&amp;quot;order_id&amp;quot;], &amp;quot;decision&amp;quot;: decision}\r\n\r\ndef collect_decisions(node_input):\r\n    # This node receives the list of worker results.\r\n    return {&amp;quot;decisions&amp;quot;: node_input}\r\n\r\nbatch_review = Workflow(\r\n    name=&amp;quot;batch_review&amp;quot;,\r\n    edges=[(START, review_case, collect_decisions)],\r\n)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;lang-py&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e67109410&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Each worker receives one case, and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;collect_decisions&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; receives the results as a list. No separate &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;JoinNode&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; is needed. The flag also works on agents—for example, to write an explanation for each case. The &lt;/span&gt;&lt;a href="https://github.com/google/adk-python/blob/main/contributing/samples/workflows/parallel_worker/agent.py" rel="noopener" target="_blank"&gt;&lt;code style="text-decoration: underline; vertical-align: baseline;"&gt;parallel_worker&lt;/code&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt; sample&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; shows both forms.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The policy calculation here is small. Concurrency is more useful when each item waits on an API or model call, but the way inputs and results move stays the same.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt; &lt;/p&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;&lt;table border="1" style="border-collapse: collapse; width: 100%;"&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;td style="width: 31.4907%;"&gt;&lt;strong&gt;Pattern&lt;/strong&gt;&lt;/td&gt;
&lt;td style="width: 31.4907%;"&gt;&lt;strong&gt;Work distributed&lt;/strong&gt;&lt;/td&gt;
&lt;td style="width: 31.4907%;"&gt;&lt;strong&gt;Collected output&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="width: 31.4907%;"&gt;&lt;strong&gt;Fan-out with &lt;code&gt;JoinNode&lt;/code&gt;&lt;/strong&gt;&lt;/td&gt;
&lt;td style="width: 31.4907%;"&gt;Different notes doing independent jobs&lt;/td&gt;
&lt;td style="width: 31.4907%;"&gt;Dictionary keyed by node name&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="width: 31.4907%;"&gt;&lt;strong&gt;Parallel worker&lt;/strong&gt;&lt;/td&gt;
&lt;td style="width: 31.4907%;"&gt;The same node for every item&lt;/td&gt;
&lt;td style="width: 31.4907%;"&gt;List in input order&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;div align="left"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt; &lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The batch size can change without changing the graph. A variable amount of work still fits inside a fixed process. See the &lt;a href="https://github.com/google/adk-python/blob/main/docs/guides/workflow/parallel_worker/index.md" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;parallel worker guide&lt;/span&gt;&lt;/a&gt;&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; for execution details.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Let results shape the next step&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Our refund process has known paths, even when a batch contains more cases or a reviewer takes longer to answer. But some work only becomes clear as we investigate.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Suppose a disputed refund reveals a second transaction. Checking it raises a delivery question that needs further investigation. Now each result can create follow-up work. A &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;dynamic node&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; can examine those results and schedule the next checks in Python.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;ADK provides &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;ctx.run_node&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; to run another node and await its result. To see how that changes orchestration, let's first express our existing refund flow this way:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;HANDLERS = {\r\n    &amp;quot;AUTO_APPROVE&amp;quot;: approve_notice,\r\n    &amp;quot;MANUAL_REVIEW&amp;quot;: escalate_to_human,\r\n    &amp;quot;DENY&amp;quot;: denial_notice,\r\n}\r\n\r\n@node(rerun_on_resume=True)\r\nasync def refund_flow(ctx, node_input):\r\n    # Step 1: fetch all three records at once.\r\n    order, payment, history = await asyncio.gather(\r\n        ctx.run_node(fetch_order,   node_input, use_sub_branch=True),\r\n        ctx.run_node(fetch_payment, node_input, use_sub_branch=True),\r\n        ctx.run_node(fetch_history, node_input, use_sub_branch=True),\r\n    )\r\n    case = order | payment | history\r\n\r\n    # Step 2: apply the refund policy — the same pure function, 0 LLM calls.\r\n    decision = refund_policy(\r\n        amount=case[&amp;quot;amount_usd&amp;quot;],\r\n        days_ago=case[&amp;quot;placed_days_ago&amp;quot;],\r\n        chargeback_open=case[&amp;quot;chargeback_open&amp;quot;],\r\n        priors=case[&amp;quot;prior_refunds_12mo&amp;quot;],\r\n    )\r\n\r\n    # Step 3: run the chosen handler, and use its output as this node\&amp;#x27;s output.\r\n    await ctx.run_node(HANDLERS[decision], case, use_as_output=True)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;lang-py&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e67176d90&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;code style="vertical-align: baseline;"&gt;asyncio.gather&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; runs the lookups together. Python combines their results, applies the policy, and runs the selected &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;handler. use_as_output=True&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; makes the handler's result the parent node's output without emitting it twice.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The dynamic version also adjusts the human-review node: after the answer arrives, it calls the chosen notice agent through &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;ctx.run_node&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;. The complete dynamic script (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;refund_dynamic.py&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;) includes that variation.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/graph-workflows-in-adk-everything-you-need.max-1000x1000_XydYu0o.png"
        
          alt="graph-workflows-in-adk-everything-you-need-to-know-04-dynamic"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;p data-block-key="z0o0k"&gt;The outer graph only needs an entry point:&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;workflow = Workflow(\r\n    name=&amp;quot;refund_dynamic&amp;quot;,\r\n    edges=[(START, refund_flow)],\r\n)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;lang-py&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e67176dd0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In the static version, the edge list shows the branches and destinations. In this version, we read &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;refund_flow&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; to see them. The &lt;/span&gt;&lt;a href="https://github.com/google/adk-python/blob/main/docs/guides/workflow/dynamic_nodes/index.md" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;dynamic node guide&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; covers this approach.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The human-review pause still works here. When the answer arrives, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;refund_flow&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; runs again from the top, but completed &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;ctx.run_node&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; calls return their recorded outputs from session history. The lookups do not repeat, as direct function calls would. Keeping side effects inside child nodes lets completed calls replay their results when the parent resumes. Each child also gets its own trace span, and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;use_sub_branch=True&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; keeps concurrent children's events on separate branches.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For this fixed refund process, the edge list remains easy to inspect. The Python version gives us a place to add the investigation logic described above: inspect a result, choose a follow-up node, and run independent checks together. A model might suggest what to investigate, while code limits the work—for example, three follow-ups per finding and two levels of investigation before human review.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Loops can still fit in a static graph. A fixed draft → check → revise process can use a conditional back-edge and a router that limits revisions. “Static” describes the possible connections; the actual path and iteration count can vary. The &lt;/span&gt;&lt;a href="https://github.com/google/adk-python/blob/main/docs/guides/workflow/graph/index.md" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;graph guide&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; covers conditional cycles.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;You can also place a dynamic node inside a static workflow, using Python for a stage that needs it while keeping the surrounding process visible.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Choose who decides what runs next&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Start with a question: &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;can you draw the possible workflow before the input arrives?&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Include branches, loops, and repeated stages. You do not need to predict the path each request will take.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt; &lt;/p&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;&lt;table border="1" style="border-collapse: collapse; width: 100%;"&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="width: 48.1356%;"&gt;
&lt;p&gt;&lt;strong&gt;What you need&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="width: 48.1356%;"&gt;&lt;strong&gt;Pattern&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="width: 48.1356%;"&gt;Independent jobs, then all their results&lt;/td&gt;
&lt;td style="width: 48.1356%;"&gt;Fan-out and fan-in&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="width: 48.1356%;"&gt;A branch chosen by fixed rules&lt;/td&gt;
&lt;td style="width: 48.1356%;"&gt;Deterministic router&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="width: 48.1356%;"&gt;A branch chosen by interpreting meaning&lt;/td&gt;
&lt;td style="width: 48.1356%;"&gt;Agent router&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="width: 48.1356%;"&gt;A person's decision before continuing&lt;/td&gt;
&lt;td style="width: 48.1356%;"&gt;&lt;code&gt;RequestInput&lt;/code&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="width: 48.1356%;"&gt;The same step across a list&lt;/td&gt;
&lt;td style="width: 48.1356%;"&gt;Parallel worker&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="width: 48.1356%;"&gt;Code that schedules work as results arrive&lt;/td&gt;
&lt;td style="width: 48.1356%;"&gt;Dynamic node&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;div align="left"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt; &lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Use an edge list when it makes those connections clear. Use dynamic orchestration when results create further work or Python expresses the control more naturally. A small, open-ended task may need only one agent and its tools.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In our refund workflow, the graph coordinates the lookups, code applies the policy, a person handles exceptions, and a model writes the reply. Graph engineering gives each a clear responsibility—and makes it easier to see how the process works.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;strong style="vertical-align: baseline;"&gt;Get started&lt;/strong&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Read:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; the &lt;/span&gt;&lt;a href="https://github.com/google/adk-python/blob/main/docs/guides/workflow/workflow/index.md" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;workflow guide&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Explore:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; the &lt;/span&gt;&lt;a href="https://github.com/google/adk-python/tree/main/contributing/samples/workflows" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;workflow samples&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Build:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; the &lt;/span&gt;&lt;a href="https://codelabs.developers.google.com/adk2/instructions#0" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;hands-on codelab&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, which applies these patterns to a marathon race-day coach.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Tue, 29 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/developers-practitioners/graph-workflows-in-adk-everything-you-need-to-know/</guid><category>Developers &amp; Practitioners</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/graph-workflows-in-adk-everything-you-need-t.max-600x600_G3dwz11.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Graph Workflows in ADK: Everything You Need to Know</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/graph-workflows-in-adk-everything-you-need-t.max-600x600_G3dwz11.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/developers-practitioners/graph-workflows-in-adk-everything-you-need-to-know/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Annie Wang</name><title>Google AI Cloud Developer Advocate</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Shangjie Chen</name><title>Software Engineer, Google Cloud AI</title><department></department><company></company></author></item><item><title>Best practices guide for customizing Gemini models via Reinforcement Learning (RL)</title><link>https://cloud.google.com/blog/topics/developers-practitioners/best-practices-guide-for-customizing-gemini-models/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Reinforcement learning (RL) has been a keystone of modern LLM post-training, but it demands large training clusters and access to model internals that external customers can't have with proprietary models like Gemini. So here at Google Cloud, we packaged it into &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;a managed RL fine-tuning service (RLFT service) &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;— you bring prompts and a reward function; we handle the infrastructure and the proprietary model internals. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Now, you can adapt Gemini with the service — teaching the model from a reward signal you define, rather than from a fixed set of labeled answers. This unlocks a class of problems that supervised fine-tuning (SFT) struggles with: tasks that are hard to demonstrate but easy to score.  &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In this guide, we will walk through practical best practices for using RL fine-tuning service. We'll start with a short tour of the RL training loop, how to decide if and when to use RL, and introduce how to get the most value from this approach.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;What is RLFT? &lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;RLFT adapts Gemini from a &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;reward signal you define&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; rather than labeled answers. Instead of authoring a large set of gold examples, you write one program that &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;scores&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; a response and the service improves the model against it — unlocking tasks that are hard to &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;demonstrate&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; but easy to &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;verify&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;: you can't hand-write the ideal SQL for every schema, but you can run the query and check the result.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_-_Single-Step_RL_Training_Loop.max-1000x1000.png"
        
          alt="1 - Single-Step RL Training Loop"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;At each training step the service generates multiple candidate responses to your prompts, scores them with your reward, and improves the model so that higher-scoring responses become more likely while it stays close to the original Gemini. The reinforcement learning that makes this work is fully managed — you never configure it. The one thing you own, and the thing that most determines your results, is the reward.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Three properties define what RLFT can and can't do:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;It &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;learns from the model's own outputs:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; It refines what the model already produces rather than copying an external target, so it tends to disturb unrelated capabilities less than SFT. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;It &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;rewards outcomes, not paths:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;   Any response that reaches a good result earns reward, which fits open-ended tasks with many valid solutions. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;It &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;amplifies existing competence: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;  It makes &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;occasional&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; success &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;reliable&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;, but it can't teach a skill the model never demonstrates.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;When to use RLFT&lt;/span&gt;&lt;/h3&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/2_-_RLFT_Approaches_-_Direct_RL_or_SFT_War.max-1000x1000.png"
        
          alt="2 - RLFT Approaches - Direct RL or SFT Warmup to Continuous RLFT"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Prompting and SFT handle most adaptation; exhaust them first. RLFT earns its keep when you can &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;grade a response but can't cheaply author it&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, when &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;SFT has plateaued&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; on the metric that matters (faithfulness, schema validity, tone), or when the task has &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;many equally valid answers&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; a single reference target would wrongly penalize. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;SFT and RLFT are complementary, not competing:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Direct RLFT&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; when the base model already succeeds part of the time — enough for the reward to tell better answers from worse ones.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Two-stage SFT → RLFT&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; when you have SFT data or the base success rate is too low for RL to gain traction. Use SFT as a short, cheap warm start — kept light, since over-fitting the demonstrations leaves less room for RL to improve — then continue into RL via &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Continuous Tuning&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, which initializes RL from the SFT checkpoint.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Across early adopters, these patterns show where RLFT delivers the most value — each scoring an outcome the business cares about but could never cheaply demonstrate.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Use cases for RLFT&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;AI-powered NPCs in games&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;What:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; In-character, on-brand dialogue held across long, multilingual, multi-turn conversations.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Problem:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Off-the-shelf models break immersion — wrong language, hallucinated items, ignored players, repetitive loops.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Objective and reward:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; A &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Gemini autorater (LLM-as-a-judge)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; scores each turn on persona, flow, and game-state syntax, penalizing format and language errors.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Results:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Loops and language drift disappeared and state syntax held, making shippable in-game characters viable at scale.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Structured entity extraction&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;What:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Pulling a set of items from unstructured documents, such as supplier invoices and shipping manifests, into structured records automatically&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Problem:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The long tail where SFT plateaus — missing required fields (recall) or inventing ones that aren't there (precision).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Objective and reward:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; A &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;rule-based precision/recall reward&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; forces every field to be grounded in the source text, not imitated from one gold answer.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Results:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Field-level accuracy rose on noisy real-world documents where tuning had stalled, turning a manual review step into an automated one.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Content moderation&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;What:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Applying intricate policies and decision trees at scale.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Problem:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Models hallucinate false positives or reward-hack with invalid formats to dodge evaluation.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Objective and reward:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; A &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Cloud Run reward&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; pairs format validation with a deterministic grader to enforce multi-step policy adherence.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Results:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The model handled complex exemption carve-outs, sharply cut false positives, and stopped reward hacking — reducing the human-escalation volume that makes moderation expensive.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Code measured by execution&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;What:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; SQL or API calls graded on whether they actually run against customer data.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Problem:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; SFT mimics one reference query and breaks on unseen proprietary schemas.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Objective &amp;amp; Reward:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; A &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;code-execution reward&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; runs the code in a secure sandbox and pays out only if it compiles, executes, and returns the correct result.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Results:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The model produced first-attempt executable queries at closed-frontier quality and lower inference cost, letting non-technical users query proprietary data in natural language.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Presentation slide generation via HTML&lt;/strong&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;What:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Multi-slide decks authored as HTML/CSS.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Problem:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Training on text alone is blind to visual quality — overflows, clipped elements, and inconsistent styling slip through unnoticed.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Objective &amp;amp; Reward:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; A &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;code-execution reward&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; renders the slides and scores visual design, layout integrity, structural completeness, and rubric adherence.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Results:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The model emitted modular, well-styled decks with cohesive themes and no layout overflow.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Where to start?&lt;/span&gt;&lt;/h3&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;A dataset.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; A diverse set of prompts with a held-out validation split is enough for a first run — confirm the loop converges and reward moves the right way, then scale. Keep train and eval strictly separated; a contaminated eval hides overfitting.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;A reward function.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Your task specification as code or configs, and the dominant driver of quality. A good reward correlates with human preference, is robust to malformed output (catch the failed parse and return a clearly negative score rather than crashing), and resists reward hacking — ensemble judges, penalize length, floor degenerate outputs, and prefer a verifiable check over a model's opinion. Validate it offline before launch.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The service handles the rest; start from the defaults, watch reward and eval curves in the console, and take the checkpoint where validation reward saturates rather than the last step.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/3_-_rlft_tutorial.gif"
        
          alt="3 - rlft_tutorial"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Get started today&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;What will you build? The tools are ready and waiting. &lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/tuning/reinforcement-tuning"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Documentation: Reinforcement Learning Fine-Tuning&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Fri, 25 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/developers-practitioners/best-practices-guide-for-customizing-gemini-models/</guid><category>AI &amp; Machine Learning</category><category>Developers &amp; Practitioners</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Best practices guide for customizing Gemini models via Reinforcement Learning (RL)</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/developers-practitioners/best-practices-guide-for-customizing-gemini-models/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Jiaqi Pan</name><title>Senior Software Engineer, Google</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Kunal Jha</name><title>Senior Product Manager, Google</title><department></department><company></company></author></item><item><title>Agent Factory recap: Agent harnesses, shifting left, and autonomous coding</title><link>https://cloud.google.com/blog/topics/developers-practitioners/agent-factory-recap-agent-harnesses-shifting-left-and-autonomous-coding/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In this episode of &lt;/span&gt;&lt;a href="https://www.youtube.com/playlist?list=PLIivdWyY5sqLXR1eSkiM5bE6pFlXC-OSs" rel="noopener" target="_blank"&gt;&lt;span style="font-style: italic; text-decoration: underline; vertical-align: baseline;"&gt;The Agent Factory&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, we explore the reality of building with autonomous agents alongside Ryan Lopopolo, a software engineer at Google Cloud and the person who coined the term &lt;/span&gt;&lt;a href="https://cloud.google.com/discover/agent-harness?e=48754805"&gt;&lt;strong style="font-style: italic; text-decoration: underline; vertical-align: baseline;"&gt;agent harness&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. From throwing out manual code editors to treating team collaboration like leveling up RPG stats, Ryan breaks down how grounding models in rich context and shifting interventions left unlocks high levels of agent autonomy.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-video"&gt;



&lt;div class="article-module article-video "&gt;
  &lt;figure&gt;
    &lt;a class="h-c-video h-c-video--marquee"
      href="https://youtube.com/watch?v=F8EZJAm9iO8"
      data-glue-modal-trigger="uni-modal-F8EZJAm9iO8-"
      data-glue-modal-disabled-on-mobile="true"&gt;

      
        &lt;img src="//img.youtube.com/vi/F8EZJAm9iO8/maxresdefault.jpg"
             alt="Harness Engineering Explained: Inside the Stack Behind Antigravity, Claude Code &amp;amp; Cursor"/&gt;
      
      &lt;svg role="img" class="h-c-video__play h-c-icon h-c-icon--color-white"&gt;
        &lt;use xlink:href="#mi-youtube-icon"&gt;&lt;/use&gt;
      &lt;/svg&gt;
    &lt;/a&gt;

    
  &lt;/figure&gt;
&lt;/div&gt;

&lt;div class="h-c-modal--video"
     data-glue-modal="uni-modal-F8EZJAm9iO8-"
     data-glue-modal-close-label="Close Dialog"&gt;
   &lt;a class="glue-yt-video"
      data-glue-yt-video-autoplay="true"
      data-glue-yt-video-height="99%"
      data-glue-yt-video-vid="F8EZJAm9iO8"
      data-glue-yt-video-width="100%"
      href="https://youtube.com/watch?v=F8EZJAm9iO8"
      ng-cloak&gt;
   &lt;/a&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This post guides you through the key ideas from our conversation. Use it to quickly recap topics or dive deeper into specific segments with links and timestamps.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;The Agent Harness - What is it?&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;Timestamp: [&lt;/span&gt;&lt;a href="https://www.youtube.com/watch?v=F8EZJAm9iO8&amp;amp;t=30s" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;00:30&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;]&lt;/span&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;An &lt;em&gt;AI agent&lt;/em&gt; as we're defining it here is a large language model (LLM) plus an &lt;/span&gt;&lt;a href="https://cloud.google.com/discover/agent-harness?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;agent harness&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Think of the harness as everything wrapped around the LLM that isn't the model itself. For example, if you're working in &lt;/span&gt;&lt;a href="https://antigravity.google/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Antigravity&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; using &lt;/span&gt;&lt;a href="https://antigravity.google/blog/gemini-3-8-flash-in-google-antigravity" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini 3.8 Flash&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, Gemini Flash is the LLM and Google Antigravity is the harness.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;While an unassisted model can answer simple questions out of the box, it can't check live conditions or interact with your workspace on its own. When a user asks a question like &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;"Why is the sky blue?"&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;, an unassisted LLM can respond without issue. However, when asked a question like &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;"Should I wear a raincoat today?"&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;, the model can't answer on its own because it lacks the necessary data. The harness catches the intent, queries live weather tools, bundles that context back into the prompt, and hands it to the model to produce an informed answer.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Ryan Lopopolo on agent harnesses and autonomous coding&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Tilde Thurium sat down with Ryan Lopopolo to discuss what it takes to run fully autonomous coding workflows in production. See the summary below!&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Coining the harness and writing zero production code&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;Timestamp: [&lt;/span&gt;&lt;a href="https://www.youtube.com/watch?v=F8EZJAm9iO8&amp;amp;t=142s" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;02:22&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;]&lt;/span&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The term &lt;em&gt;agent harness&lt;/em&gt; grew out of Ryan's extensive work on autonomous coding agents, culminating in a &lt;/span&gt;&lt;a href="https://openai.com/index/harness-engineering/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;February 2026 essay&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; on leveraging coding models in an agent-first world. Ryan shared that he hasn't opened a traditional code editor since May of last year, maintaining that streak through his transition into &lt;/span&gt;&lt;a href="https://cloud.google.com/"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. In this paradigm, engineers no longer author or review individual lines of syntax; instead, they operate at the level of natural language specifications and inspect the final artifacts, such as pull requests, documents, and spreadsheets. Then they determine whether the end result meets organizational standards.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-pull_quote"&gt;&lt;div class="uni-pull-quote h-c-page"&gt;
  &lt;section class="h-c-grid"&gt;
    &lt;div class="uni-pull-quote__wrapper h-c-grid__col h-c-grid__col--8 h-c-grid__col-m--6 h-c-grid__col-l--6
      h-c-grid__col--offset-2 h-c-grid__col-m--offset-3 h-c-grid__col-l--offset-3"&gt;
      &lt;div class="uni-pull-quote__inner-wrapper h-c-copy h-c-copy"&gt;
        &lt;q class="uni-pull-quote__text"&gt;Harness engineering is the study and the practice of putting a model into an environment where it can succeed. If you don&amp;#x27;t do that work, you end up doing what I call &amp;#x27;prompt and pray&amp;#x27;.&lt;/q&gt;

        
          &lt;cite class="uni-pull-quote__author"&gt;
            
            
              &lt;span class="uni-pull-quote__author-meta"&gt;
                
                  &lt;strong class="h-u-font-weight-medium"&gt;Ryan Lopopolo&lt;/strong&gt;&lt;br /&gt;
                
                
              &lt;/span&gt;
            
          &lt;/cite&gt;
        
      &lt;/div&gt;
    &lt;/div&gt;
  &lt;/section&gt;
&lt;/div&gt;

&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Context curation and lazy prompting&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;Timestamp: [&lt;/span&gt;&lt;a href="https://youtu.be/F8EZJAm9iO8?si=dZGjT_rG5fiLeVfm&amp;amp;t=215" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;03:35&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;]&lt;/span&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Upfront harness investment pays off by allowing engineers to become lazy prompters. When the repository contains structured documentation, clear interfaces, and discoverable tools, you do not need to paste walls of text into a prompt box every morning &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;"I aspire to be an incredibly lazy prompter. If I have done the job to give the model the tools and context it needs to ground itself, I don't need to write a long prompt. It figures it out."&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The model uses its harness to pull relevant context, allowing it to navigate large codebases and execute complex tasks without oversight.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Shifting left: engineering best practices as autonomous guardrails&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;Timestamp: [&lt;/span&gt;&lt;a href="https://www.youtube.com/watch?v=F8EZJAm9iO8&amp;amp;t=300s" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;05:00&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;]&lt;/span&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When an agent fails, developers face a whole spectrum of interventions. The most common reflex is to fiddle with the prompt or retry, but that never scales across a team.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;"The simplest, smooth-brain, stupidest intervention I can think of is literally just: try my prompt again without changing anything else. But shifting left means moving interventions earlier into the development lifecycle where they are cheapest and automated: from prompts, to repo docs, to linters, to tests, and all the way to upstream evals."&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Instead of hoping the model guesses right on the next turn, shifting left embeds standards directly into the environment. Linters, tests, and &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;AGENTS.md&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; files act as durable memory and enforcement of what you think good looks like, making sure the agent stays on the rails without needing constant hand-holding.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Leveraging established tools and determinism&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;Timestamp: [&lt;/span&gt;&lt;a href="https://youtu.be/F8EZJAm9iO8?si=0K18aWfjWSkEv0Ws&amp;amp;t=410" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;06:50&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;]&lt;/span&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Agents shine when they're handed tools that already mirror patterns heavily represented in pre-training data. Pairing models with standard command-line interfaces moves reasoning into determinism, shifting the burden of context aggregation away from the model and onto reliable tools.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Ryan also shared an environmental design trick for context efficiency: structuring markdown files so link anchors sit directly beneath their corresponding prose blocks rather than inline. This prevents context clutter and mitigates "lost in the middle" retrieval issues. Because Ryan operates exclusively by reviewing end-state artifacts, keeping documentation readable allows him to easily inspect execution runs:&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;"I want to be able to look at the pull request and review it. If it made a bad decision, I need to know where it went off the rails so I can whack the agent on the head and make sure it does not make that same mistake again."&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Long horizons and expanding the agentic loop&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;Timestamp: [&lt;/span&gt;&lt;a href="https://youtu.be/F8EZJAm9iO8?si=o3r4mie7FkoYLTTS&amp;amp;t=586" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;09:45&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;]&lt;/span&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The central challenge of harness engineering is ensuring that agents cohere over long time horizons. Because human organizations produce software through iterative refinement rather than single-shot prompts, agent workflows must mirror that cadence. Harness engineering uses tightly scoped, reviewable pull requests to narrow the agent's state space. Stacking these high-confidence changes end-to-end allows supervisors to gradually expand the loop size, building trust until agents can autonomously execute large-scale initiatives, including entire language migrations &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Curating agent teams like RPG stats&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;Timestamp: [&lt;/span&gt;&lt;a href="https://youtu.be/F8EZJAm9iO8?si=1l25hbXKZjjKoLsh&amp;amp;t=718" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;11:58&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;]&lt;/span&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Rather than divvying up sprint tasks based on individual specialties, having a diverse team contribute to an agent turns it into a central producer of work that carries everyone's strengths. Ryan compared leveling up an agent's capabilities to building out a character sheet:&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;"[It's like] building out the stats of your RPG character. I get a new person on the team who is a React architect and boom! The attention that they pay is able to bump out the stats in front-end architecture and performance."&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With that collective expertise baked into the environment, the agent can autonomously classify incoming work and activate the exact skills it needs on demand, operating as both a backend architect and a front-end specialist.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Accruing leverage in tools, not custom harnesses&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;Timestamp: [&lt;/span&gt;&lt;a href="https://youtu.be/F8EZJAm9iO8?si=giil4EL8zs3RDif-&amp;amp;t=878" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;14:38&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;]&lt;/span&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For developers wondering whether to build their own custom agent harness, Ryan offered clear advice: don't build one from scratch. Standard harnesses already provide the foundational primitives: file reading, grep search, and command execution. Over-scaffolding an agent with rigid, bespoke frameworks creates technical debt and leads to sunk-cost traps when frontier models advance.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;"If you focus all of your efforts on improving quality on tools and context, you can freely adopt the newest models as they come out and you'll be constantly accruing leverage into a bit of the system that will never become obsolete."&lt;/em&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Operating Google Cloud and eliminating capability overhang&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;Timestamp: [&lt;/span&gt;&lt;a href="https://youtu.be/F8EZJAm9iO8?si=MlXgcYa5QcQqOogM&amp;amp;t=1007" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;16:47&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;]&lt;/span&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Discussing his work at &lt;/span&gt;&lt;a href="https://cloud.google.com/"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, Ryan outlined his motivation to eliminate &lt;em&gt;capability overhang&lt;/em&gt;: the delta between what frontier AI models are theoretically capable of and how much useful work is currently extracted in production. Because the cloud functions as a massive, programmable surface, equipping agents with direct interfaces to Google Cloud lets them manage and deploy infrastructure effectively, turning raw model capability into tangible enterprise utility.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Continually updating your priors on AI&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;Timestamp: [&lt;/span&gt;&lt;a href="https://youtu.be/F8EZJAm9iO8?si=NSFTXZjs6QvgCFiM&amp;amp;t=1093" rel="noopener" target="_blank"&gt;&lt;span style="font-style: italic; text-decoration: underline; vertical-align: baseline;"&gt;18:13&lt;/span&gt;&lt;/a&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;]&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The speed of AI development requires engineers and teams to actively unlearn old limitations and constantly reassess what these models can achieve. "&lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;It's very important to continually be updating what you think is possible with these lovely tools that we have&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;," Ryan urged. What broke six months ago often runs effortlessly on today's frontier models. Rather than getting locked into rigid workflows, developers should build around the two highly extensible interfaces that will remain relevant across every model upgrade: tools and context.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;"Agents will always need context in order to do that last mile adaptation into what you think good is. And as you can continue to... shift it to the left, in terms of increasingly capable tools which act as a form of memory and enforcement of what you think good looks like, you'll continually be amazed as the models are able to do more and more interesting things for you over time."&lt;/em&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;How To Build A Custom Harness&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Next, &lt;/span&gt;&lt;a href="https://www.linkedin.com/in/billyjacobson/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Billy Jacobson&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; started us off by showing how developers can customize their own agent harnesses for specific tasks.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Under the hood: Why build a custom harness?&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;Timestamp: [&lt;/span&gt;&lt;a href="https://www.youtube.com/watch?v=F8EZJAm9iO8&amp;amp;t=1184s&amp;amp;pp=0gcJCWMAwfN6Pr3D" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;19:44&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;]&lt;/span&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Before jumping into code, Billy unpacked why developers should understand the mechanics of a harness rather than treating it like a black box. Recalling advice from an engineering mentor that &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;"You can just use the framework, but a great engineer will really understand the framework"&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;, Billy explained that building a &lt;/span&gt;&lt;a href="https://cloud.google.com/discover/agent-harness?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;harness&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; yourself is the best way to debug what happens when an agent breaks. You can evaluate three core design decisions for every workflow:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Looping&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: How many iterations should the agent run, and what conditions trigger an exit state?&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Tools&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: What specific tools should the agent access, and when and how should it invoke them?&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Memory&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: How important is conversational and operational memory, and when should it be retrieved or compacted?&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Linear Agent Harness: Deterministic Single-Pass Execution&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;Timestamp: [&lt;/span&gt;&lt;a href="https://youtu.be/F8EZJAm9iO8?si=teRfnEVBrXQa0DBl&amp;amp;t=1234" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;21:34&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;]&lt;/span&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Billy demonstrated a minimalist &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/architecture/choose-design-pattern-agentic-ai-system#sequential-pattern"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;linear harness&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; designed for deterministic workflows where looping is unnecessary. This pattern is ideal for targeted inspections, file transformations, or single-turn data analyses where you want a high level of determinism and need the agent to perform the exact same execution flow every single time.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Closed-Loop Agent Harness: Iterative Test-Driven Repair&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;Timestamp: [&lt;/span&gt;&lt;a href="https://youtu.be/F8EZJAm9iO8?si=SZK9VsaI6x4FpKXT&amp;amp;t=1365" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;22:45&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;]&lt;/span&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When tasks demand active bug fixing and refactoring, a &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/architecture/choose-design-pattern-agentic-ai-system#loop-pattern"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;closed-loop harness&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; provides the iterative reasoning required to reach a verified resolution. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In this demo, Billy showcased an agent that applies an automated code edit to address a failing requirement, and the harness executes the unit test suite against the updated codebase. If the tests fail, the runtime captures standard failure logs and detailed stack traces, feeding those error diagnostics directly back into the agent's working memory. The process repeats continuously until all unit tests pass, backed by a five-iteration ceiling to prevent infinite loops and runaway execution costs.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Guardrail Harness with Google's Agent Development Kit (ADK)&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;Timestamp: [&lt;/span&gt;&lt;a href="https://youtu.be/F8EZJAm9iO8?si=iyoPxuMlAkYkrrls&amp;amp;t=1405" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;23:25&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;]&lt;/span&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For developers who require custom behavior without rewriting core orchestration plumbing from scratch, Google's &lt;/span&gt;&lt;a href="https://adk.dev/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Agent Development Kit &lt;span style="vertical-align: baseline;"&gt;(ADK)&lt;/span&gt;&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; provides scaffolding with automated memory management and execution safeguards. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Billy walked through an example that leverages ADK's native context compaction to summarize older conversational turns, preventing context window bloat during extended debugging runs. Custom interception hooks inspect and filter shell actions before execution, automatically stopping high-risk operations such as recursive file deletions, database drops, or unauthorized remote git pushes. This architecture gives teams fine-grained control over tool execution boundaries while avoiding the maintenance burden of bespoke harness frameworks.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;The 3-Layer Agent Dev Stack: Gemini 3.8 Flash, Google Antigravity, and Google Skills&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;em&gt;&lt;span style="vertical-align: baseline;"&gt;Timestamp: [&lt;/span&gt;&lt;a href="https://youtu.be/F8EZJAm9iO8?si=L0T5reEiQBXxZ86H&amp;amp;t=1527" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;25:27&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;]&lt;/span&gt;&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Next up, &lt;/span&gt;&lt;a href="https://www.linkedin.com/in/smithakolan/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Smitha Kolan&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; broke down why coding agents do not always require heavier reasoning models, emphasizing that high performance stems from balancing the three layers of the agent stack: &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Model&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Harness&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, and &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Knowledge&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;"Your coding agent doesn't need a smarter model. It needs a better stack: model, harness, and knowledge. When all three click into place, everything changes."&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;She then walked through the three tools she's been loving recently, one for each layer of the stack.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Layer 1 | Model | &lt;/strong&gt;&lt;a href="https://antigravity.google/blog/gemini-3-8-flash-in-google-antigravity" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini 3.8 Flash&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; High-frequency agentic loops run between 20 and 60 sequential hops per task (inspecting files, updating functions, and executing unit tests). Because latency and API costs compound across iterations, a lightweight, responsive model like Gemini 3.8 Flash makes real-time agent loops practical without running up a massive bill.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Layer 2 | Harness | &lt;/strong&gt;&lt;a href="https://antigravity.google/" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Google Antigravity&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt; with &lt;/strong&gt;&lt;a href="https://antigravity.google/docs/boost/" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;/boost&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; Default Antigravity handles standard navigation and component creation. On top of that, the &lt;/span&gt;&lt;code&gt;&lt;span style="vertical-align: baseline;"&gt;/boost&lt;/span&gt;&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; command spins up an orchestrator that coordinates specialized sub-agents in parallel and concludes with an independent audit pass before modifying files.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Layer 3 | Knowledge | &lt;/strong&gt;&lt;a href="https://github.com/google/skills/tree/main" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Google Skills Repository&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; With over 19,000 GitHub stars and 100+ curated domain packages across Google Cloud, Firebase, Flutter, and Maps, this harness-agnostic repository injects precise domain context on demand, preventing agents from guessing cloud configurations &lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Your turn to build&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Building effective coding agents requires moving past the reflex of simply swapping in larger models. As Ryan Lopopolo's philosophy of harness engineering illustrates, true developer leverage is achieved by shifting best practices to the left and investing in rich tools, deterministic verifiers, and well-curated context that survive model upgrades. When combined with fast inference models, structured orchestration harnesses, and modular domain knowledge, agents evolve from conversational novelties into dependable, autonomous engineering partners.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Ready to put it into practice? Explore the tools and resources covered in this episode:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://antigravity.google/blog/gemini-3-8-flash-in-google-antigravity" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini 3.8 Flash&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;: Fast, low-cost model for multi-hop agent loops.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://antigravity.google/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Antigravity&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (&lt;/span&gt;&lt;a href="https://antigravity.google/docs/boost/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;/boost&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;): Orchestrator harness for parallel coding agents and verification.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://github.com/google/skills/tree/main" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Skills&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;: Modular domain knowledge for Google Cloud, Firebase, and Flutter.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://adk.dev/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Agent Development Kit (ADK)&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;: Custom harness middleware for safety guardrails and memory compaction.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://youtu.be/F8EZJAm9iO8?si=Uptrs898iW1XIRC6" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Full episode video&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;: The full interview and live Factory Floor code demos.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Connect with us&lt;/span&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Smitha Kolan&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; →&lt;/span&gt; &lt;a href="https://www.linkedin.com/in/smithakolan/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;LinkedIn&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; |&lt;/span&gt; &lt;a href="https://www.youtube.com/@smithakolan" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;YouTube&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; |&lt;/span&gt; &lt;a href="https://x.com/smithakolan" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;X&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Luke Schlangen&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; → &lt;/span&gt;&lt;a href="https://www.linkedin.com/in/lukeschlangen/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;LinkedIn&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; |&lt;/span&gt; &lt;a href="https://www.luke.mn/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;website&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Ryan Lopopolo&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; → &lt;/span&gt;&lt;a href="https://www.linkedin.com/in/ryanlopopolo/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;LinkedIn&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; | &lt;/span&gt;&lt;a href="https://hyperbo.la/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Hyperbola&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; | &lt;/span&gt;&lt;a href="https://x.com/_lopopolo" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;X&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Tilde&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Thurium &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;→&lt;/span&gt; &lt;a href="https://www.linkedin.com/in/annthurium/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;LinkedIn&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Billy Jacobson&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; →&lt;/span&gt; &lt;a href="https://www.linkedin.com/in/billyjacobson/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;LinkedIn&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; | &lt;/span&gt;&lt;a href="https://x.com/billyjacobson" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;X&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Mollie Pettit&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; → &lt;/span&gt;&lt;a href="https://www.linkedin.com/in/molliepettit/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;LinkedIn&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; | &lt;/span&gt;&lt;a href="https://dev.to/molliepettit" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;dev.to&lt;/span&gt;&lt;/a&gt; | &lt;a href="https://x.com/MollzMP" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;X&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; |&lt;/span&gt; &lt;a href="https://bsky.app/profile/mollzmp.bsky.social" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Bluesky&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Thu, 24 Sep 2026 23:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/developers-practitioners/agent-factory-recap-agent-harnesses-shifting-left-and-autonomous-coding/</guid><category>Developers &amp; Practitioners</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/agent-factory-recap-agent-harness.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Agent Factory recap: Agent harnesses, shifting left, and autonomous coding</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/agent-factory-recap-agent-harness.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/developers-practitioners/agent-factory-recap-agent-harnesses-shifting-left-and-autonomous-coding/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Mollie Pettit</name><title>Developer Relations Engineer</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Smitha Kolan</name><title>Senior Developer Relations</title><department></department><company></company></author></item><item><title>How Google Cloud Networking Supports Your Fluid Compute Choices for AI Workloads</title><link>https://cloud.google.com/blog/topics/developers-practitioners/how-google-cloud-networking-supports-your-fluid-compute-choices-for-ai-workloads/</link><description>&lt;div class="block-paragraph"&gt;&lt;p data-block-key="1cnbu"&gt;The availability of resources for AI workloads can be challenging across the industry, especially accelerators. This can slow your AI workload deployment if it’s built around a specific type of accelerator. The concept of &lt;b&gt;fluid compute&lt;/b&gt; allows you to design your AI deployment with several options based on available resources that can fit your use case.&lt;/p&gt;&lt;p data-block-key="e2jpo"&gt;In this blog, we will explore how Google Cloud networking supports your AI workloads and considerations that are relevant to your choice of accelerator (GPU or TPU), as the backend networking component configuration is not exactly the same.&lt;/p&gt;&lt;h2 data-block-key="4araf"&gt;The resource options&lt;/h2&gt;&lt;p data-block-key="7qk9r"&gt;After deciding the type of work you want to achieve with your AI deployment, another important component is the actual hardware to get this done. In this case, we want to run inference for a private LLM, and the target is the NVIDIA B200 GPU family which is available in the &lt;a href="https://cloud.google.com/compute/docs/gpus"&gt;A4 VMs&lt;/a&gt; (a4-highgpu-8g).&lt;/p&gt;&lt;p data-block-key="6uc8e"&gt;Now we have identified what we want to get done and a possible compute option, but the challenge is: is this available?&lt;/p&gt;&lt;p data-block-key="57nfi"&gt;To get access to resources, there are several options which include:&lt;/p&gt;&lt;ul&gt;&lt;li data-block-key="ch7k0"&gt;&lt;a href="https://docs.cloud.google.com/compute/docs/instances/about-flex-start-vms"&gt;Dynamic Workload Scheduler (Flex-start VM)&lt;/a&gt;: Queues workloads until all required accelerator nodes are available at the same time, provisioning them together and running non-preemptibly for up to seven days.&lt;/li&gt;&lt;li data-block-key="qniq"&gt;&lt;a href="https://docs.cloud.google.com/compute/docs/instances/future-reservations-calendar-mode-overview"&gt;Dynamic Workload Scheduler (calendar mode)&lt;/a&gt;: Enables reserving accelerator capacity 1 to 90 days in advance with guaranteed start and end times, ideal for scheduled pre-training runs and benchmarking.&lt;/li&gt;&lt;li data-block-key="datn1"&gt;&lt;a href="https://docs.cloud.google.com/compute/docs/instances/future-reservations-overview"&gt;Future reservations&lt;/a&gt;: Guarantees access to committed hardware in a specified zone beginning at a specific future date.&lt;/li&gt;&lt;li data-block-key="dn02u"&gt;&lt;a href="https://cloud.google.com/compute/docs/instances/reservations-overview"&gt;Flex reservations&lt;/a&gt;: Offers short-term commitment windows to secure scarce accelerator nodes without multi-year lock-in.&lt;/li&gt;&lt;li data-block-key="60m4b"&gt;&lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/how-to/node-auto-provisioning#enable-workload-level"&gt;Dynamic node auto-provisioning and ComputeClasses&lt;/a&gt;: In Google Kubernetes Engine (GKE), defining multi-family fallback lists within ComputeClasses allows the cluster to automatically attempt provisioning alternative accelerator types if primary pools face regional constraints.&lt;/li&gt;&lt;li data-block-key="3p8rj"&gt;&lt;a href="https://cloud.google.com/compute/docs/instances/spot"&gt;Spot VMs&lt;/a&gt;: Delivers surplus compute at substantial discounts for fault-tolerant, checkpointed batch jobs.&lt;/li&gt;&lt;/ul&gt;&lt;p data-block-key="6fdjm"&gt;Read more on this in the blog &lt;a href="https://medium.com/google-cloud/never-run-out-of-compute-a-practical-guide-to-gke-resource-obtainability-eab4dea059ca" target="_blank"&gt;&lt;i&gt;Never Run Out of Compute: A Practical Guide to GKE Resource Obtainability&lt;/i&gt;&lt;/a&gt;.&lt;/p&gt;&lt;h2 data-block-key="4nu88"&gt;Networking your choices&lt;/h2&gt;&lt;p data-block-key="9vs9m"&gt;The networking component of the accelerator varies based on your choice, so let's explore four configurations: standard networking, accelerated GPU networking (&lt;a href="https://docs.cloud.google.com/compute/docs/gpus/gpudirect"&gt;TCPX/TCPXO&lt;/a&gt; and &lt;a href="https://cloud.google.com/blog/products/networking/rdma-rocev2-for-ai-workloads-on-google-cloud?e=48754805"&gt;RoCEv2&lt;/a&gt;), TPU networking, and Cloud Run.&lt;/p&gt;&lt;h3 data-block-key="9rjjn"&gt;&lt;b&gt;Standard networking&lt;/b&gt;&lt;/h3&gt;&lt;ul&gt;&lt;li data-block-key="fgtoe"&gt;&lt;b&gt;Supported accelerators:&lt;/b&gt; NVIDIA T4 (&lt;a href="https://docs.cloud.google.com/compute/docs/gpus#n1-gpus"&gt;N1 series&lt;/a&gt;), NVIDIA L4 (&lt;a href="https://docs.cloud.google.com/compute/docs/gpus#l4-gpus"&gt;G2 series&lt;/a&gt;), NVIDIA A100 (&lt;a href="https://docs.cloud.google.com/compute/docs/accelerator-optimized-machines#a2-vms"&gt;A2 machine series&lt;/a&gt; single-node and multi-node), Cloud TPU v3, and &lt;a href="https://cloud.google.com/tpu/docs/v5e"&gt;Cloud TPU v5e&lt;/a&gt; (single-host/standalone slices).&lt;/li&gt;&lt;li data-block-key="7ub8d"&gt;&lt;b&gt;Architecture:&lt;/b&gt; Nodes communicate over the primary &lt;a href="https://cloud.google.com/vpc/docs/vpc"&gt;Virtual Private Cloud (VPC)&lt;/a&gt; network using the &lt;a href="https://cloud.google.com/compute/docs/networking/using-gvnic"&gt;Google Virtual NIC (gVNIC)&lt;/a&gt; over standard TCP/IP.&lt;/li&gt;&lt;li data-block-key="qrc9"&gt;&lt;b&gt;Workload fit:&lt;/b&gt; Provides straightforward portability across Google Cloud compute environments, supporting distributed data preprocessing, decoupled pipeline stages, independent inference replicas, and computer vision workloads using standard VPC routing and network policies.&lt;/li&gt;&lt;/ul&gt;&lt;h3 data-block-key="1h2el"&gt;&lt;b&gt;Accelerated GPU Networking (TCPX/TCPXO and RoCEv2)&lt;/b&gt;&lt;/h3&gt;&lt;p data-block-key="cg8on"&gt;Distributed training and multi-node inference require specialized multi-rail network fabrics to handle massive parameter exchanges and collective communications.&lt;/p&gt;&lt;p data-block-key="62qt2"&gt;&lt;b&gt;GPUDirect-TCPX and TCPXO Fabrics&lt;/b&gt;&lt;/p&gt;&lt;ul&gt;&lt;li data-block-key="8c575"&gt;&lt;b&gt;Supported accelerators&lt;/b&gt;: NVIDIA H100 (&lt;a href="https://docs.cloud.google.com/compute/docs/accelerator-optimized-machines#a3-high-vms"&gt;A3 High VMs&lt;/a&gt; with 4 rails) and NVIDIA H100 Mega (&lt;a href="https://docs.cloud.google.com/compute/docs/accelerator-optimized-machines#a3-mega-vms"&gt;A3 Mega VMs&lt;/a&gt; with 8 rails).&lt;/li&gt;&lt;li data-block-key="1ipe4"&gt;&lt;b&gt;Architecture:&lt;/b&gt; Uses custom &lt;a href="https://docs.cloud.google.com/compute/docs/gpus/gpudirect#a3-high-and-a3-edge"&gt;GPUDirect-TCPX&lt;/a&gt; (4 dedicated VPCs) and &lt;a href="https://docs.cloud.google.com/compute/docs/gpus/gpudirect#a3-mega"&gt;GPUDirect-TCPXO&lt;/a&gt; (8 dedicated VPCs) offload engines to achieve high-throughput multi-rail GPU communication over standard Ethernet infrastructure without requiring native RDMA hardware.&lt;/li&gt;&lt;li data-block-key="22e9k"&gt;&lt;b&gt;Deployment blueprints:&lt;/b&gt; These multi-VPC topologies can be deployed in many ways including using pre-built blueprints from the &lt;a href="https://docs.cloud.google.com/cluster-toolkit/docs/setup/cluster-blueprint"&gt;Cluster Toolkit&lt;/a&gt;.&lt;/li&gt;&lt;/ul&gt;&lt;p data-block-key="3fe9a"&gt;&lt;b&gt;RoCEv2 Fabrics (VM and Bare Metal)&lt;/b&gt;&lt;/p&gt;&lt;ul&gt;&lt;li data-block-key="57rtk"&gt;&lt;b&gt;Supported accelerators&lt;/b&gt;: NVIDIA H200 (&lt;a href="https://docs.cloud.google.com/compute/docs/gpus#h200-gpus"&gt;A3 Ultra VMs&lt;/a&gt;), NVIDIA B200 (&lt;a href="https://docs.cloud.google.com/compute/docs/gpus#b200-gpus"&gt;A4 VMs&lt;/a&gt;), NVIDIA GB200 NVL72 (&lt;a href="https://docs.cloud.google.com/compute/docs/gpus#gb200-gpus"&gt;A4X VMs&lt;/a&gt;), and NVIDIA GB300 (&lt;a href="https://docs.cloud.google.com/compute/docs/gpus#gb300-gpus"&gt;A4X Max Bare Metal&lt;/a&gt;).&lt;/li&gt;&lt;li data-block-key="3bcf8"&gt;&lt;b&gt;Zonal network profiles:&lt;/b&gt; RoCEv2 operates over a dedicated RDMA VPC attached to a specialized zonal network profile: VM instances (A3 Ultra, A4, A4X) use the &lt;a href="https://docs.cloud.google.com/vpc/docs/rdma-network-profiles#roce-supported-features"&gt;ZONE-vpc-roce&lt;/a&gt; profile, while Bare Metal instances (such as A4X Max) utilize the dedicated &lt;a href="https://docs.cloud.google.com/vpc/docs/rdma-network-profiles#roce-metal-supported-features"&gt;ZONE-vpc-roce-metal&lt;/a&gt; bare-metal profile.&lt;/li&gt;&lt;li data-block-key="5sks7"&gt;&lt;b&gt;Rail-aligned fabrics:&lt;/b&gt; This dedicated VPC is isolated strictly for GPU communication and contains subnets mapped directly to the accelerator NICs. The backend is rail-aligned, with support for Jumbo Frames (&lt;a href="https://cloud.google.com/vpc/docs/mtu"&gt;MTU 8896&lt;/a&gt;), delivering non-blocking multi-terabit bandwidth with minimal cross-rail interference.&lt;/li&gt;&lt;li data-block-key="7ifut"&gt;&lt;b&gt;Automated plumbing with GKE&lt;/b&gt; &lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/how-to/allocate-network-resources-dra"&gt;&lt;b&gt;Dynamic Resource Allocation Network (DRANET)&lt;/b&gt;&lt;/a&gt;: When deploying these GPUs on GKE, the GKE managed DRANET can be used to automatically provision additional networks and assign drivers that map the RDMA network interfaces to the GPU. These can then be assigned and consumed directly in your workload pods using standard Kubernetes resource claims.&lt;/li&gt;&lt;li data-block-key="2eg66"&gt;&lt;b&gt;Turnkey deployment&lt;/b&gt;: You can deploy this entire end-to-end stack—including RDMA VPCs, MTU tuning, and DRA drivers—using automated blueprints from the &lt;a href="https://cloud.google.com/cluster-toolkit/docs/overview"&gt;Cluster Toolkit&lt;/a&gt;.&lt;/li&gt;&lt;/ul&gt;&lt;h3 data-block-key="4h9mf"&gt;&lt;b&gt;TPU Networking&lt;/b&gt;&lt;/h3&gt;&lt;ul&gt;&lt;li data-block-key="2ep8f"&gt;&lt;b&gt;Supported accelerators&lt;/b&gt;: &lt;a href="https://cloud.google.com/tpu/docs/v4"&gt;Cloud TPU v4&lt;/a&gt;, &lt;a href="https://cloud.google.com/tpu/docs/v5p"&gt;Cloud TPU v5p&lt;/a&gt;, &lt;a href="https://cloud.google.com/tpu/docs/v5e"&gt;Cloud TPU v5e&lt;/a&gt; (multi-host Pod slices), &lt;a href="https://cloud.google.com/tpu/docs/v6e"&gt;Cloud TPU v6e&lt;/a&gt; (Trillium), and &lt;a href="https://docs.cloud.google.com/tpu/docs/tpu7x"&gt;TPU7x&lt;/a&gt; (Ironwood).&lt;/li&gt;&lt;li data-block-key="6hk6v"&gt;&lt;b&gt;Inter-chip interconnect (ICI)&lt;/b&gt;: Inside a TPU Pod or slice, chips communicate directly over dedicated, ultra-low-latency optical links organized in 2D or 3D torus meshes, bypassing traditional network stacks entirely.&lt;/li&gt;&lt;li data-block-key="dabji"&gt;&lt;b&gt;Optical circuit switches&lt;/b&gt; (OCS): In TPU v4 and TPU v5p SuperPods, software-reconfigurable OCS units dynamically change physical network topologies, route around faulty trays, and provision custom-sized accelerator slices without manual recabling.&lt;/li&gt;&lt;li data-block-key="2cpgc"&gt;&lt;b&gt;Multi-NIC architecture&lt;/b&gt; (TPU v6e and Higher): While earlier TPU generations relied on ICI within a slice and single-NIC for host traffic, Cloud TPU v6e (Trillium) and TPU7x introduce a native multi-NIC architecture where worker nodes isolate standard Kubernetes management traffic onto a primary VPC while using secondary dedicated VPCs configured for high-throughput TPU data and cross-slice communication.&lt;/li&gt;&lt;li data-block-key="93dce"&gt;&lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/how-to/allocate-network-resources-dra#use-non-rdma-interfaces-tpu"&gt;&lt;b&gt;DRANET for TPU&lt;/b&gt;&lt;/a&gt;&lt;b&gt; deployments&lt;/b&gt;: When deploying these TPUs on GKE, the GKE managed DRANET can be used to automatically provision additional networks and assign drivers for TPU communication. These can then be assigned and consumed directly in your workload pods using standard Kubernetes resource claims.&lt;/li&gt;&lt;li data-block-key="6v7ae"&gt;&lt;b&gt;Data-center network (DCN) Multislice&lt;/b&gt;: For models scaling beyond an individual TPU slice, &lt;a href="https://docs.cloud.google.com/tpu/docs/multislice-introduction"&gt;Cloud TPU Multislice&lt;/a&gt; connects multiple independent ICI meshes over Google's high-speed Jupiter Data Center Network utilizing these dedicated multi-NIC paths.&lt;/li&gt;&lt;/ul&gt;&lt;h3 data-block-key="bsskn"&gt;&lt;b&gt;Cloud Run&lt;/b&gt;&lt;/h3&gt;&lt;ul&gt;&lt;li data-block-key="7kmlq"&gt;&lt;b&gt;Supported accelerators:&lt;/b&gt; NVIDIA L4 (G2 series) and NVIDIA RTX PRO 6000 (Blackwell) on &lt;a href="https://cloud.google.com/run/docs/configuring/services/gpu"&gt;Cloud Run GPU services&lt;/a&gt;.&lt;/li&gt;&lt;li data-block-key="2npad"&gt;&lt;b&gt;Direct VPC egress&lt;/b&gt;: Binds serverless containers directly to your private VPC network using sub-minute IP allocation via &lt;a href="https://cloud.google.com/run/docs/configuring/vpc-direct-vpc"&gt;Direct VPC Egress&lt;/a&gt;, enabling secure, low-latency access to internal data lakes, databases, and private APIs without traversing the public internet or requiring legacy connector VMs.&lt;br/&gt;&lt;/li&gt;&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/how-google-cloud-networking-supports-your-.max-1000x1000.jpg"
        
          alt="how-google-cloud-networking-supports-your-fluid-compute-choices-networks"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;h2 data-block-key="1cnbu"&gt;&lt;b&gt;Summary&lt;/b&gt;&lt;/h2&gt;&lt;p data-block-key="3oqnr"&gt;Google Cloud networking options support various accelerator types. When using fluid compute you can adjust your network setup to support the best design to optimise your workloads performance.&lt;/p&gt;&lt;h2 data-block-key="8sopm"&gt;&lt;b&gt;Next Steps&lt;/b&gt;&lt;/h2&gt;&lt;p data-block-key="f0grd"&gt;Take a deeper dive into Google Cloud AI infrastructure and networking architectures with these resources:&lt;/p&gt;&lt;ul&gt;&lt;li data-block-key="d4450"&gt;Blog: &lt;a href="https://cloud.google.com/blog/topics/ai-infrastructure/best-practices-for-dynamic-capacity-management?e=48754805"&gt;Dynamic capacity management for AI infrastructure&lt;/a&gt;&lt;/li&gt;&lt;li data-block-key="a7eab"&gt;Tutorial: &lt;a href="https://discuss.google.dev/t/how-to-build-an-elastic-scalable-llm-inference-platform-on-gke-using-fluid-compute/388108" target="_blank"&gt;How to build an elastic, scalable LLM Inference Platform on GKE using Fluid Compute&lt;/a&gt;&lt;/li&gt;&lt;li data-block-key="9rt2j"&gt;Blog: &lt;a href="https://cloud.google.com/blog/products/networking/how-google-cloud-networking-supports-your-ai-workloads"&gt;How Google Cloud Networking Supports Your AI Workloads&lt;/a&gt;&lt;/li&gt;&lt;/ul&gt;&lt;p data-block-key="8ir6m"&gt;Want to ask a question, find out more, or share a thought? Please connect with me on &lt;a href="https://www.linkedin.com/in/ammett/" target="_blank"&gt;LinkedIn&lt;/a&gt;.&lt;/p&gt;&lt;/div&gt;</description><pubDate>Thu, 24 Sep 2026 13:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/developers-practitioners/how-google-cloud-networking-supports-your-fluid-compute-choices-for-ai-workloads/</guid><category>Networking</category><category>Developers &amp; Practitioners</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/how-google-cloud-networking-supports-your-fl.max-600x600_VSTWN6E.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>How Google Cloud Networking Supports Your Fluid Compute Choices for AI Workloads</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/how-google-cloud-networking-supports-your-fl.max-600x600_VSTWN6E.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/developers-practitioners/how-google-cloud-networking-supports-your-fluid-compute-choices-for-ai-workloads/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Ammett Williams</name><title>Developer Relations Engineer</title><department></department><company></company></author></item><item><title>A guide to speeding up your video processing with AlphaEvolve</title><link>https://cloud.google.com/blog/topics/developers-practitioners/how-to-speed-up-your-video-processing-with-alphaevolve/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In real-time streaming, every millisecond counts. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For example, at 30 frames per second (fps), developers have a strict frame budget of just 33.3 ms (and only 16.6 ms at 60 fps) to ingest camera frames, run neural segmentation, apply shaders, and composite output. Exceeding that budget by even a fraction of a millisecond leads to dropped frames and stuttering. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Manual optimization is notoriously tedious — requiring weeks of analyzing flame graphs and hand-tuning low-level code in Swift, C++, or Metal. While standard AI coding assistants can generate boilerplate, they can’t optimize  against target hardware, benchmark real-world latency, or ensure optimizations preserve visual fidelity.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Autonomous, closed-loop evolutionary optimization changes this paradigm. Tools like&lt;/span&gt; &lt;a href="https://cloud.google.com/blog/products/ai-machine-learning/alphaevolve-is-available-for-everyone?e=0&amp;amp;utm_source=gemini"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;AlphaEvolve&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; pair cloud-scale model reasoning with local hardware execution, and we’re already seeing real-world impact. In partnership with Google,&lt;/span&gt; &lt;a href="https://www.doit.com/about?utm_source=gemini" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;DoIt&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; used AlphaEvolve to autonomously optimize production Swift code in a live macOS streaming app, uncovering performance headroom that manual profiling missed (read the full&lt;/span&gt;&lt;a href="https://medium.com/google-cloud/running-alphaevolve-on-your-own-code-f8aeebceb4d0?utm_source=gemini" rel="noopener" target="_blank"&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;technical writeup&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;).&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;While this post focuses on video pipelines, the split-loop pattern applies anywhere performance matters — from microservice throughput and database queries to ML tensor pipelines and embedded systems. In every case, the formula is the same: pair Gemini code generation in the cloud with your domain-specific benchmark harness and automated quality gates.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Today, we’ll show you how to use AlphaEvolve to speed up video processing—and apply these principles to your own performance bottlenecks:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Understanding the split-loop architecture: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;How AlphaEvolve decouples managed cloud generation (Gemini model ensemble on Google Cloud) from local evaluation (e.g. compiling and timing native Swift/Metal code).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Evaluator craft and quality gates:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; How to construct scoring functions using metrics like Structural Similarity Index (SSIM) to prevent evolutionary loops from gaming the benchmark (e.g., skipping rendering entirely to go fast).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Autonomous algorithmic discovery:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; How Gemini-driven evolutionary search can autonomously discover unprompted framework APIs and make intelligent engineering trade-offs (e.g., frame-caching limits).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Setting realistic performance boundaries: &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;How to measure code optimization against physical hardware floors.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;1. Understanding AlphaEvolve’s split-loop architecture &lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;AlphaEvolve runs a closed-loop evolutionary process: given a seed program and a custom scoring function, a mixture of Gemini models proposes code variations, executes the scoring function against each candidate, keeps the highest-performing code, and iteratively climbs toward an optimal solution over multiple generations.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_OsAwmXL.max-1000x1000.jpg"
        
          alt="1"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;A core architectural advantage of AlphaEvolve is its clean separation into two halves:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;The generation half (Google Cloud managed service):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Contains the prompt sampler, Gemini model ensemble, and program database. Google Cloud handles the scale, prompt orchestration, and generation mechanics.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;The evaluation half (customer managed compute):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Scoring code quality is strictly domain-specific. You own the evaluator module entirely, running it on your own hardware or target architecture (in this case, macOS running native Swift code).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;While AlphaEvolve is Python-first on the cloud generation side, evaluation can be written in any language. The custom evaluator compiles each Swift candidate using swift and executes it against a standard reference webcam clip.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;2. Evaluator craft and quality gates&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;An automated optimization loop like AlphaEvolve never actually "sees" your video stream. It only sees the numeric fitness score your evaluator returns. If your evaluation metric has a blind spot, evolutionary code generation will aggressively exploit it.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In our early runs, a naive fitness score weighted toward raw latency produced an astonishing speedup: the model simply bypassed blur rendering entirely and returned unmodified frames in 0 ms.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Structural Similarity Index Measure (SSIM)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To prevent the model from gaming your benchmark, try building a two-tiered scoring function that pairs throughput with structural fidelity metrics like &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Structural Similarity Index (SSIM)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;speedup = baseline_ms_per_frame / candidate_ms_per_frame\r\nssim    = mean_ssim_vs_golden\r\n\r\n#Disqualify any candidate falling below visual threshold\r\n\r\n\r\nif ssim &amp;lt; 0.98 or worst_frame_ssim &amp;lt; 0.95:\r\n    return {&amp;quot;speedup&amp;quot;: -1e12}   # Disqualified\r\n\r\nreturn {&amp;quot;speedup&amp;quot;: speedup, &amp;quot;ssim&amp;quot;: ssim}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e74107410&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;What does this give you?&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;The ability to test against worst-case clips:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Never benchmark on static frames or blank cameras. Candidate code can easily pass an average SSIM gate on static backgrounds while failing completely during quick head turns.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;You can track the minimum, not just the mean:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Enforce both an average threshold and a per-frame floor to catch dropped frames or delayed mask updates.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Autonomous algorithmic discovery:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Most developers use generative AI for local micro-optimizations (e.g., inlining helper functions, unrolling loops, or tweaking memory pools). But when given architectural room, the evolutionary loop can discover systemic optimizations on its own.&lt;/span&gt;&lt;/p&gt;
&lt;h4&gt;&lt;strong style="vertical-align: baseline;"&gt;Engineering lessons:&lt;/strong&gt;&lt;/h4&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Provide framework context, not isolated loops:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Include public SDK headers, interface definitions, or API reference symbols in the prompt or retrieval harness. An LLM cannot adopt a sequence-aware subsystem if its context window only contains an isolated frame-processing callback.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Expose multi-frame lifecycle hooks:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Let your candidate code maintain a bounded state across executions (e.g., historical masks or cache timestamps) rather than enforcing pure, stateless functions.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Let quality gates police the trade-offs:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; When AlphaEvolve introduced temporal mask caching, it initially cached masks too aggressively, causing noticeable trailing artifacts. Because our SSIM gate penalized drift during motion, the search converged on a production-ready cache window without manual parameter tuning.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Setting realistic performance boundaries&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;A common pitfall in performance engineering is optimizing in the dark. If you achieve a 2x speedup, is that an incredible achievement, or did you leave another 3x on the table?&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In real-time media, total frame time splits into two distinct categories:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Mutable software overhead:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Memory allocations, buffer format conversions, thread context switches, and API dispatch friction.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Immutable hardware floors:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Raw Neural Engine inference latency, GPU shader compute time, and hardware display synchronization.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;To make the most of AlphaEvolve, developers should measure against theoretical maximum headroom&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Before running optimization loops, here’s a few principles to keep in mind: &lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Build a "no-op" pipeline:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Strip out Swift/C++ orchestration, data marshalling, and frame conversions. Dispatch only the pre-warmed ML model and bare GPU pass on a dummy buffer. The resulting time is your physical hardware lower bound.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Calculate your addressable ceiling:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Your total possible optimization potential is:&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/original_images/2_u73NadS.jpg"
        
          alt="2"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;3. &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Score against the hardware gap:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Instead of arbitrary speedup multiples, measure optimization efficiency:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/3_Y7cUASN.max-1000x1000.jpg"
        
          alt="3"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Get started &lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;All benchmark code, test clips, evaluation scripts, and raw candidate logs are open source:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;GitHub repository:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://github.com/SaschaHeyer/gen-ai-livestream/tree/main/alphaevolve/examples/camera-background-blur" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;AlphaEvolve Camera Background Blur Example&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Detailed technical write-up of our case study with DoIt:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://medium.com/google-cloud/running-alphaevolve-on-your-own-code-f8aeebceb4d0" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Running AlphaEvolve on Your Own Code&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Wed, 23 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/developers-practitioners/how-to-speed-up-your-video-processing-with-alphaevolve/</guid><category>Developers &amp; Practitioners</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>A guide to speeding up your video processing with AlphaEvolve</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/developers-practitioners/how-to-speed-up-your-video-processing-with-alphaevolve/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Anant Nawalgaria</name><title>Group AI Product Manager &amp; Engineer, Google</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Sascha Heyer</name><title>Principal AI Lead, DoIt</title><department></department><company></company></author></item><item><title>The DevFest Community Workshop Experience: Building Real Agents Together</title><link>https://cloud.google.com/blog/topics/developers-practitioners/the-devfest-community-workshop-experience-building-real-agents-together/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This week we kicked off the DevFest season in North America at Google Hudson Square in New York City with 80 engineers packed into the room. Typical technical workshops hand you a finished repo, tell you to blindly paste blocks of code into your terminal, and hope nothing crashes. You walk away with green checkmarks, but your brain stays on autopilot.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We've introduced a completely different experience called &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;Workbench&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Workbench focuses on understanding core ideas and architectural models rather than obsessing over syntax and code snippets. Instead of getting bogged down in boilerplate, engineers spent the day grappling with the actual mental models behind graph engineering, self-evolving architectures, and automated self-patching harnesses.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;A glimpse into the Workshop Experience&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;At the DevFest Community Workshop, we spent one intense day building long-running, self-evolving multi-agent systems powered by Google's agentic stack. Ricky Robinett, Senior Director of Developer Marketing, kicked off the day by diagnosing why so many engineering teams hit a wall with agents. Ricky broke down why prompt engineering fails as a safety mechanism: English is just a probabilistic suggestion, not an execution boundary. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Right after Ricky, Rachel Francois, Google Developer Groups (GDG) North America Program Lead, took the stage alongside GDG Brooklyn organizers to welcome the community and spotlight the power of local developer chapters. They set the tone for the entire day, reminding everyone that building durable software works best as a team sport where engineers share real-world patterns and build local networks that outlast any single framework.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Getting hands on with labs&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Annie Wang &amp;amp; Christina Lin, Americas DevRel Team members, led the morning lab that put those runtime ideas to work. Attendees explored Google's Agent Development Kit (ADK), Veo 3.1, Memory Bank on Gemini Enterprise Agent Platform, and RAG Engine on Gemini Enterprise Agent Platform. Through Workbench, developers grasped the principle of separating state from active compute for long running tasks. Workflows paused cleanly mid-execution, waited out asynchronous human approvals, and resumed without running up idle compute costs.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;After lunch, Logan Hennessy, Americas Developer Relations Engineer (DRE), and Kartik Derasari, Google Developer Expert (GDE), led a lab using auction history as insight for better bidding strategy. Attendees worked through the architecture by integrating BigQuery data into autonomous data engineering pipelines, reasoning about deterministic bidding logic and adding eval-gated, self-patching harnesses that catch spend anomalies and update runtime execution safely.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Between lab blocks, we ran fast-paced speed quizzes where developers raced to lock in their answers as quickly as possible. Screens flashed, fingers flew across keyboards, and seconds made the difference between topping the leaderboard or dropping five spots. Nothing beats watching a room full of serious engineers completely lose their cool over a live quiz leaderboard.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Join a DevFest Community Workshop this fall&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;New York was only round one. We are taking this exact experience on tour to five more cities this fall. Find your city and grab your seat before spots fill up:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://rsvp.withgoogle.com/events/devfest-extended-sunnyvale" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Sunnyvale on September 30&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://rsvp.withgoogle.com/events/devfest-extended-dc" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Washington DC on October 6&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://goo.gle/devfest-extended-atlanta" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Atlanta on October 30&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (as a part of DevFest Atlanta)&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://rsvp.withgoogle.com/events/devfest-extended-seattle" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Seattle on November 4&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://rsvp.withgoogle.com/events/devfest-extended-boston" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Boston on November 10&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Fri, 18 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/developers-practitioners/the-devfest-community-workshop-experience-building-real-agents-together/</guid><category>Developers &amp; Practitioners</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/devfest-community-workshop-experience-hero.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>The DevFest Community Workshop Experience: Building Real Agents Together</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/devfest-community-workshop-experience-hero.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/developers-practitioners/the-devfest-community-workshop-experience-building-real-agents-together/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Christina Lin</name><title>Developer Relations Engineering Manager</title><department></department><company></company></author></item><item><title>Best practices for handling cloud reliability incidents</title><link>https://cloud.google.com/blog/topics/developers-practitioners/cloud-reliability-incident-handling-best-practices/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Cloud outages can range from global service disruptions to issues isolated to a specific region, zone, or even just your project, workload or application. If you suspect a Google Cloud Platform outage is impacting your services, we recommend you follow a structured “Verify→ Investigate→Report→Resolve→Review" workflow to resolve it. And before that outage occurs, you should also have &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;prepared&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; your environment for an eventual disruption by designing for failure, and actively practicing the steps you need to take to restore service. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In this blog, we summarize the key reliability incident handling best practices to help you design and practice your reliability incident response capabilities and minimize impact. Rather than an exhaustive guide, this is meant as a primer on only the most important practices for advisory purposes. Please note that we do not cover additional practices specific to security incidents here. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Beyond the base steps covered here, you may want to also &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;explore how AI agents and tools are starting to transform incident handling. Check out &lt;/span&gt;&lt;a href="https://sre.google/prodcast/transcripts/sre-prodcast-04-09/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;this episode of the Prodcast&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, where Googlers explore the latest trends of &lt;/span&gt;&lt;a href="https://sre.google/prodcast/transcripts/sre-prodcast-04-09/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;leveraging agentic AI in Site Reliability Engineering&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; (SRE) to detect issues early and prevent disruptions. Try&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/cloud-assist/investigations"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Assist investigations&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, or explore &lt;/span&gt;&lt;a href="https://github.com/google/skills" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Agent Skills&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/mcp/supported-products"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;remote managed MCP servers&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to give you another set of tools for quickly pinpointing an issue. Before getting into these advanced techniques, we focus below on the foundational steps to good incident handling.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;1. Prepare&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Long before things start to go sideways, you should have spent significant time preparing for an outage along at least four dimensions: design, data, playbooks and training.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Design&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Think ahead and mitigate future incidents by designing automated response actions, like a load balancer shifting traffic away from slow or unresponsive instances, or by automating as much of your incident response playbook as possible. Review &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/architecture/framework"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;designs&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; of all critical applications to automate as many actions as possible to accelerate response and recovery.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Data&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: When a disruption occurs, having meaningful data at your fingertips vastly improves response capabilities. Use &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/logging/docs/overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Logging&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/trace/docs"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Trace&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/monitoring/docs/monitoring-overview"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Monitoring&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, or other third-party observability tools, and replicate that data to a redundant stack in a separate location from the systems being observed. Make sure, in advance of any incident, that time stamps are synced across your observability streams for easy correlation, or know how to do that on-demand during an outage, when time is of the essence.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Playbook&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: A well-thought-out playbook documenting your incident response processes, including crystal clear role and responsibility definitions for all personas, is paramount to efficient incident response. Who is responsible to do what? Who needs to be notified or mobilized for each type of disruption? How can they be reached? What tools and data are available? How are results communicated? How do teams hand over to the next shift during long running incidents? etc. Conduct a simulated incident response and critically review every step to find where your playbook needs clarification. Without clear responsibilities, mitigation inevitably takes longer.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Training&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Hopefully, service disruptions are rare events. To ensure your staff knows and remembers how to react, they need to retrain on the process several times per year&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; by running simulated cross-team incident response drills. A retrospective&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; on the simulated exercise will help identify warranted improvements.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;2. Verify&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Despite your best efforts, sooner or later, a service disruption will occur, which you can detect via any number of mechanisms:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Observability tools (&lt;/span&gt;&lt;a href="https://docs.cloud.google.com/docs/observability"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google tools&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; or third-party tools)&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://docs.cloud.google.com/unified-maintenance/docs/overview?_gl=1*1028q22*_ga*ODU2MjY4NzUyLjE3NzM0MTg2Mjk.*_ga_WH2QY8WWF5*czE3NzM2ODQ4MTckbzQkZzEkdDE3NzM2ODUwMjUkajEyJGwwJGgw"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Unified Maintenance Management&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; notifications for planned maintenance&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://www.google.com/search?q=https://console.cloud.google.com/service-health"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Personalized Service Health&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; notifications managed with alert policies&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Proactive customer monitoring by Google&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Now, you need to determine what broke and who should ultimately fix the problem:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Google, e.g., a bug, code roll-out, hardware failure, etc.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;You, e.g., a configuration change, elevated load, quota ceiling, etc.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Third party, e.g., a directory hosted by a different cloud provider&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If Google has declared an incident and started working to fix the problem, estimate whether you can possibly reestablish service sooner, for example by failing over to a secondary stack (see the ‘Typical Causes’ table below). You can determine whether Google has declared an incident and will provide a fix by consulting:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://cloud.google.com/service-health"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Personalized Service Health&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Check this first.&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Personalized Service Health shows incidents specifically relevant to your projects and regions, distinguishing between incident types:. &lt;/span&gt;&lt;/p&gt;
&lt;ul style="list-style-type: circle;"&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Emerging Incidents: Google has received an alert, on-callers are investigating, impact is yet unknown&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Confirmed Incidents: Google has investigated and found customers are impacted&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Located within the Google Cloud console, Personalized Service Health often displays limited-scope incidents that don't appear on the public dashboard. Personalized Service Health also offers a mobile client for Android and iOS smartphones, assuming you can use your work ID and credentials on the phone.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://g.co/kgs/j2BVWVE" rel="noopener" target="_blank"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Cloud Assist&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, which is &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/devops-sre/gemini-cloud-assist-integrated-with-personalized-service-health?e=4875480"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;integrated with Personalized Service Health&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, so you can use it to query that information in natural language.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://status.cloud.google.com/"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Cloud Service Health dashboard&lt;/strong&gt;&lt;/a&gt;&lt;strong style="vertical-align: baseline;"&gt;:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; This is the public-facing non-authenticated web page for broad, severe incidents affecting many customers. Limited blast radius disruptions are &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;not&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; externalized to the public. All its content is available in Personalized Service Health as well. If ever Personalized Service Health goes down, Cloud Service Health serves as an alternative channel built on a separate infrastructure.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Known Issues:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; In the console, navigate to &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Support &amp;gt; Cases&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, view a case, and u&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;se the resource selector on the console toolbar to find the specific cloud resource you’re interested in. Then click &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Known issues&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; If your issue matches one listed here, you can link a support case to it, so you will receive automatic updates in your case record. If you don’t find a match, open a new support case. Google will automatically match the case to a related incident, as soon as one is declared.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Google declared incidents are updated as new information becomes available, so check back regularly, or set up a Personalized Service Health alert policy to be notified each time new information becomes available.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If you host cloud resources in multiple clouds, a good practice is to check early on whether the problem occurs for multiple cloud providers. If so, the problem is likely external to the providers and caused either by you or by a third-party service that your application interacts with.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;3. Investigate&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To determine the blast radius within your cloud footprint of Google-declared reliability incidents, first check Personalized Service Health updates for a description of the technical problem. Knowing what to look for will allow you to map your blast radius and decide on suitable contingency actions quicker.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If Google hasn’t declared an incident, try to rule out configuration errors or issues within your environment by checking:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Cloud Monitoring:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Look for spikes in error rates (e.g. 5xx errors), increased latency, or drops in traffic in your dashboards.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Cloud Logs:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Use &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Log Explorer&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; to look for specific error messages like &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;DEADLINE_EXCEEDED&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;SERVICE_UNAVAILABLE&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;, or specific API errors.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Quotas:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Ensure you haven't hit a project quota (e.g., CPU, API rate limits), which can often mimic the behavior of an outage.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Change history:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Check your log of recently applied changes. Not all problems manifest immediately, but proximity on a timeline can be a powerful indicator of causality, even if it’s not proof. Also check whether Google rolled out any updates just before the symptoms started. See the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/unified-maintenance/docs/overview?_gl=1*1028q22*_ga*ODU2MjY4NzUyLjE3NzM0MTg2Mjk.*_ga_WH2QY8WWF5*czE3NzM2ODQ4MTckbzQkZzEkdDE3NzM2ODUwMjUkajEyJGwwJGgw"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Unified Maintenance Management&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; interface in Cloud Hub.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Absent a clear culprit, such as a traffic spike or a DDOS attack, and if symptoms manifested immediately after rolling out a change, a good strategy is to back out that change and attempt to return to a last known good configuration. &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;4. Report&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If the Cloud Service Health and Personalized Service Health dashboards are green but your metrics show a failure, you must report it to Google. &lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Determine priority:&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;ul&gt;
&lt;li aria-level="2" style="list-style-type: circle; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;P1 (Critical):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Your production service is unusable or severely impacted with no workaround.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="2" style="list-style-type: circle; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;P2 (High):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Significant impact or degradation, but a workaround may exist.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="2" style="list-style-type: circle; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;See &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/support/docs/best-practices#setting_the_priority_and_escalating"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;guidance on setting priority&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/support/docs/best-practices#describing_your_issue"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;guidance on describing your issue&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;File a case:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Go to &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Support &amp;gt; Cases &amp;gt; Create Case&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; in the console.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;ul&gt;
&lt;li aria-level="2" style="list-style-type: circle; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Explain&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; quantifiable business impact to rationalize the submitted priority and prevent it from being reset when Cloud Support prioritizes cases. A clear and accurate rationale helps!&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Essential information to include:&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;ul&gt;
&lt;li aria-level="2" style="list-style-type: circle; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Project ID&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; and affected &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;region/zone&lt;/strong&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="2" style="list-style-type: circle; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Timestamps&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; (when it started and if it's ongoing) with a clearly labeled timezone&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="2" style="list-style-type: circle; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Specific error messages&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; or log snippets&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="2" style="list-style-type: circle; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Scope:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Is it affecting all users/systems, or a specific subset/location?&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;Escalation for Premium/Enhanced support&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If you have a &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Premium&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; or &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Enhanced&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; support plan and a P1 case is not receiving the attention it requires, use the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/support/docs/best-practices#escalating"&gt;&lt;strong style="text-decoration: underline; vertical-align: baseline;"&gt;Escalate&lt;/strong&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; button within the support case in the console. This alerts a support manager to investigate and rectify the situation.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;5. Resolve&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By taking these steps, you are well on your way to resolving the outage. In the meantime, here are some ways to mitigate the impact of the outage and communicate with impacted stakeholders.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;While waiting for a resolution:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Communicate:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Notify your stakeholders and customers. Transparency helps manage expectations and reduces duplicate internal reports.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Fail over:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; If you have a multi-regional architecture, consider shifting traffic to a healthy region. As a best practice, first &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;ensure that the disruption is at the infrastructure level and not at your workload level.&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Check for workarounds:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; While working on a permanent fix, Google often posts temporary workarounds in the Service Health Dashboard updates, or in Personalized Service Health updates.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Consider your regulatory reporting requirements&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Know whether your organization is subject to regulatory reporting requirements, and what the required deadlines are for both initial and follow-up reporting. Google Cloud prepares Incident Reports for incidents that meet certain criteria — see details &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/service-health/docs/get-incident-reports"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;here&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for how to get those reports. Premium Support customers can also request an Incident Summary, which is an Incident Report customized to your account’s specific hosting location, time stamps, etc.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h4&gt;&lt;span style="vertical-align: baseline;"&gt;De-escalation and closure&lt;/span&gt;&lt;/h4&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Once systems are stable, Google downgrades the severity levels and deactivates the active on-call escalation chain. Google only closes an incident in Personalized Service Health when it has taken all the mitigation steps covering all impacted customers. Your specific services might be restored sooner than the incident closure time, if other customers are restored later than you. The incident is officially closed on the Google Cloud Status Dashboard when systems have run stably for a designated auto-close duration. Verify that your services are operating normally at this point. And if your incident responders aren’t compensated for extra time spent on the incident, find a way to thank them.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;6. Review&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;After the problem has been fixed and operations have returned to a normal, steady state, it’s time to conduct a &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/architecture/framework/reliability/conduct-postmortems"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;post-mortem analysis&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to identify how your team can respond better in future service disruptions. A “blameless” approach is essential to surfacing meaningful and impactful improvements that can be made to your incident response process. Ask questions like:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;What went well?&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;What could we have done better?&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Where did we get lucky?&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Where did we get unlucky?&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Then decide what changes can be made to improve your playbook, tools and training.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;At Google, we often publish a &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;post-mortem&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; or &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Incident Report&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; for major outages, available via Personalized Service Health. Review this to understand the root cause and adjust your own disaster recovery plans to prevent or reduce future impact. Customers with a Premium Support plan can request an &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Incident Summary&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; for a Google-caused incident they were impacted by and for which they opened a P1 case. An Incident Summary is an Incident Report customized for &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;your&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; environment (e.g., start and end times of impact).&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Typical causes, comms and prevention strategies&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To help you prepare and plan ahead, here’s an overview of some typical incidents based&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; on the symptoms reported in Cloud Service Health and Personalized Service Health along with guidance on what Google communications to expect, and some generic mitigation or prevention strategies you can build into your playbooks.&lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;/p&gt;
&lt;div align="left"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;&lt;table&gt;&lt;colgroup&gt;&lt;col/&gt;&lt;col/&gt;&lt;col/&gt;&lt;col/&gt;&lt;/colgroup&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Blast radius&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Typical cause&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Comms&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Strategy&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Single zone or region.&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Subset of products.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Typical of a software problem triggered by a rollout. Learning points:&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;- Understand the location scope (zones and regions) of your workload&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;- Products can depend on other products&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Major incidents are communicated via Cloud Service Health.&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Major and Minor (by number of customers, not severity) incidents are communicated via Personalized Service Health.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Highly localized incidents are not communicated via Cloud Service Health or Personalized Service Health.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Fail over, if so configured, but verify the health of the secondary stack first.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Single zone.&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Most or all products.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Typical of a power or cooling issue.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Check Cloud Service Health and Personalized Service Health.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Fail over to a different zone, if so configured.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Single region.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Most or all products.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Typical of a backbone networking infrastructure issue &lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Check Cloud Service Health and Personalized Service Health.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Fail over to a different region, if so configured.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Control plane issue for a product&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Typical of a late detected issue&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Communicated via Personalized Service Health if significant customer impact is verified.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Look for workarounds. Wait for Google to fix. Fail over, if so configured.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Multi-regional issue with a global product&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Rare but possible, typically detected quickly. Learnings: Mitigation options can be limited. Try regional variants, alternative products with similar functionality&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Check Cloud Service Health and Personalized Service Health.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Wait for Google to fix. In the meantime, verify via Google Comms and your own investigation that this is truly Google’s problem to fix.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Capacity / Stockout issue&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;System-level demand exceeding capacity in the product/location/model. (Cloud is designed to scale, but limits always exist, so proper planning is advised)&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Error message. No incident will be declared.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Place reservations for predicted capacity needs (if cost is acceptable). Flexibility in zone placement can also help.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Quota exhaustion&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Difficult / inaccurate prediction of traffic&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Error message. No incident will be declared.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: top; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Review consumption trends against ceiling regularly.&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Go deeper&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;This document offers only a condensed summary of key points. If you have an active Premium Support contract with Google Cloud, reach out to your account team for a deeper review of your response plans. For a comprehensive treatise on how to build reliable services and how to respond to incidents, we strongly recommend Google’s &lt;/span&gt;&lt;a href="https://sre.google/sre-book/table-of-contents/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;SRE Book&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, which is available as a free download. A new version of the SRE book is releasing ~Oct 2026 and will be available for purchase on O’Reilly Media. We’re also working on a future primer that explores AI-supported incident handling in-depth — stay tuned!&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Tue, 15 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/developers-practitioners/cloud-reliability-incident-handling-best-practices/</guid><category>DevOps &amp; SRE</category><category>Management Tools</category><category>Google Cloud Consulting</category><category>Developers &amp; Practitioners</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/sre.max-600x600.jpg" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Best practices for handling cloud reliability incidents</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/sre.max-600x600.jpg</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/developers-practitioners/cloud-reliability-incident-handling-best-practices/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Flemming Christensen</name><title>Product Manager &amp; Technical Solutions Engineer, Google Cloud</title><department></department><company></company></author></item><item><title>Introducing the Google Cloud Developer Plugin for AI Coding Agents</title><link>https://cloud.google.com/blog/topics/developers-practitioners/introducing-the-google-cloud-developer-plugin-for-ai-coding-agents/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Agent skills fit well alongside documentation and remote MCP servers as ways of enabling the success of your AI workflows. They reduce context window usage for certain use cases, and they're straightforward to install. However, you might have noticed that managing individual skills can be unwieldy, or that some skills are most useful when they act alongside other skills or MCP servers toward the same goal. That's where plugins come in to help.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Today, we're thrilled to announce a new Google Cloud plugin for AI coding agents! Designed as installable bundles, agent plugins equip the AI agent of your choice with skills and tools to be more effective on Google Cloud.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Solving the tool coupling problem&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;As you expand your usage of coding agents, you might find that they become significantly more capable when they use related skills in tandem or with complementary context and tooling. For example, an agent analyzing infrastructure is more effective when combining domain knowledge, workflow recommendations, and the ability to interact with a live environment together.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Plugins solve this coupling challenge by packaging related capabilities into cohesive, installable bundles. This allows you to take advantage of both broad foundational capabilities and deep, product-specific tools without managing complex dependencies. In the &lt;/span&gt;&lt;a href="https://g.dev/cloud/agent-plugins" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Agent Skills repository&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, you'll start to see the following for Google Cloud popping up over time:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Foundational plugins:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Essential platform-wide guidance for things like documentation discovery, project configuration, and architectural design.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Domain-focused plugins:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Specialized knowledge and best practices for technical areas in the context of Google Cloud.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For this release, we've started with a foundational plugin that supports agent functionality for all Google Cloud users, focusing on making it easier for agents to retrieve Google Cloud-related skills, make use of official documentation, and handle programmatic interactions with Google Cloud.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Built on an open standard&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We've also built our plugin in compliance with the &lt;/span&gt;&lt;a href="https://g.dev/cloud/agent-plugins-specification" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Agent Plugins specification&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, an open, vendor-neutral standard for packaging Agent Skills and Model Context Protocol (MCP) servers into portable, interoperable units. Rather than requiring developers to maintain different configurations and wrappers for every AI assistant, the Agent Plugins standard provides a unified manifest and directory structure.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Our Google Cloud plugin adopts this standard to ensure that developers across a variety of AI coding environments get consistent, high-quality access to tools that help them succeed with Google Cloud. That includes not only the plugins we talk about today, but all other plugins published to the Google Agent Skills repository as well.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Let's take a look at the flagship plugin that we've just published in the Google Agent Skills repository: &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;google-cloud-developer&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;. This plugin exists to help agents successfully navigate the fundamentals of interacting with Google Cloud: things like authentication, authorization, managing projects, and guardrails for gcloud CLI operations. This plugin also bundles configuration for the &lt;/span&gt;&lt;a href="https://g.dev/cloud/dk-mcp-connect" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Developer Knowledge MCP server&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, which gives agents up-to-date grounding in Google's official developer documentation.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Plugin in action: Project onboarding and identity authentication&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To see how this plugin works, consider a situation where you're bootstrapping a new project as part of working on a script. With the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;google-cloud-developer&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; plugin installed, you can prompt your agent:&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;I'm brand new to this platform, and I need to get an account and a first project with billing set up. Then, I need my local machine authenticated so a script that I'm writing can call the APIs as a service identity instead of as me.&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Environment awareness:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The agent silently runs background checks against your live environment for prerequisites like CLI availability and potential existing projects or organizations.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Review:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The agent considers IAM best practices to avoid risks that might be assumed as part of the prompt, like accidental key leaks or git commits.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Interaction with guardrails:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The agent outlines a workflow roadmap and offers to act on those steps before modifying any resources.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/screenshot_plugin_blog_post.max-1000x1000.png"
        
          alt="screenshot_plugin_blog_post"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Installing Google Cloud plugins&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Because Google Cloud plugins are available from the open &lt;/span&gt;&lt;a href="https://g.dev/cloud/agent-plugins" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Agent Skills&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; repository and adhere to the standard Agent Plugins layout, adding them to your environment is straightforward. For example, here's how you'd install the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;google-cloud-developer&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; plugin:&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Antigravity CLI&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Install the plugin directly via the CLI using its path in the Google Agent Skills repository:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;agy plugin install https://github.com/google/skills/plugins/cloud/google-cloud-developer&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e679959d0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;p data-block-key="9v8as"&gt;Enable the Developer Knowledge API in your Google Cloud project by using the &lt;a href="https://docs.cloud.google.com/sdk/docs/install-sdk"&gt;gcloud CLI&lt;/a&gt;:&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;gcloud services enable developerknowledge.googleapis.com --project=&amp;lt;YOUR_PROJECT_ID&amp;gt;&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e75118e10&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Claude Code&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Add the Google plugins marketplace, then install the plugin:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;claude plugin marketplace add google/skills\r\nclaude plugin install google-cloud-developer@google-plugins&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e67c35510&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;p data-block-key="9v8as"&gt;Enable the Developer Knowledge API in your Google Cloud project by using the &lt;a href="https://docs.cloud.google.com/sdk/docs/install-sdk"&gt;gcloud CLI&lt;/a&gt;:&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;gcloud services enable developerknowledge.googleapis.com --project=&amp;lt;YOUR_PROJECT_ID&amp;gt;&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e673847d0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;p data-block-key="9v8as"&gt;Create an API key for the Developer Knowledge API by following the instructions &lt;a href="https://developers.google.com/knowledge/quickstart#create-secure-key" target="_blank"&gt;Create and secure the API key&lt;/a&gt;. Save the key you create in a secure location.&lt;/p&gt;&lt;p data-block-key="10jks"&gt;Export the API key to your environment before starting Claude Code:&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;export DEVELOPERKNOWLEDGE_API_KEY=&amp;lt;YOUR_API_KEY&amp;gt;&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e67385e50&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;strong style="vertical-align: baseline;"&gt;Note:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; The &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;DEVELOPERKNOWLEDGE_API_KEY&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; environment variable needs to be set in the environment before you start Claude Code. Consider adding this export to your shell's startup script (e.g. &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;.bashrc&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;.zshrc&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;) for convenience.&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Codex CLI&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Add the Google plugins marketplace, then install the plugin:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;codex plugin marketplace add google/skills\r\ncodex plugin add google-cloud-developer@google-plugins&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e67317ad0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;p data-block-key="9v8as"&gt;Enable the Developer Knowledge API in your Google Cloud project by using the &lt;a href="https://docs.cloud.google.com/sdk/docs/install-sdk"&gt;gcloud CLI&lt;/a&gt;:&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;gcloud services enable developerknowledge.googleapis.com --project=&amp;lt;YOUR_PROJECT_ID&amp;gt;&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e66fdfa90&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;p data-block-key="9v8as"&gt;Create an API key for the Developer Knowledge API by following the instructions &lt;a href="https://developers.google.com/knowledge/quickstart#create-secure-key" target="_blank"&gt;Create and secure the API key&lt;/a&gt;. Save the key you create in a secure location.&lt;/p&gt;&lt;p data-block-key="8ch63"&gt;Enable authenticated access to the Developer Knowledge MCP server by updating ~/.codex/config.toml (or your project's .codex/config.toml) to include the following lines:&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;[mcp_servers.developer-knowledge]\r\n  url = &amp;quot;https://developerknowledge.googleapis.com/mcp&amp;quot;\r\n  env_http_headers = { &amp;quot;X-Goog-Api-Key&amp;quot; = &amp;quot;DEVELOPERKNOWLEDGE_API_KEY&amp;quot; }&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e66fdc610&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph"&gt;&lt;p data-block-key="9v8as"&gt;Export the API key to your environment before starting Codex:&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;export DEVELOPERKNOWLEDGE_API_KEY=&amp;lt;YOUR_API_KEY&amp;gt;&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e66fdda90&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;strong&gt;Note&lt;/strong&gt;: The &lt;code&gt;DEVELOPERKNOWLEDGE_API_KEY&lt;/code&gt; environment variable needs to be set in the environment before you start Codex. Consider adding this export to your shell's startup script (e.g. &lt;code&gt;.bashrc&lt;/code&gt;, &lt;code&gt;.zshrc&lt;/code&gt;) for convenience.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Note:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; After performing this configuration update, a status of &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;not logged in&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; for the Developer Knowledge MCP server is expected and doesn't block access to the server.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Next Steps&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If you're already a Google Cloud user, try the above installation steps to set up your agent for success. We think you'll like what you see! For those who want a more guided approach, our new &lt;/span&gt;&lt;a href="https://g.dev/cloud/agent-plugins-codelab-agy" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;codelab&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; will walk you through the installation and initial exploration of the plugin in Antigravity.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If you're new to Google Cloud, you can also get started with instructions &lt;/span&gt;&lt;a href="https://g.dev/cloud/dev-setup" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;in our documentation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; to set yourself up for local development.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The most curious readers can also take a deeper look at the plugins and agent skills available to use today in the &lt;/span&gt;&lt;a href="https://g.dev/cloud/agent-plugins" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Agent Skills&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; repository.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Thu, 10 Sep 2026 10:53:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/developers-practitioners/introducing-the-google-cloud-developer-plugin-for-ai-coding-agents/</guid><category>Developers &amp; Practitioners</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/agent-plugins-blog-cover.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Introducing the Google Cloud Developer Plugin for AI Coding Agents</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/agent-plugins-blog-cover.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/developers-practitioners/introducing-the-google-cloud-developer-plugin-for-ai-coding-agents/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Jonathan Lee</name><title>Content Strategist</title><department></department><company></company></author></item><item><title>Power agent hubs or custom harnesses with the Antigravity SDK in one toolkit</title><link>https://cloud.google.com/blog/topics/developers-practitioners/power-agent-hubs-or-custom-harnesses-with-the-antigravity-sdk/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Enterprise agent adoption isn’t one-size-fits-all. While many teams will opt for managed commercial platforms, such as &lt;/span&gt;&lt;a href="https://cloud.google.com/products/gemini-enterprise-agent-platform"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Gemini Enterprise Agent Platform&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for turnkey agent deployment and governance, developers with bespoke workflows or custom execution engines often choose to build their own lightweight agent hubs.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;If you are building a centralized agent hub from the ground up, you need tools that run predictably, log everything, and stay in their sandbox. The &lt;/span&gt;&lt;a href="https://antigravity.google/product/antigravity-sdk" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Antigravity SDK&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; gives you the exact runtime engine used in Antigravity 2.0 and the Antigravity CLI, adding declarative safety policies, real-time telemetry, and stateful multi-turn persistence straight into your application. When the core runtime updates, your SDK agents get those optimizations automatically. That's why today, we're breaking down how the Antigravity SDK powers a complete multi-agent control plane.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;How Antigravity comes together&lt;/span&gt;&lt;/h3&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/01-agy-harness.max-1000x1000.png"
        
          alt="01-agy-harness"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;A multi-agent control plane monitors and manages LLM workloads. It shows you exactly what the agent is thinking, which tools it calls, and how it stores state.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/02-agents-dashboard.max-1000x1000.png"
        
          alt="02-agents-dashboard"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;It consists of two critical components:&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;1. Antigravity SDK agent core&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: The runtime that manages model interactions (like Gemini 3.1 Pro and Gemini 3.8 Flash), runs tools, generates thinking traces, and executes skills.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;2. Observability and telemetry middleware&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: An event-driven layer powered by Antigravity SDK Lifecycle Hooks. It intercepts agent actions like step starts, thinking updates, and tool calls, and streams telemetry over WebSockets to your dashboard.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Use case: Multi-agent monitoring and interactive control&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Let's explore a scenario where an organization is building or maintains a custom agent hub and wants to integrate Antigravity SDK-powered agents. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;The problem&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: An operations engineer needs to monitor multiple active agents (e.g., &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;gemini-pro-agent&lt;/code&gt;&lt;code style="vertical-align: baseline;"&gt;, &lt;/code&gt;&lt;code style="vertical-align: baseline;"&gt;github-agent&lt;/code&gt;&lt;code style="vertical-align: baseline;"&gt;, &lt;/code&gt;&lt;code style="vertical-align: baseline;"&gt;email-agen&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;t&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;) performing background research, document summarization, and task scheduling. Traditionally, observing agent progress requires:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Tailing fragmented console logs across multiple terminal windows&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Manually inspecting JSON transcripts to diagnose stuck or failing tool calls&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Lack of visibility into which Skills or MCP connectors are loaded for a given agent session&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Difficulty tracking cumulative token usage and execution latency&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;The solution&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: This post walks through each one: the streaming API for real-time observation, lifecycle hooks for telemetry and interception, the policy engine for steering, skills for capability management, and session state for persistence. With an SDK-powered dashboard, operators get a single view into what every agent is doing.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/03-agy-bespoke-agent-hub.max-1000x1000.jpg"
        
          alt="03-agy-bespoke-agent-hub"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;What happens behind the scenes?&lt;br/&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;When an operator or dashboard interacts with an Antigravity agent, the runtime coordinates execution through five core mechanisms:&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Session initialization and state attachment (&lt;/strong&gt;&lt;strong style="vertical-align: baseline;"&gt;save_dir&lt;/strong&gt;&lt;strong style="vertical-align: baseline;"&gt; &amp;amp; &lt;/strong&gt;&lt;strong style="vertical-align: baseline;"&gt;conversation_id&lt;/strong&gt;&lt;strong style="vertical-align: baseline;"&gt;):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;The runtime initializes or reattaches to a session, binding execution to a root &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;save_dir&lt;/code&gt;&lt;code style="vertical-align: baseline;"&gt;.&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; Multi-turn trajectory logs, tool receipts, and artifacts are preserved under &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;traj-&amp;lt;conversation_id&amp;gt;&lt;/code&gt;&lt;code style="vertical-align: baseline;"&gt; &lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;for persistent auditability.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Skill resolution (&lt;/strong&gt;&lt;strong style="vertical-align: baseline;"&gt;skills_paths&lt;/strong&gt;&lt;strong style="vertical-align: baseline;"&gt;):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Domain-specific capabilities and instructions are resolved directly from filesystem paths pointing to &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;SKILL.md&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; bundles, dynamically augmenting the agent's system prompt without an external registry.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Concurrent stream generation (&lt;/strong&gt;&lt;strong style="vertical-align: baseline;"&gt;ChatResponse&lt;/strong&gt;&lt;strong style="vertical-align: baseline;"&gt;):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;The runtime exposes three concurrent async iterators over the single model response:&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;ul&gt;
&lt;li aria-level="2" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;code style="vertical-align: baseline;"&gt;response&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; (yields visible text tokens)&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="2" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;code style="vertical-align: baseline;"&gt;response.thoughts&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; (yields internal chain-of-thought reasoning deltas)&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="2" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;code style="vertical-align: baseline;"&gt;response.tool_calls&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; (yields typed &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;ToolCall&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; events containing &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;.name&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; and &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;.args&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;)&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Declarative sandboxing and built-in tool execution:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;When the agent performs workspace operations, built-in tools (&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;list_directory&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;find_file&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;search_directory&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;view_file&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;create_file&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;, &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;edit_file&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;) execute strictly within configured &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;workspaces&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; directories governed by safety policies (such as &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;policy.workspace_only()&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;).&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Telemetry interception via lifecycle hooks:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Decorated async hook functions (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;@hooks.on_session_start&lt;/code&gt;&lt;code style="vertical-align: baseline;"&gt;, &lt;/code&gt;&lt;code style="vertical-align: baseline;"&gt;@hooks.pre_tool_call_decide&lt;/code&gt;&lt;code style="vertical-align: baseline;"&gt;, &lt;/code&gt;&lt;code style="vertical-align: baseline;"&gt;@hooks.post_tool_call&lt;/code&gt;&lt;code style="vertical-align: baseline;"&gt;, &lt;/code&gt;&lt;code style="vertical-align: baseline;"&gt;@hooks.on_session_end&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;) intercept agent transitions in real time, validating or modifying tool calls and broadcasting telemetry payloads over WebSockets to the live dashboard.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The Antigravity SDK organizes these responsibilities into four core building blocks:&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;1. Modular capabilities with Skills&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Skills provide reusable, domain-specific instruction bundles and reference assets that agents load dynamically. Rather than managing an in-memory registry, skills are resolved directly from filesystem directories containing a &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;SKILL.md&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; file:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;from google.antigravity import Agent, LocalAgentConfig\r\n\r\n# Pass directory paths containing SKILL.md bundles directly to config.\r\n# The runtime dynamically resolves and injects them into the prompt.\r\nconfig = LocalAgentConfig(\r\n    model=&amp;quot;gemini-3.8-flash&amp;quot;,\r\n    system_instructions=(\r\n        &amp;quot;You are an enterprise operations assistant equipped with &amp;quot;\r\n        &amp;quot;specialized operational skills.&amp;quot;\r\n    ),\r\n    skills_paths=[&amp;quot;./skills/research&amp;quot;, &amp;quot;./skills/code_review&amp;quot;],\r\n)\r\n\r\nasync with Agent(config) as agent:\r\n    response = await agent.chat(&amp;quot;Analyze the deployment logs.&amp;quot;)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e743be310&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;2. Sandboxed built-in tools and workspace scoping&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The SDK provides production-ready file and workspace tools out of the box, which removes the need to write custom filesystem wrappers. When paired with&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt; &lt;/code&gt;&lt;code style="vertical-align: baseline;"&gt;workspaces&lt;/code&gt;&lt;code style="vertical-align: baseline;"&gt; &lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;and declarative safety policies, tools are strictly confined to authorized directories:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;from google.antigravity import Agent, LocalAgentConfig, types\r\nfrom google.antigravity.policies import policy\r\n\r\nconfig = LocalAgentConfig(\r\n    model=&amp;quot;gemini-3.8-flash&amp;quot;,\r\n    # Selectively enable built-in tools via CapabilitiesConfig\r\n    capabilities=types.CapabilitiesConfig(\r\n        enabled_tools=[\r\n            types.BuiltinTools.LIST_DIR,       # &amp;quot;list_directory&amp;quot;\r\n            types.BuiltinTools.FIND_FILE,      # &amp;quot;find_file&amp;quot;\r\n            types.BuiltinTools.SEARCH_DIR,     # &amp;quot;search_directory&amp;quot;\r\n            types.BuiltinTools.VIEW_FILE,      # &amp;quot;view_file&amp;quot;\r\n            types.BuiltinTools.CREATE_FILE,    # &amp;quot;create_file&amp;quot;\r\n            types.BuiltinTools.EDIT_FILE,      # &amp;quot;edit_file&amp;quot;\r\n        ]\r\n    ),\r\n    # Enforce filesystem isolation: operations outside these paths are blocked\r\n    workspaces=[&amp;quot;./workspace&amp;quot;],\r\n    policies=[policy.workspace_only()],\r\n)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e75212250&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;3. Session isolation and trajectory persistence (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;save_dir&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; &amp;amp; &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;conversation_id&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;)&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;State persistence in the Antigravity SDK is managed through declarative configuration rather than an external database. Specifying a &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;save_dir&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; establishes a root directory where full turn trajectories, tool receipts, and artifacts are preserved under &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;traj-&amp;lt;conversation_id&amp;gt;&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;from google.antigravity import Agent, LocalAgentConfig\r\n\r\nconfig = LocalAgentConfig(\r\n    model=&amp;quot;gemini-3.8-flash&amp;quot;,\r\n    # Root directory storing all conversation trajectories\r\n    save_dir=&amp;quot;./storage/sessions&amp;quot;,\r\n    # Supply conversation_id to reattach to an existing trajectory;\r\n    # omit it to let the SDK mint a new ID on the first turn.\r\n    conversation_id=&amp;quot;ops-session-20260820-001&amp;quot;,\r\n)\r\n\r\nasync with Agent(config) as agent:\r\n    # Resumes prior context and continues the multi-turn session seamlessly\r\n    response = await agent.chat(&amp;quot;Summarize the issues identified in the last turn.&amp;quot;)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e75213e50&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;4. Real-time telemetry and interception with lifecycle hooks&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Lifecycle hooks allow dashboards and monitoring engines to observe and steer every stage of execution. Using decorated async functions, you can stream status updates over WebSockets, inspect tool parameters, and enforce human-in-the-loop approvals before tools run:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;from google.antigravity import Agent, LocalAgentConfig, types\r\nfrom google.antigravity.hooks import hooks\r\n\r\n# 1. Session start &amp;amp; end telemetry\r\n@hooks.on_session_start\r\nasync def on_session_start():\r\n    broadcast_to_dashboard({&amp;quot;type&amp;quot;: &amp;quot;STATUS&amp;quot;, &amp;quot;status&amp;quot;: &amp;quot;RUNNING&amp;quot;})\r\n\r\n@hooks.on_session_end\r\nasync def on_session_end():\r\n    broadcast_to_dashboard({&amp;quot;type&amp;quot;: &amp;quot;STATUS&amp;quot;, &amp;quot;status&amp;quot;: &amp;quot;IDLE&amp;quot;})\r\n\r\n# 2. Intercept tool calls before execution (human-in-the-loop / audit gate)\r\n@hooks.pre_tool_call_decide\r\nasync def intercept_tool(tool_call: types.ToolCall) -&amp;gt; types.HookResult:\r\n    broadcast_to_dashboard({\r\n        &amp;quot;type&amp;quot;: &amp;quot;TOOL_CALL&amp;quot;,\r\n        &amp;quot;tool&amp;quot;: tool_call.name,\r\n        &amp;quot;args&amp;quot;: tool_call.args,\r\n    })\r\n    # Return HookResult to approve or block execution\r\n    return types.HookResult(allow=True)\r\n\r\n# 3. Post-execution tool receipts\r\n@hooks.post_tool_call\r\nasync def record_tool_result(result):\r\n    broadcast_to_dashboard({\r\n        &amp;quot;type&amp;quot;: &amp;quot;TOOL_RESULT&amp;quot;,\r\n        &amp;quot;tool&amp;quot;: result.name,\r\n        &amp;quot;error&amp;quot;: getattr(result, &amp;quot;error&amp;quot;, None),\r\n    })\r\n\r\nconfig = LocalAgentConfig(\r\n    model=&amp;quot;gemini-3.8-flash&amp;quot;,\r\n    hooks=[on_session_start, on_session_end, intercept_tool, record_tool_result],\r\n)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e752110d0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Get started&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Get started with your own enterprise agent control plane using the following resources:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://antigravity.google/docs/sdk/overview" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Antigravity SDK Quick Start&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;a href="https://github.com/google-antigravity/antigravity-sdk-python" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Antigravity github repository&lt;/span&gt;&lt;/a&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Tue, 08 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/developers-practitioners/power-agent-hubs-or-custom-harnesses-with-the-antigravity-sdk/</guid><category>Developers &amp; Practitioners</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Power agent hubs or custom harnesses with the Antigravity SDK in one toolkit</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/developers-practitioners/power-agent-hubs-or-custom-harnesses-with-the-antigravity-sdk/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Wei Yih Yap</name><title>Forward Deployed Engineer, Google Cloud</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Paul Datta</name><title>Practice Customer Engineer, Google Cloud</title><department></department><company></company></author></item><item><title>Spanner migrations: Automating dual-write with Antigravity CLI for minimal disruption</title><link>https://cloud.google.com/blog/topics/developers-practitioners/using-antigravity-cli-to-streamline-dual-write-database-migration/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When Google's Finance Engineering team needed to modernize their legacy data layer, they chose &lt;/span&gt;&lt;a href="https://cloud.google.com/spanner?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Spanner&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;, a globally distributed, strongly consistent, multi-model database with high availability capabilities. But migrating to Spanner without taking production services offline was a daunting engineering challenge: As the internal team responsible for the application, we needed to manually rewrite dual-write logic across dozens of Data Access Objects (DAOs), a process that is slow and prone to human error. Further, doing so without disruption would have required implementing multi-phase dual-write architectures across every DAO in our codebase. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To solve this, we took an alternative approach: We built an automated refactoring pipeline powered by Antigravity CLI in headless mode. This helped us accelerate our migration velocity significantly while maintaining strict data parity in our staging environments as we prepare for production. &lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;The challenge: Anatomy of a dual-write migration&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When migrating high-throughput production services where financial accuracy is essential, simple cutover scripts do not work. You must verify that both the legacy datastore and Spanner receive identical writes simultaneously until all the historical data backfills and verifications are complete.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We structured our migration across three distinct phases:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Historical backfill:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Copying existing historical records to Spanner while maintaining referential integrity.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Dual-write / dual-read implementation:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Modifying every DAO to write mutations to both the primary store and Cloud Spanner in parallel during the migration window.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Automated API verification and parity checking:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Intercepting RPC traffic and verifying end-to-end that every write lands with byte-for-byte equivalence across both stores.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/1_-_Dual_Write_Architecture.max-1000x1000.png"
        
          alt="1 - Dual Write Architecture"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The architectural pattern is clean, but at our scale, we began to encounter friction. That’s because each DAO requires:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;A dedicated &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;MutationConverter&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; class mapping complex domain models to Spanner schema columns&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;Dual-write branch handling and rollback or error-reporting logic&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;A suite of unit tests verifying both primary and Spanner writes using fake time sources and test doubles (&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;FakeTimeSource&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;)&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Performing these identical, high-precision code changes across 30+ DAOs by hand would have taken months of engineering time.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;The solution: Standardized mutation converter patterns&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To verify that our automation pipeline could reliably generate clean code, we first standardized our DAO refactoring pattern around a decoupled &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;MutationConverter&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; interface.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Instead of embedding raw Spanner table names and column assignments directly inside core DAO business logic, we isolate Spanner schema translation into dedicated converter units:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;// Example of the standardized pattern generated by our pipeline\r\n\r\ntype BpcTransferAmountsMutationConverter interface {\r\n    ToInsertMutation(entity *model.BpcTransferAmount) (*spanner.Mutation, error)\r\n    ToUpdateMutation(entity *model.BpcTransferAmount) (*spanner.Mutation, error)\r\n}\r\n\r\ntype bpcTransferAmountsMutationConverterImpl struct {\r\n    tableName string\r\n}\r\n\r\nfunc (c *bpcTransferAmountsMutationConverterImpl) ToInsertMutation(entity *model.BpcTransferAmount) (*spanner.Mutation, error) {\r\n    if entity == nil {\r\n        return nil, errors.New(&amp;quot;entity cannot be nil&amp;quot;)\r\n    }\r\n    \r\n    // Map domain fields to Cloud Spanner table schema\r\n    cols := []string{&amp;quot;TransferId&amp;quot;, &amp;quot;AmountCents&amp;quot;, &amp;quot;CurrencyCode&amp;quot;, &amp;quot;LastModifiedTimestamp&amp;quot;}\r\n    vals := []interface{}{\r\n        entity.TransferId,\r\n        entity.AmountCents,\r\n        entity.CurrencyCode,\r\n        spanner.CommitTimestamp, // Use Spanner commit timestamps\r\n    }\r\n    \r\n    return spanner.Insert(c.tableName, cols, vals), nil\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e753f6910&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By establishing a rigid, deterministic contract between the DAO and the Spanner SDK (&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;spanner.Mutation&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;), we created an exact target specification that an AI coding agent could reason about and generate reliably.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Why use Antigravity&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;CLI in headless mode?&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Interactive AI chat interfaces in IDEs work well for exploratory coding, but they are poorly suited for systematic, multi-file code updates across an entire codebase. When you need to apply repeatable refactoring to dozens of targets without missing edge cases, you need automated workflows.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We addressed this by building an orchestration script (&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;migration_ui.py&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;) that runs Antigravity CLI in headless mode (&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;-p&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;).&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Headless mode lets Antigravity run directly inside shell scripts, continuous integration pipelines, and background automation jobs without requiring manual terminal prompts. This approach helped us scale our work in three key ways:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Deterministic prompt architectures:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; We treated our prompts as version-controlled engineering artifacts. We codified precise rules handling common Spanner edge cases — such as timestamp serialization, nullability conversions, mutation ambiguity, and &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;FakeTimeSource&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; test injection — directly into reusable prompt templates.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Batch execution and automated verification:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Our orchestration script takes a target DAO name as input, retrieves the existing single-write source code and schema, and feeds it to headless Antigravity alongside our structural conventions. Antigravity generates the new converter, the refactored dual-write DAO, and corresponding unit tests. The script then runs &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;blaze test&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;. If a linter error or test assertion fails, the error log feeds directly back into Antigravity for self-correction.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Overnight execution at scale:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Because the loop runs unattended, engineers can queue up 10 DAOs at the end of the day. By morning, the pipeline generates, tests, and validates 10 clean changelists ready for human code review.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Results and key takeaways for cloud engineers&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Combining Spanner's distributed database primitives with Antigravity CLI's headless automation produced clear benefits across our engineering organization:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Significant reduction in migration effort&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: DAO dual-write migrations that previously required extensive manual coding and testing were completed and reviewed in a fraction of the time &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Highly reliable data migration:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Because every generated DAO adhered to the exact same tested &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;MutationConverter&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; pattern and underwent automated unit testing against Spanner test doubles, we sustained high data fidelity during our extensive migration testing. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Focus on higher-value engineering:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Engineers avoided repetitive boilerplate refactoring, giving them time to focus on data modeling, architectural resilience, and performance optimization.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Three tips for your next database migration&lt;/span&gt;&lt;/h3&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Decouple schema translation first:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Before writing migration scripts, define a strict interface (like our &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;MutationConverter&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;) that isolates your new cloud database SDK requirements from your existing business logic. AI agents work best when given clear, bounded design patterns.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Move from interactive chat to headless automation:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; When executing repetitive refactoring across more than three or four files, invest in scripted, headless workflows. Treating prompt inputs and test verifications as automated build steps help maintain quality and consistency.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Let the build system act as your guardrail:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Connect your AI generation loop directly to your build and test harness (&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;bazel test&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; or &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;go test&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;). This lets the model fix compile and assertion errors before a developer reviews the code.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Get started&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Whether you’re migrating financial systems or building cloud-native applications from scratch, Spanner and Antigravity provide a foundation for scalable software development.&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Explore Cloud Spanner:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Learn more about Spanner's distributed architecture &lt;/span&gt;&lt;a href="https://cloud.google.com/spanner/docs"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Spanner documentation&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Discover Gemini for Developers:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; See how AI-assisted coding and headless CLI automation can assist your engineering workflows at &lt;/span&gt;&lt;a href="https://cloud.google.com/use-cases/ai-for-developers"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud AI for Developers&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Fri, 04 Sep 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/developers-practitioners/using-antigravity-cli-to-streamline-dual-write-database-migration/</guid><category>AI &amp; Machine Learning</category><category>Cloud Migration</category><category>Developers &amp; Practitioners</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Spanner migrations: Automating dual-write with Antigravity CLI for minimal disruption</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/developers-practitioners/using-antigravity-cli-to-streamline-dual-write-database-migration/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Sachin Mathapati</name><title>Application Engineer</title><department></department><company></company></author></item><item><title>Not All LLM Workloads Are Equal: Benchmarking TPU Performance on Classification vs. Generation</title><link>https://cloud.google.com/blog/topics/developers-practitioners/not-all-llm-workloads-are-equal-benchmarking-tpu-performance-on-classification-vs-generation/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Moving Large Language Models (LLMs) from experimental prototypes into enterprise production exposes a critical truth: your infrastructure dictates both your performance ceilings and your unit economics. Standard hardware benchmarks often ignore a fundamental reality—not all LLM requests stress the silicon in the same way. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In this post, we dive into a comprehensive benchmarking exercise comparing Gemma 3 12B and Gemma 3 27B on Google Cloud TPU v6e to answer a crucial architectural question: &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;How does TPU infrastructure actually perform when tasked with structurally distinct workloads at scale?&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Key Findings and Suggestions&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Before diving into the methodology, here are the critical takeaways for architects deploying Gemma 3 on TPU v6e:&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;The Generation Performance Wall&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For decode-heavy generation tasks, the Gemma 3 27B model hits a strict performance wall past 64 concurrent users, plateauing at a 4.12x normalized throughput multiplier at 128 users. In contrast, the 12B model scales up to an 8.19x multiplier. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;strong&gt;Suggestion&lt;/strong&gt;: If your workload requires high-concurrency generation, downsize to the 12B model, or set strict pod-autoscaling limits capping concurrent requests at 64 per replica for the 27B model.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;The Classification Parity&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For prefill-heavy classification tasks, model parameter size matters significantly less. Both the 12B and 27B models achieve similar peak scaling (around 6.0x to 6.4x normalized throughput at 128 users) without saturating the TPUs.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;strong&gt;Suggestion&lt;/strong&gt;: You can safely deploy larger, more capable models for summarization or classification workflows without paying a throughput penalty. The average --max-num-seqs or --max-model-len should be kept judiciously based on the average user load and average tokens per request, without which there might be request drops.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Designing Around the Wall&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Hardware saturation manifests as severe latency spikes and silent request dropouts. To mitigate this, do not rely on standard CPU/Memory scaling triggers. Instead, scale based on  End-to-End (E2E) latency metrics, and implement aggressive vLLM bucket padding optimizations (VLLM_TPU_BUCKET_PADDING_GAP) to conserve memory.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;The Architecture Setup&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The inference stack can be divided into three core pillars:&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;1. Infrastructure: GKE &amp;amp; TPU&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The foundation of our deployment is a Google Kubernetes Engine (GKE) Autopilot cluster. Connected to this is a single-host TPU v6e node pool configured with a 2x2 chip topology.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;2. Software &amp;amp; Tools: vllm&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For the serving framework, we leveraged vllm via &lt;/span&gt;&lt;a href="https://github.com/vllm-project/tpu-inference" rel="noopener" target="_blank"&gt;&lt;span style="vertical-align: baseline;"&gt;vllm-project/tpu-inference&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;3. Models: Gemma 3 12B and 27B&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We evaluated two highly capable open-weights models: Gemma 3 12B and Gemma 3 27B. These models were accessed via HuggingFace.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;The Workloads: Classification vs. Generation&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Not all LLM requests stress the system equally. We benchmarked two distinct scenarios: Classification and Generation, across 16, 32, 64, and 128 concurrent users:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Classification (High Input, Low Output):&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; This use case mimics an e-commerce compliance task. The prompt includes large blocks of product rules, item descriptions, and OCR-extracted text. The output is exceptionally small—typically just classifying an item as "Allow" or "Prohibit". Input Sequence Length (ISL) is ~4,000 tokens and Output Sequence Length (OSL) is ~10 tokens.&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Generation (Low/Medium Input, High Output): &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;This use case mimics long-form text generation. The prompt requests a detailed, analytical policy brief on the future of AI in the labor market. The model spends the majority of its time decoding and streaming out hundreds of tokens. Input Sequence Length (ISL) is 500 tokens and Output Sequence Length (OSL) is ~1,000 tokens.&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Results and Observations&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We measured metrics like Throughput (requests/sec), End-to-End Latency and the results provided some fascinating insights into how parameter size and hardware bandwidth interact. To ensure architectural consistency, every benchmark was executed using the &lt;/span&gt;&lt;a href="https://github.com/vllm-project/tpu-inference" rel="noopener" target="_blank"&gt;&lt;span style="vertical-align: baseline;"&gt;vllm-project/tpu-inference&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; hardware plugin, leveraging a standardized global serving configuration of &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;max-model-len=128000&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;,&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; max-num-batched-tokens=8192&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;,&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;and&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; max-num-seqs=512&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Generation Scaling Divergence&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In Generation tasks, both models perform similarly up to 64 concurrent users. However, at 128 concurrent users, the Gemma 3 12B model shows significantly better scaling, achieving an 8.19x normalized throughput multiplier compared to a 4.12x plateau for the Gemma 3 27B model (normalized against the Gemma 3 12B baseline at 16 users). This suggests that the larger 27B model hits memory or compute limits much earlier under high generation loads.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt; &lt;/p&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;&lt;table border="1" style="border-collapse: collapse; width: 96.2054%; height: 206px;"&gt;
&lt;thead&gt;
&lt;tr style="background-color: #d2e3fc; text-align: center;"&gt;
&lt;td style="width: 33.3738%;"&gt;&lt;strong style="vertical-align: baseline;"&gt;Concurrent Users&lt;/strong&gt;&lt;/td&gt;
&lt;td style="width: 33.3738%;"&gt;&lt;strong style="vertical-align: baseline;"&gt;Gemma 3 12B Throughput (req/s)&lt;/strong&gt;&lt;/td&gt;
&lt;td style="width: 33.3738%;"&gt;&lt;strong style="vertical-align: baseline;"&gt;Gemma 3 27B Throughput (req/s)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="width: 33.3738%;"&gt;&lt;span style="vertical-align: baseline;"&gt;16 users&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 33.3738%;"&gt;&lt;span style="vertical-align: baseline;"&gt;1.00 x&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 33.3738%;"&gt;&lt;span style="vertical-align: baseline;"&gt;1.05 x&lt;/span&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="width: 33.3738%;"&gt;&lt;span style="vertical-align: baseline;"&gt;32 users&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 33.3738%;"&gt;&lt;span style="vertical-align: baseline;"&gt;1.98 x&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 33.3738%;"&gt;&lt;span style="vertical-align: baseline;"&gt;1.97 x&lt;/span&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="width: 33.3738%;"&gt;&lt;span style="vertical-align: baseline;"&gt;64 users&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 33.3738%;"&gt;&lt;span style="vertical-align: baseline;"&gt;2.96 x&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 33.3738%;"&gt;&lt;span style="vertical-align: baseline;"&gt;4.00 x&lt;/span&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="width: 33.3738%;"&gt;&lt;span style="vertical-align: baseline;"&gt;128 users&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 33.3738%;"&gt;&lt;span style="vertical-align: baseline;"&gt;8.19 x&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 33.3738%;"&gt;&lt;span style="vertical-align: baseline;"&gt;4.12 x&lt;/span&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/generation_scaling.max-1000x1000.jpg"
        
          alt="generation_scaling"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-aside"&gt;&lt;dl&gt;
    &lt;dt&gt;aside_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;title&amp;#x27;, &amp;#x27;Pro Tip → Metrics Inflation at High Concurrency&amp;#x27;), (&amp;#x27;body&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e7414e5d0&amp;gt;), (&amp;#x27;btn_text&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;href&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;image&amp;#x27;, None)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Classification Performance Parity&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In Classification tasks, there is negligible difference in scaling behavior between the Gemma 3 12B and Gemma 3 27B models. Both models operate efficiently within the hardware's capacity and scale well, reaching peak normalized throughputs of approximately 6.04x to 6.37x at 128 concurrent users (normalized against the Gemma 3 12B baseline at 16 users).&lt;/span&gt;&lt;/p&gt;
&lt;p&gt; &lt;/p&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;&lt;table border="1"&gt;
&lt;thead&gt;
&lt;tr style="text-align: center; background-color: #d2e3fc;"&gt;
&lt;td style="border: 1px solid #000000; padding: 16px;"&gt;&lt;strong style="vertical-align: baseline;"&gt;Concurrent Users&lt;/strong&gt;&lt;/td&gt;
&lt;td style="border: 1px solid #000000; padding: 16px;"&gt;&lt;strong style="vertical-align: baseline;"&gt;Gemma 3 12B Throughput (req/s)&lt;/strong&gt;&lt;/td&gt;
&lt;td style="border: 1px solid #000000; padding: 16px;"&gt;&lt;strong style="vertical-align: baseline;"&gt;Gemma 3 27B Throughput (req/s)&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="border: 1px solid #000000; padding: 16px;"&gt;&lt;span style="vertical-align: baseline;"&gt;16 users&lt;/span&gt;&lt;/td&gt;
&lt;td style="border: 1px solid #000000; padding: 16px;"&gt;&lt;span style="vertical-align: baseline;"&gt;1.00 x&lt;/span&gt;&lt;/td&gt;
&lt;td style="border: 1px solid #000000; padding: 16px;"&gt;&lt;span style="vertical-align: baseline;"&gt;0.76x&lt;/span&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="border: 1px solid #000000; padding: 16px;"&gt;&lt;span style="vertical-align: baseline;"&gt;32 users&lt;/span&gt;&lt;/td&gt;
&lt;td style="border: 1px solid #000000; padding: 16px;"&gt;&lt;span style="vertical-align: baseline;"&gt;1.18x&lt;/span&gt;&lt;/td&gt;
&lt;td style="border: 1px solid #000000; padding: 16px;"&gt;&lt;span style="vertical-align: baseline;"&gt;1.53x&lt;/span&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="border: 1px solid #000000; padding: 16px;"&gt;&lt;span style="vertical-align: baseline;"&gt;64 users&lt;/span&gt;&lt;/td&gt;
&lt;td style="border: 1px solid #000000; padding: 16px;"&gt;&lt;span style="vertical-align: baseline;"&gt;2.04x&lt;/span&gt;&lt;/td&gt;
&lt;td style="border: 1px solid #000000; padding: 16px;"&gt;&lt;span style="vertical-align: baseline;"&gt;3.15x&lt;/span&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="border: 1px solid #000000; padding: 16px;"&gt;&lt;span style="vertical-align: baseline;"&gt;128 users&lt;/span&gt;&lt;/td&gt;
&lt;td style="border: 1px solid #000000; padding: 16px;"&gt;&lt;span style="vertical-align: baseline;"&gt;6.37x&lt;/span&gt;&lt;/td&gt;
&lt;td style="border: 1px solid #000000; padding: 16px;"&gt;&lt;span style="vertical-align: baseline;"&gt;6.04x&lt;/span&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/classification_perf_table_image2.max-1000x1000.png"
        
          alt="classification_perf_table_image2"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Latency Threshold Analysis&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;End-to-End (E2E) latency exhibits different scaling behaviors depending on the model size and task. When using identical serving hyperparameters (--max-num-seqs=512), the Gemma 3 12B model's Classification latency roughly doubles when moving from 32 users to 64 users, indicating resource contention. However, for the larger Gemma 3 27B model, Classification latency remains relatively flat between 32 and 64 users before doubling at the 128-user mark. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt; &lt;/p&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;&lt;table border="1" style="border-collapse: collapse; width: 97.2684%; height: 182px;"&gt;
&lt;thead&gt;
&lt;tr style="background-color: #d2e3fc; text-align: center;"&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;strong&gt;Model&lt;/strong&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;strong&gt;Task&lt;/strong&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;strong&gt;16 Users&lt;/strong&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;strong&gt;32 Users&lt;/strong&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;strong&gt;64 Users&lt;/strong&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;strong&gt;128 Users&lt;/strong&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;Gemma 3 12B&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;Generation&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;1.00x&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;1.13x&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;1.40x&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;1.70x&lt;/span&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;Gemma 3 12B&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;Classification&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;1.00x&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;0.99x&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;1.79x&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;2.90x&lt;/span&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;Gemma 3 27B&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;Generation&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;1.20x&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;1.68x&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;2.93x&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;3.33x&lt;/span&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;Gemma 3 27B&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;Classification&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;1.20x&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;1.95x&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;1.95x&lt;/span&gt;&lt;/td&gt;
&lt;td style="width: 16.6204%;"&gt;&lt;span style="vertical-align: baseline;"&gt;3.88x&lt;/span&gt;&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/latency_threshold_image3.max-1000x1000.png"
        
          alt="latency threshold_image3"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-aside"&gt;&lt;dl&gt;
    &lt;dt&gt;aside_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;title&amp;#x27;, &amp;#x27;A Crucial TPU Optimization Technique&amp;#x27;), (&amp;#x27;body&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e677a5ed0&amp;gt;), (&amp;#x27;btn_text&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;href&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;image&amp;#x27;, None)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Conclusion&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Benchmarking Gemma 3 12B and 27B models on Google Cloud TPU v6e architecture reveals that raw parameter count is not the sole predictor of inference performance; rather, the interaction between the serving framework, hardware topology, and workload token ratios dictates efficiency. For generation tasks (low input, high output), the 12B model proves superior at high concurrency, sustaining an 8.19x relative throughput multiplier where the 27B model saturates at 4.12x. Conversely, for prefill-heavy classification tasks, both models perform similarly, allowing organizations to deploy larger models without a severe scaling penalty. Our evaluation also mapped exact hardware saturation thresholds—such as End-to-End latency doubling at 64 users for classification and hitting a cliff at 128 users for generation—enabling precise, data-driven auto-scaling triggers rather than costly over-provisioning. Ultimately, achieving these peak metrics requires aggressive tuning of vllm parameters, such as adjusting batched tokens and configuring TPU-specific bucket padding to prevent compute waste, proving that cost-effective AI infrastructure must strictly align model selection and serving configurations to the unique input/output profiles of production workloads.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Ready to scale your LLM workloads?&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Don't let unoptimized infrastructure bottleneck your enterprise AI rollouts. Now that you know how different workload shapes impact hardware saturation, it's time to put these insights into practice:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;strong&gt;Use these benchmarks to right-size your production architecture&lt;/strong&gt;. &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;Safely leverage the larger Gemma 3 27B for prefill-heavy classification tasks without a throughput penalty, but consider switching to the 12B model to maintain linear scaling for decode-heavy generation at high concurrency.&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;strong&gt;Deploy using &lt;/strong&gt;&lt;/span&gt;&lt;strong&gt;&lt;a href="https://docs.cloud.google.com/kubernetes-engine/docs/tutorials/serve-vllm-tpu"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Kubernetes Engine (GKE) &lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; with TPU v6e node pools to build a highly scalable, managed AI foundation and dedicated &lt;/span&gt;&lt;a href="https://github.com/vllm-project/tpu-inference" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;vllm-project/tpu-inference&lt;/span&gt;&lt;/a&gt;&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;&lt;strong&gt; hardware plugin&lt;/strong&gt;. Alternatively, you can also deploy via &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/gemini-enterprise-agent-platform/models/open-models/vllm/use-vllm-tpu"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Model Garden on Gemini Enterprise Agent Platform&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; or you can spin up &lt;/span&gt;&lt;a href="https://github.com/AI-Hypercomputer/tpu-recipes/tree/main/inference/trillium/vLLM" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;TPU VMs&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for serving Gemma 3 models.&lt;/span&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Have you encountered similar performance walls in your own production deployments? Share your scaling strategies, ask questions, and join the discussion in the &lt;/span&gt;&lt;a href="https://www.googlecloudcommunity.com/" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Google Cloud Community forums&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt; &lt;/p&gt;&lt;/div&gt;</description><pubDate>Fri, 04 Sep 2026 15:36:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/developers-practitioners/not-all-llm-workloads-are-equal-benchmarking-tpu-performance-on-classification-vs-generation/</guid><category>Developers &amp; Practitioners</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/hero_2_fvBime0.max-600x600.jpg" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Not All LLM Workloads Are Equal: Benchmarking TPU Performance on Classification vs. Generation</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/hero_2_fvBime0.max-600x600.jpg</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/developers-practitioners/not-all-llm-workloads-are-equal-benchmarking-tpu-performance-on-classification-vs-generation/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Rupjit Chakraborty</name><title>AI Engineer</title><department></department><company></company></author></item><item><title>Announcing the Google Gen AI SDK for Kotlin 1.0: Idiomatic multiplatform access to Gemini</title><link>https://cloud.google.com/blog/topics/developers-practitioners/announcing-the-google-gen-ai-sdk-for-kotlin-10-idiomatic-multiplatform-access-to-gemini/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Integrating modern generative AI capabilities into Kotlin applications shouldn't require juggling raw HTTP clients or bridging disparate Java libraries. Today, we're excited to announce the &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;1.0 release of the Google Gen AI SDK for Kotlin&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;google-genai-kotlin&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;). You can dive right into the code, explore runnable samples, and star the project today on &lt;/span&gt;&lt;a href="https://github.com/googleapis/kotlin-genai" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GitHub at &lt;/span&gt;&lt;code style="text-decoration: underline; vertical-align: baseline;"&gt;googleapis/kotlin-genai&lt;/code&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Built from the ground up as a &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Kotlin Multiplatform (KMP)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; library, the SDK brings idiomatic Kotlin paradigms (including first-class &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Coroutines&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;, asynchronous &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;Flow&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; streaming, and immutable data classes with named and default parameters) to developers specifically targeting the &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;JVM&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; (backend services, serverless functions, desktop)&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The SDK provides a unified surface to interact with both the &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Gemini Developer API&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; (Google AI Studio) and the &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Gemini Enterprise Agent Platform&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; (on Google Cloud) with minimal configuration tweaks.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Although this SDK &lt;span style="vertical-align: baseline;"&gt;is a Kotlin Multiplatform library, using it directly from a mobile app is blocked for&lt;/span&gt;&lt;a href="https://github.com/googleapis/kotlin-genai#google-gen-ai-kotlin-sdk" rel="noopener" target="_blank"&gt;&lt;span style="vertical-align: baseline;"&gt; &lt;/span&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;security reasons&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. Instead, use &lt;/span&gt;&lt;a href="https://firebase.google.com/products/firebase-ai-logic" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Firebase AI Logic&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt; for direct access to Gemini models from mobile apps.&lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;1. Getting started: Adding the dependency&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The SDK is published to Maven Central under &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;com.google.genai:google-genai-kotlin&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Kotlin Multiplatform (KMP)&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For multiplatform applications, add the dependency to your &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;commonMain&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; source set:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;// build.gradle.kts\r\nkotlin {\r\n    sourceSets {\r\n        commonMain.dependencies {\r\n            implementation(&amp;quot;com.google.genai:google-genai-kotlin:1.0.0&amp;quot;)\r\n        }\r\n    }\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e670d2590&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Standard JVM projects&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For single-platform Kotlin projects, Gradle automatically selects the optimal variant via Gradle Module Metadata:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;// build.gradle.kts\r\ndependencies {\r\n    implementation(&amp;quot;com.google.genai:google-genai-kotlin:1.0.0&amp;quot;)\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e6705add0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;2. Unary and streaming text generation and chat&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The primary entry point is the &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;Client&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; class. It manages HTTP connections and authentication automatically based on your environment variables (&lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;GEMINI_API_KEY&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; or &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;GOOGLE_API_KEY&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; for Google AI Studio, and &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;GOOGLE_GENAI_USE_ENTERPRISE=true&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; with standard Google Cloud Application Default Credentials).&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Single prompt request with Gemini Flash&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Using Kotlin's &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;use&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; extension ensures the client's underlying network engine and HTTP connections are released cleanly:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;import com.google.genai.kotlin.Client\r\nimport kotlinx.coroutines.runBlocking\r\n\r\nfun main() = runBlocking {\r\n    Client().use { client -&amp;gt;\r\n        val response = client.models.generateContent(\r\n            model = &amp;quot;gemini-flash-latest&amp;quot;,\r\n            text = &amp;quot;Explain quantum entanglement in two sentences.&amp;quot;\r\n        )\r\n\r\n        println(response.text)\r\n    }\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e67059bd0&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Low-latency streaming with Coroutines &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;Flow&lt;/code&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For interactive UIs and responsive CLI tools, &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;generateContentStream&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; returns a cold Kotlin Coroutine &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;Flow&amp;lt;GenerateContentResponse&amp;gt;&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;, delivering token chunks in real time:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;import com.google.genai.kotlin.Client\r\nimport kotlinx.coroutines.runBlocking\r\n\r\nfun main() = runBlocking {\r\n    Client().use { client -&amp;gt;\r\n        val responseFlow = client.models.generateContentStream(\r\n            model = &amp;quot;gemini-flash-latest&amp;quot;,\r\n            text = &amp;quot;Outline the key architectural patterns for microservices on Google Cloud.&amp;quot;\r\n        )\r\n\r\n        responseFlow.collect { chunk -&amp;gt;\r\n            chunk.text?.let { print(it) }\r\n        }\r\n        println()\r\n    }\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e6705ad10&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Multi-turn conversations (chat)&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Managing conversation history manually across request turns can become tedious. The SDK includes a dedicated &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;chats&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; service that automatically maintains context, appends turns, formats conversation history, and handles function calling:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;import com.google.genai.kotlin.Client\r\nimport com.google.genai.kotlin.types.Content\r\nimport com.google.genai.kotlin.types.GenerateContentConfig\r\nimport kotlinx.coroutines.runBlocking\r\n\r\nfun main() = runBlocking {\r\n    Client().use { client -&amp;gt;\r\n        val config = GenerateContentConfig(\r\n            systemInstruction = Content.fromText(&amp;quot;You are an expert Google Cloud Solutions Architect.&amp;quot;)\r\n        )\r\n\r\n        // Create a multi-turn chat session\r\n        val chat = client.chats.create(\r\n            model = &amp;quot;gemini-flash-latest&amp;quot;,\r\n            config = config\r\n        )\r\n\r\n        // Turn 1\r\n        val firstResponse = chat.sendMessage(&amp;quot;We are designing an event-driven ingestion pipeline on Google Cloud.&amp;quot;)\r\n        println(&amp;quot;Gemini: ${firstResponse.text}\\n&amp;quot;)\r\n\r\n        // Turn 2: context from the first turn is included automatically\r\n        val secondResponse = chat.sendMessage(&amp;quot;Which managed messaging service should we choose: Pub/Sub or Kafka?&amp;quot;)\r\n        println(&amp;quot;Gemini: ${secondResponse.text}\\n&amp;quot;)\r\n    }\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e6705b710&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;You can also use &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;chat.sendMessageStream(...)&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt; for streaming multi-turn chat responses.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;3. Multimodal analysis grounded with Google Search&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Gemini's multimodal reasoning is especially effective when combined with external verification. For instance, when analyzing technical, medical, or scientific diagrams, you can attach &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Google Search Grounding&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; to cross-check factual claims against live web sources.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;import com.google.genai.kotlin.Client\r\nimport com.google.genai.kotlin.types.*\r\nimport java.io.File\r\nimport kotlinx.coroutines.runBlocking\r\n\r\nfun main() = runBlocking {\r\n    Client().use { client -&amp;gt;\r\n        val imageBytes = File(&amp;quot;src/main/resources/medical_diagram.png&amp;quot;).readBytes()\r\n\r\n        val content = Content(\r\n            parts = listOf(\r\n                Part(inlineData = Blob(mimeType = &amp;quot;image/png&amp;quot;, data = imageBytes)),\r\n                Part(text = &amp;quot;Is this anatomical diagram accurate? Verify labels against authoritative medical sources.&amp;quot;)\r\n            )\r\n        )\r\n\r\n        // Enable Google Search as a grounding tool\r\n        val config = GenerateContentConfig(\r\n            tools = listOf(Tool(googleSearch = GoogleSearch()))\r\n        )\r\n\r\n        val response = client.models.generateContent(\r\n            model = &amp;quot;gemini-flash-latest&amp;quot;,\r\n            content = content,\r\n            config = config\r\n        )\r\n\r\n        println(&amp;quot;=== Analysis ===&amp;quot;)\r\n        println(response.text)\r\n\r\n        // Inspect citations and search queries\r\n        val grounding = response.groundingMetadata\r\n        println(&amp;quot;\\n=== Search Queries Executed ===&amp;quot;)\r\n        grounding?.webSearchQueries?.forEach { println(&amp;quot;- $it&amp;quot;) }\r\n\r\n        println(&amp;quot;\\n=== Grounding Sources ===&amp;quot;)\r\n        grounding?.groundingChunks?.mapNotNull { it.web }?.forEach { source -&amp;gt;\r\n            println(&amp;quot;- ${source.title}: ${source.uri}&amp;quot;)\r\n        }\r\n    }\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e67059490&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;4. Visual generation and conversational editing: The Gemini 3 image family&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The SDK provides full support for Google's latest image generation models (popularly known as the &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;Nano Banana&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; series of models on leaderboards).&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Generating and Saving an Image&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Generated image bytes are delivered directly in the response parts as a &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;Blob&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;import com.google.genai.kotlin.Client\r\nimport java.io.File\r\nimport kotlinx.coroutines.runBlocking\r\n\r\nfun main() = runBlocking {\r\n    Client().use { client -&amp;gt;\r\n        val response = client.models.generateContent(\r\n            model = &amp;quot;gemini-3.1-flash-image&amp;quot;, // Nano Banana 2\r\n            text = &amp;quot;A photorealistic blueprint of an eco-friendly modern datacenter, isometric view, 4k&amp;quot;\r\n        )\r\n\r\n        val imagePart = response.parts?.firstOrNull { it.inlineData != null }\r\n        imagePart?.inlineData?.data?.let { bytes -&amp;gt;\r\n            File(&amp;quot;datacenter_blueprint.png&amp;quot;).writeBytes(bytes)\r\n            println(&amp;quot;Image generated and saved successfully.&amp;quot;)\r\n        }\r\n    }\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e6705b610&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h3&gt;&lt;span style="vertical-align: baseline;"&gt;Conversational image-to-image editing&lt;/span&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;You can pass existing images and conversational edit instructions in the same request:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;val originalImage = File(&amp;quot;input.png&amp;quot;).readBytes()\r\n\r\nval editPrompt = Content(\r\n    parts = listOf(\r\n        Part(inlineData = Blob(mimeType = &amp;quot;image/png&amp;quot;, data = originalImage)),\r\n        Part(text = &amp;quot;Change the daylight illumination to a dramatic twilight skyline with illuminated windows.&amp;quot;)\r\n    )\r\n)\r\n\r\nval editResponse = client.models.generateContent(\r\n    model = &amp;quot;gemini-3-pro-image&amp;quot;, // Nano Banana Pro\r\n    content = editPrompt\r\n)&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e67058810&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;5. Real-time bidirectional interaction with Gemini Live&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For low-latency voice, audio, and live multimodal interactions, the SDK supports the &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Gemini Live API&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; via persistent WebSocket connections using &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;client.live.connect(...)&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;import com.google.genai.kotlin.Client\r\nimport com.google.genai.kotlin.types.AudioTranscriptionConfig\r\nimport com.google.genai.kotlin.types.LiveConnectConfig\r\nimport kotlinx.coroutines.launch\r\nimport kotlinx.coroutines.runBlocking\r\n\r\nfun main() = runBlocking {\r\n    Client().use { client -&amp;gt;\r\n        val model = if (client.enterprise) &amp;quot;gemini-live-2.5-flash-native-audio&amp;quot;\r\n                    else &amp;quot;gemini-3.1-flash-live-preview&amp;quot;\r\n\r\n        val config = LiveConnectConfig(\r\n            outputAudioTranscription = AudioTranscriptionConfig()\r\n        )\r\n\r\n        // Establish real-time bidirectional WebSocket session\r\n        client.live.connect(model, config).use { session -&amp;gt;\r\n            println(&amp;quot;Connected to Gemini Live session!&amp;quot;)\r\n\r\n            // Launch collector for server messages (audio and text transcriptions)\r\n            val receiveJob = launch {\r\n                session.receive().collect { serverMessage -&amp;gt;\r\n                    serverMessage.serverContent?.outputTranscription?.text?.let { text -&amp;gt;\r\n                        print(text)\r\n                    }\r\n                }\r\n            }\r\n\r\n            // Stream real-time text (or raw PCM audio blobs via session.sendRealtimeInput(audio = ...))\r\n            session.sendRealtimeInput(text = &amp;quot;Hello Gemini! Give me a 5-second motivational quote.&amp;quot;)\r\n\r\n            // When finished, clean up\r\n            receiveJob.cancel()\r\n            session.closeSession()\r\n        }\r\n    }\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e67059950&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;6. Structured tool and function calling&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When building agentic workflows or bridging LLMs with backend microservices, developers can pass structured JSON schemas via &lt;/span&gt;&lt;code style="vertical-align: baseline;"&gt;FunctionDeclaration&lt;/code&gt;&lt;span style="vertical-align: baseline;"&gt;. The model will intelligently select when to invoke the tool:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;val telemetryTool = FunctionDeclaration(\r\n    name = &amp;quot;getDatacenterMetrics&amp;quot;,\r\n    description = &amp;quot;Fetch real-time CPU and thermal telemetry for a Google Cloud region&amp;quot;,\r\n    parameters = Schema(\r\n        type = Type.OBJECT,\r\n        properties = mapOf(&amp;quot;region&amp;quot; to Schema(type = Type.STRING)),\r\n        required = listOf(&amp;quot;region&amp;quot;)\r\n    )\r\n)\r\n\r\nval response = client.models.generateContent(\r\n    model = &amp;quot;gemini-flash-latest&amp;quot;,\r\n    text = &amp;quot;Check telemetry for europe-west1&amp;quot;,\r\n    config = GenerateContentConfig(\r\n        tools = listOf(Tool(functionDeclarations = listOf(telemetryTool)))\r\n    )\r\n)\r\n\r\nresponse.functionCalls?.firstOrNull()?.let { call -&amp;gt;\r\n    println(&amp;quot;Model triggered tool: ${call.name} with arguments: ${call.args}&amp;quot;)\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e67059410&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Additionally, when using the chats service, you can take advantage of Automatic Function Calling (AFC), which means that functions declared in the chat conversation can be invoked automatically and transparently by the SDK on your behalf, as you can see in the following example:&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-code"&gt;&lt;dl&gt;
    &lt;dt&gt;code_block&lt;/dt&gt;
    &lt;dd&gt;&amp;lt;ListValue: [StructValue([(&amp;#x27;code&amp;#x27;, &amp;#x27;fun main() = runBlocking {\r\n    // A mocked function\r\n    val getWeather = callableFunction(&amp;quot;get_weather&amp;quot;, paramName = &amp;quot;city&amp;quot;) { city: String -&amp;gt;\r\n        &amp;quot;18 degrees and sunny in $city&amp;quot;\r\n    }\r\n\r\n    Client().use { client -&amp;gt;\r\n        val chat = client.chats.create(\r\n            model = &amp;quot;gemini-flash-latest&amp;quot;,\r\n            automaticFunctionCalling = AutomaticFunctionCalling(getWeather),\r\n        )\r\n\r\n        // SDK calls get_weather if needed in this conversation\r\n        println(chat.sendMessage(&amp;quot;What is the weather in Zurich?&amp;quot;).text)\r\n    }\r\n}&amp;#x27;), (&amp;#x27;language&amp;#x27;, &amp;#x27;&amp;#x27;), (&amp;#x27;caption&amp;#x27;, &amp;lt;wagtail.rich_text.RichText object at 0x7f7e6705b350&amp;gt;)])]&amp;gt;&lt;/dd&gt;
&lt;/dl&gt;&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;What's next?&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With the 1.0 release of the Google Gen AI SDK for Kotlin, Kotlin developers across backend server ecosystems (Ktor, Spring Boot, Quarkus, Micronaut) now have a clean, multiplatform foundation for building generative AI applications.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To learn more and get started, check out the following resources:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;GitHub Repository:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Check out the source, stars, and discussions at &lt;/span&gt;&lt;a href="https://github.com/googleapis/kotlin-genai" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;github.com/googleapis/kotlin-genai&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Documentation and Samples:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; Explore the &lt;/span&gt;&lt;a href="https://github.com/googleapis/kotlin-genai/tree/main/examples" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;Kotlin Gen AI sample suite&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Feedback:&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; File issues, suggest features, or submit pull requests directly on &lt;/span&gt;&lt;a href="https://github.com/googleapis/kotlin-genai" rel="noopener" target="_blank"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;GitHub&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;We look forward to seeing what you build with Kotlin and Gemini!&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Thu, 03 Sep 2026 09:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/developers-practitioners/announcing-the-google-gen-ai-sdk-for-kotlin-10-idiomatic-multiplatform-access-to-gemini/</guid><category>Developers &amp; Practitioners</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/google-genai-sdk-kotlin.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Announcing the Google Gen AI SDK for Kotlin 1.0: Idiomatic multiplatform access to Gemini</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/google-genai-sdk-kotlin.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/developers-practitioners/announcing-the-google-gen-ai-sdk-for-kotlin-10-idiomatic-multiplatform-access-to-gemini/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Guillaume Laforge</name><title>Developer Advocate</title><department></department><company></company></author></item><item><title>Simplify your resilience testing strategy with Fault Injection Testing</title><link>https://cloud.google.com/blog/products/networking/introducing-google-cloud-fault-injection-testing-in-preview/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When databases fail and network paths falter, you still need your mission-critical cloud services to stay online. Yet guaranteeing high availability has become increasingly difficult because of the complexity of modern distributed systems. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;To help you maintain availability and reliability during adverse events, we’re announcing Fault Injection Testing in preview. Fault Injection Testing is designed to help developers and architects automate failure testing to ensure predictable behavior during disruptions. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;By deliberately introducing faults into your environment, you can verify your safety mechanisms &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;before&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; an actual outage impacts your customers.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Why native resilience testing matters&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Unlike in self-hosted data centers, cloud applications offer less direct access to underlying infrastructure to facilitate failover testing.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Without native tools to prove your application can survive a failure, you risk a critical gap in your reliability strategy that exposes you to several risks:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Damaged trust and reputation&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Frequent failures or poor performance lead to customer dissatisfaction and long-term damage to your brand's image.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Compliance and regulatory penalties&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: For many industries, particularly financial institutions, failing to prove disaster recovery capabilities can lead to non-compliance, audits, and fines.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Migration delays&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Large-scale migrations often stop when teams cannot verify that critical applications will remain stable during a zone failure.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;How Fault Injection Testing works&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Fault Injection Testing allows you to run experiments by creating experiment templates. These templates act as blueprints, defining the specific fault to be injected and the resources that will be targeted for the experiment.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In this public preview, you can test two primary failure scenarios:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Failover Cloud SQL&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: This fault triggers a failover of a high availability Cloud SQL instance from the primary zone to a standby zone.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Degrade application traffic&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: &lt;span style="vertical-align: baseline;"&gt;This allows you to selectively add latency and HTTP error codes through an Application Load Balancer.  &lt;/span&gt;&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Before any fault is injected, Fault Injection Testing performs an automated dry run. This read-only simulation checks your permissions and provides an up-to-date list of every resource that will be affected. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Once you verify the scope, you can manually start the injection. The duration you defined in the template will run its course, and the faults will be reverted at the expiration of the timer.  &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;During the experiment, you can verify that your application is behaving as you planned.  If things do not go as planned, you can use the stop and revert capability to immediately halt the experiment and begin restoring resources to their normal state.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;During preview, we recommend as a best practice to use&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; Fault Injection Testing (&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;FIT) in a non-production environment. Preview is an opportunity to get early access to learn how the service fits and complements your existing testing practices, and to provide us with your feedback to improve the product as well!&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Built for the enterprise&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Partners like KeyBank and Servier are already using Fault Injection Testing to validate their deployments. By using native fault injection, these organizations can approximate demanding failure scenarios — such as zonal outages — to help ensure their services remain stable.&lt;/span&gt;&lt;/p&gt;
&lt;h3&gt;&lt;strong style="vertical-align: baseline;"&gt;Get started with Fault Injection Testing&lt;/strong&gt;&lt;/h3&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Fault Injection Testing is available through the Google Cloud console, the gcloud CLI, and REST APIs.&lt;/span&gt;&lt;/p&gt;
&lt;ol&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Request preview access&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;:&lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt; &lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;Talk to your Google Cloud Account Team to add your project to the preview.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Enable the API&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Search for "Fault Testing API" in your Google Cloud console and select enable.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Assign roles&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Ensure your team has the &lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt;roles/faulttesting.operator&lt;/span&gt;&lt;span style="vertical-align: baseline;"&gt; role to configure and run experiments.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: decimal; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Run your first dry run&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Create a template for a Cloud SQL or load balancer resource in a non-production environment and execute a dry run to see the potential impact.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For more details on implementation, talk to your account team, or view the &lt;/span&gt;&lt;a href="https://docs.cloud.google.com/fault-injection-testing"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;User Guide for Fault Injection Testing&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;</description><pubDate>Wed, 26 Aug 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/networking/introducing-google-cloud-fault-injection-testing-in-preview/</guid><category>Security &amp; Identity</category><category>Developers &amp; Practitioners</category><category>Networking</category><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Simplify your resilience testing strategy with Fault Injection Testing</title><description></description><site_name>Google</site_name><url>https://cloud.google.com/blog/products/networking/introducing-google-cloud-fault-injection-testing-in-preview/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Toby Owen</name><title>Group Product Manager, Google Cloud</title><department></department><company></company></author></item><item><title>How Uber improves network reliability while unblocking cloud migration</title><link>https://cloud.google.com/blog/products/networking/uber-de-risks-hybrid-ai-with-cloud-interconnect/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Uber has a lot in common with the cities it serves. Both are always changing and growing, both must carefully manage the resulting traffic to prevent congestion and sprawl.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Uber has continuously evolved its technical strategies to manage its expanding network, and this careful planning and constant evolution helps ensure that application traffic across its entire platform runs smoothly. Ultimately, maintaining a reliable, high-scale platform that operates seamlessly at any given time is key to preserving user trust.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;One important solution in this effort has been &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/networking/cross-cloud-network-enhancements-for-distributed-workloads/?e=48754805"&gt;&lt;span style="text-decoration: underline; vertical-align: baseline;"&gt;application awareness on Cloud Interconnect&lt;/span&gt;&lt;/a&gt;&lt;span style="vertical-align: baseline;"&gt;. An industry-first tool for application prioritization across hybrid networks, application awareness on Cloud Interconnect has helped Uber prioritize critical traffic to ensure business continuity during potential network congestion events. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Uber acted as an early design partner for application awareness on Cloud Interconnect, helping ensure that this capability met the demands of Uber’s global-scale operations. It not only improved Uber’s daily operations, it also gave Uber the confidence to move forward with a Google Cloud migration, with confidence that there would be less risk of service interruptions during switchovers. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In this post, we’ll explain the features Uber most sought and why, the inner workings of application awareness on Cloud Interconnect, and how it can help other organizations as well.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Prioritizing critical traffic&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;When migrating distributed, hybrid, or multicloud applications at a global scale, network reliability becomes a primary concern. Even the most worthwhile migrations may not seem worth it if such migrations interrupt ongoing service. For organizations like Uber, moving vast amounts of data to support large data analytics workload — including emerging AI use cases — can saturate network links, resulting in increased reliability risk for their critical application traffic. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With standard cloud interconnect approaches, enterprises typically apply simple bandwidth overprovisioning to meet extreme infrastructure needs. But with today's hybrid cloud demands, and given the size of an organization like Uber, overprovisioning network capacity for peak usage is often too costly and unreliable. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;The shortcomings of overprovisioning only become magnified with the integration of cutting-edge AI innovations. Uber needs systems in place that can take on massive data transfers without congesting its network and protecting the performance of business-critical applications.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With the benefit of application awareness on Cloud Interconnect, including the four major features of application awareness — traffic handling, congestion response, latency management, and cost efficiency — Uber was able to achieve the networking optimization its modern tech stack requires.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/aai_concept_value_prop_with_without_pictur.max-1000x1000.jpg"
        
          alt="aai concept value prop with_without picture"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Starting with a private preview, Uber deployed this feature across its infrastructure, beginning with Google Cloud Interconnect deployments in Phoenix, Arizona, and Ashburn, Virginia. Application awareness on Cloud Interconnect allows Uber to classify and prioritize end-user application traffic over less time-sensitive data using DSCP marking and configured queuing profiles.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;In the following chart, we look at the four key features of application awareness on Cloud Interconnect, how they differ from legacy approaches, and how they help provide better operational continuity for organizations like Uber. &lt;br/&gt;&lt;br/&gt;&lt;/span&gt;&lt;/p&gt;
&lt;div align="left"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;
&lt;div style="color: #5f6368; overflow-x: auto; overflow-y: hidden; width: 100%;"&gt;&lt;table&gt;&lt;colgroup&gt;&lt;col/&gt;&lt;col/&gt;&lt;col/&gt;&lt;/colgroup&gt;
&lt;thead&gt;
&lt;tr&gt;
&lt;th scope="col" style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Feature&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Standard interconnect solutions&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;th scope="col" style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Application awareness on Cloud Interconnect&lt;/strong&gt;&lt;/p&gt;
&lt;/th&gt;
&lt;/tr&gt;
&lt;/thead&gt;
&lt;tbody&gt;
&lt;tr&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Traffic handling&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;All traffic treated equally (first-in, first-out)&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Traffic classified into six distinct traffic classes&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Congestion response&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;High-priority application traffic may be dropped during bursts&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Business-critical traffic is protected via strict priority or bandwidth sharing policies&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Latency management&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Unpredictable latency for high priority applications&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Predictable and consistent low-latency for time-sensitive workloads&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;tr&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Cost efficiency&lt;/strong&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Requires expensive overprovisioning to absorb peaks&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;td style="vertical-align: middle; border: 1px solid #000000; padding: 16px;"&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Efficient bandwidth utilization and lower TCO&lt;/span&gt;&lt;/p&gt;
&lt;/td&gt;
&lt;/tr&gt;
&lt;/tbody&gt;
&lt;/table&gt;&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;/div&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Uber's key takeaways&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;For Uber, the business value of being able to prioritize business-critical traffic on its networks by deploying application awareness on Cloud Interconnect was immediate. And in doing so, Uber has also created a blueprint that other enterprises with similar hybrid cloud challenges can replicate. The core elements of that blueprint include:&lt;/span&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Ensuring business continuity&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Uber can decide in real time which application traffic to prioritize during major, high-traffic events. This means that mission critical applications stay up and running during even extreme events (both planned and unplanned). Uber leadership has called application awareness on Cloud Interconnect important for its global operations. &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Efficient bandwidth utilization&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: Instead of blindly overprovisioning bandwidth to prevent congestion, application awareness allows Uber to better utilize their existing Cloud Interconnect capacity aligned with their expected network bandwidth needs. The result is lower total cost of ownership for network infrastructure.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Unblocked workload migration&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;: By protecting critical applications from network congestion, Uber was able to migrate significant workloads to Google Cloud and, in the process, dramatically reduce operational overhead.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;"Application awareness on Cloud Interconnect was the key that unlocked our ability to migrate more strategic workloads to Google Cloud and is critical for maintaining service reliability during peak global demand. By allowing us to intelligently prioritize traffic, it helps us ensure that we can protect our higher priority services and make our infrastructure more efficient, lowering our total cost of ownership. This wasn't just a feature deployment; it was a deep engineering partnership that delivered a solution critical to our business." &lt;/span&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;– &lt;/span&gt;&lt;strong style="font-style: italic; vertical-align: baseline;"&gt;Harry Liu&lt;/strong&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;, Director of Engineering, Uber&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Securing network reliability for AI and beyond&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;As more enterprises integrate cloud-based AI models, distributed applications, and data analytics, it's becoming a business imperative to be ready to handle the massive data transfers that follow. But in doing so, they also have to ensure they never compromise the reliability of their critical applications. &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;With application awareness on Cloud Interconnect, Uber demonstrated that moving beyond simple bandwidth overprovisioning to protect business-critical traffic was an essential step to building the stability required to embrace modern hybrid and multicloud strategies.&lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;You can read our blog about &lt;/span&gt;&lt;a href="https://cloud.google.com/blog/products/networking/cross-cloud-network-enhancements-for-distributed-workloads/"&gt;&lt;span style="font-style: italic; text-decoration: underline; vertical-align: baseline;"&gt;the potential of Cloud Interconnect across industries&lt;/span&gt;&lt;/a&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt; to learn more about what the service can bring to your organization, and if you’re ready to explore more, our team of networking and industry &lt;/span&gt;&lt;a href="https://cloud.google.com/contact/form?e=48754805"&gt;&lt;span style="font-style: italic; text-decoration: underline; vertical-align: baseline;"&gt;experts are ready to help&lt;/span&gt;&lt;/a&gt;&lt;span style="font-style: italic; vertical-align: baseline;"&gt;.&lt;/span&gt;&lt;/p&gt;&lt;/div&gt;
&lt;div class="block-related_article_tout"&gt;





&lt;div class="uni-related-article-tout h-c-page"&gt;
  &lt;section class="h-c-grid"&gt;
    &lt;a href="https://cloud.google.com/blog/topics/telecommunications/vodafone-gen-ai-enhances-network-lifecycle/"
       data-analytics='{
                       "event": "page interaction",
                       "category": "article lead",
                       "action": "related article - inline",
                       "label": "article: {slug}"
                     }'
       class="uni-related-article-tout__wrapper h-c-grid__col h-c-grid__col--8 h-c-grid__col-m--6 h-c-grid__col-l--6
        h-c-grid__col--offset-2 h-c-grid__col-m--offset-3 h-c-grid__col-l--offset-3 uni-click-tracker"&gt;
      &lt;div class="uni-related-article-tout__inner-wrapper"&gt;
        &lt;p class="uni-related-article-tout__eyebrow h-c-eyebrow"&gt;Related Article&lt;/p&gt;

        &lt;div class="uni-related-article-tout__content-wrapper"&gt;
          &lt;div class="uni-related-article-tout__image-wrapper"&gt;
            &lt;div class="uni-related-article-tout__image" style="background-image: url('')"&gt;&lt;/div&gt;
          &lt;/div&gt;
          &lt;div class="uni-related-article-tout__content"&gt;
            &lt;h4 class="uni-related-article-tout__header h-has-bottom-margin"&gt;How Vodafone is using gen AI to enhance network life cycle&lt;/h4&gt;
            &lt;p class="uni-related-article-tout__body"&gt;Vodafone and Google Cloud deployed generative AI to unlock new levels of efficiency, creativity, and customer satisfaction through networ...&lt;/p&gt;
            &lt;div class="cta module-cta h-c-copy  uni-related-article-tout__cta muted"&gt;
              &lt;span class="nowrap"&gt;Read Article
                &lt;svg class="icon h-c-icon" role="presentation"&gt;
                  &lt;use xmlns:xlink="http://www.w3.org/1999/xlink" xlink:href="#mi-arrow-forward"&gt;&lt;/use&gt;
                &lt;/svg&gt;
              &lt;/span&gt;
            &lt;/div&gt;
          &lt;/div&gt;
        &lt;/div&gt;
      &lt;/div&gt;
    &lt;/a&gt;
  &lt;/section&gt;
&lt;/div&gt;

&lt;/div&gt;</description><pubDate>Wed, 26 Aug 2026 16:00:00 +0000</pubDate><guid>https://cloud.google.com/blog/products/networking/uber-de-risks-hybrid-ai-with-cloud-interconnect/</guid><category>Customers</category><category>Cloud Migration</category><category>Developers &amp; Practitioners</category><category>Hybrid &amp; Multicloud</category><category>Networking</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/image1_fsLq9RR.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>How Uber improves network reliability while unblocking cloud migration</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/image1_fsLq9RR.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/products/networking/uber-de-risks-hybrid-ai-with-cloud-interconnect/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Jean He</name><title>Distinguished Engineer, Uber</title><department></department><company></company></author><author xmlns:author="http://www.w3.org/2005/Atom"><name>Gopinath Balakrishnan</name><title>Principal Architect, Google Cloud</title><department></department><company></company></author></item><item><title>Your chance to start building AI agents from the absolute basics</title><link>https://cloud.google.com/blog/topics/developers-practitioners/your-chance-to-start-building-ai-agents-from-the-absolute-basics/</link><description>&lt;div class="block-paragraph_advanced"&gt;&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;Have you been hearing a lot about "AI agents" lately but aren't sure how to actually start building them? You don't need a background in machine learning or years of software experience to get started. The best way to learn is by doing, which is why we built &lt;/span&gt;&lt;strong style="vertical-align: baseline;"&gt;Agent Valley&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt;.   &lt;/span&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong style="vertical-align: baseline;"&gt;Agent Valley&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; is a free, 5-week live learning series designed to take you from scratch to building your very own hands-on agent systems. And instead of staring at boring terminal lines, you’ll be building and playing inside a tiny, low-poly virtual world!&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Meet your instructor&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;You’ll be learning directly from Annie Wang, one of our top Google DevRel Engineers. She designed this course from the ground up to be fully hands-on, interactive, and beginner-friendly. If you want to learn how AI systems are built by the people actually designing them at Google, this is your chance.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;How we'll learn together&lt;/span&gt;&lt;/h2&gt;
&lt;p&gt;&lt;span style="vertical-align: baseline;"&gt;You’ll learn by building in a split-screen workspace on your laptop. On Day 1, you'll describe and summon a custom low-poly companion that serves as your play character and save file. As you guide your companion through the valley's five districts, a live Runtime Inspector sits right beside the game, showing you exactly what the AI is thinking, deciding, and costing in real-time. Setup is completely zero-stress. Google will provide the environment for running these exercises, so you can dive straight into building.&lt;/span&gt;&lt;/p&gt;
&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Agent 101 Live with 5 modular sessions (Jump in anytime!) &lt;/span&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Week 1: The Summoning Grove (CONTROL)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; · Get started by summoning your companion and learning how to keep its memory and traits consistent across a conversation.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Week 2: The Buildyard (DECOMPOSE)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; · Learn how to break a big project down so multiple AI assistants can work together in parallel without stepping on each other's toes.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Week 3: Market Street (COORDINATE)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; · Open up a virtual shop! You'll learn how to write reliable code so transactions and returns go smoothly without crashing.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Week 4: The Archive (REMEMBER)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; · Give your companion a memory. Learn how to help your agent remember past details without getting confused or making things up.&lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;li aria-level="1" style="list-style-type: disc; vertical-align: baseline;"&gt;
&lt;p role="presentation"&gt;&lt;strong style="vertical-align: baseline;"&gt;Week 5: The Night Market (LIVE)&lt;/strong&gt;&lt;span style="vertical-align: baseline;"&gt; · The grand finale. Learn how to make your agent react live to events in the world (like fireworks or stage lights) while keeping the system fast and affordable.              &lt;/span&gt;&lt;/p&gt;
&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;
&lt;div class="block-image_full_width"&gt;






  
    &lt;div class="article-module h-c-page"&gt;
      &lt;div class="h-c-grid"&gt;
  

    &lt;figure class="article-image--large
      
      
        h-c-grid__col
        h-c-grid__col--6 h-c-grid__col--offset-3
        
        
      "
      &gt;

      
      
        
        &lt;img
            src="https://storage.googleapis.com/gweb-cloudblog-publish/images/agent-valley-roadmap-2160x2700.max-1000x1000.png"
        
          alt="agent-valley-roadmap-2160x2700"&gt;
        
        &lt;/a&gt;
      
    &lt;/figure&gt;

  
      &lt;/div&gt;
    &lt;/div&gt;
  




&lt;/div&gt;
&lt;div class="block-paragraph_advanced"&gt;&lt;h2&gt;&lt;span style="vertical-align: baseline;"&gt;Join the livestream    &lt;/span&gt;&lt;/h2&gt;
&lt;ul&gt;
&lt;li role="presentation"&gt;&lt;span style="vertical-align: baseline;"&gt;5 Tue starting Sep 1 · 10:00 AM (Pacific Time)&lt;/span&gt;&lt;/li&gt;
&lt;li role="presentation"&gt;Anyone new to AI agents who wants to learn by coding and playing.                                                       &lt;/li&gt;
&lt;li role="presentation"&gt;RSVP Here: &lt;a href="https://goo.gle/agent101" rel="noopener" style="font-family: -apple-system, BlinkMacSystemFont, 'Segoe UI', Roboto, Oxygen, Ubuntu, Cantarell, 'Open Sans', 'Helvetica Neue', sans-serif;" target="_blank"&gt;&lt;span style="vertical-align: baseline;"&gt;goo.gle/agent101&lt;/span&gt;&lt;/a&gt;&lt;/li&gt;
&lt;/ul&gt;&lt;/div&gt;</description><pubDate>Wed, 26 Aug 2026 09:07:00 +0000</pubDate><guid>https://cloud.google.com/blog/topics/developers-practitioners/your-chance-to-start-building-ai-agents-from-the-absolute-basics/</guid><category>Developers &amp; Practitioners</category><media:content height="540" url="https://storage.googleapis.com/gweb-cloudblog-publish/images/Agent_Valley_Hero_Blog.max-600x600.png" width="540"></media:content><og xmlns:og="http://ogp.me/ns#"><type>article</type><title>Your chance to start building AI agents from the absolute basics</title><description></description><image>https://storage.googleapis.com/gweb-cloudblog-publish/images/Agent_Valley_Hero_Blog.max-600x600.png</image><site_name>Google</site_name><url>https://cloud.google.com/blog/topics/developers-practitioners/your-chance-to-start-building-ai-agents-from-the-absolute-basics/</url></og><author xmlns:author="http://www.w3.org/2005/Atom"><name>Christina Lin</name><title>Developer Relations Engineering Manager</title><department></department><company></company></author></item></channel></rss>