Skip to the evolution relay
EvoCloudView timeline

Google Cloud / 2008 to 2025

The work
changes hands.

Six milestones changed what people operated directly and what the platform could carry for them.

Begin with 2008
Google Cloud capability relay through SiliconA chronological flow chart showing how managed work changes shape. App Engine runtime instances converge on BigQuery. Application workloads running as Kubernetes Pods can exchange data with BigQuery while the Kubernetes control plane schedules and reconciles those Pods. Kubernetes fans work into four Vertex AI lifecycle nodes, and Vertex AI feeds four Gemini inputs. From the Gemini model node, four workload paths fan out to four Ironwood TPU processors. Products remain separate services.2008MANAGED RUNTIMEApp Engine2011MANAGED DATABigQuery2014ORCHESTRATIONKubernetes2021ML LIFECYCLEVertex AI2023MULTIMODAL MODELGemini2025INFERENCE SILICONIronwood

April 2008

01 / 06

Runtime

App Engine

Run code without managing servers

Google Cloud capability relay through RuntimeA chronological flow chart showing how managed work changes shape. App Engine runtime instances converge on BigQuery. Application workloads running as Kubernetes Pods can exchange data with BigQuery while the Kubernetes control plane schedules and reconciles those Pods. Kubernetes fans work into four Vertex AI lifecycle nodes, and Vertex AI feeds four Gemini inputs. From the Gemini model node, four workload paths fan out to four Ironwood TPU processors. Products remain separate services.2008MANAGED RUNTIMEApp Engine2011MANAGED DATABigQuery2014ORCHESTRATIONKubernetes2021ML LIFECYCLEVertex AI2023MULTIMODAL MODELGemini2025INFERENCE SILICONIronwood
Capability map after 2008
Before
Provision, scale, and operate servers for a web application.
After
Deploy application code to a managed runtime.

Developers focused on the application while Google managed the runtime.

Read source

General availability, 2011

02 / 06

Data

BigQuery

Query data without managing infrastructure

Google Cloud capability relay through DataA chronological flow chart showing how managed work changes shape. App Engine runtime instances converge on BigQuery. Application workloads running as Kubernetes Pods can exchange data with BigQuery while the Kubernetes control plane schedules and reconciles those Pods. Kubernetes fans work into four Vertex AI lifecycle nodes, and Vertex AI feeds four Gemini inputs. From the Gemini model node, four workload paths fan out to four Ironwood TPU processors. Products remain separate services.2008MANAGED RUNTIMEApp Engine2011MANAGED DATABigQuery2014ORCHESTRATIONKubernetes2021ML LIFECYCLEVertex AI2023MULTIMODAL MODELGemini2025INFERENCE SILICONIronwood
Capability map after 2011
Before
Build and maintain infrastructure for large analytical queries.
After
Submit queries to a managed data warehouse.

Teams wrote queries instead of maintaining analytics clusters.

Read source

First commit, June 2014

03 / 06

Orchestration

Kubernetes

Keep application pods near their declared state

Google Cloud capability relay through OrchestrationA chronological flow chart showing how managed work changes shape. App Engine runtime instances converge on BigQuery. Application workloads running as Kubernetes Pods can exchange data with BigQuery while the Kubernetes control plane schedules and reconciles those Pods. Kubernetes fans work into four Vertex AI lifecycle nodes, and Vertex AI feeds four Gemini inputs. From the Gemini model node, four workload paths fan out to four Ironwood TPU processors. Products remain separate services.2008MANAGED RUNTIMEApp Engine2011MANAGED DATABigQuery2014ORCHESTRATIONKubernetes2021ML LIFECYCLEVertex AI2023MULTIMODAL MODELGemini2025INFERENCE SILICONIronwood
Capability map after 2014
Before
Manually coordinate container placement, recovery, and scaling.
After
Declare the state that Kubernetes controllers should maintain.

Controllers kept applications close to their declared state.

Read source

General availability, May 2021

04 / 06

ML lifecycle

Vertex AI

Manage training, deployment and monitoring

Google Cloud capability relay through ML lifecycleA chronological flow chart showing how managed work changes shape. App Engine runtime instances converge on BigQuery. Application workloads running as Kubernetes Pods can exchange data with BigQuery while the Kubernetes control plane schedules and reconciles those Pods. Kubernetes fans work into four Vertex AI lifecycle nodes, and Vertex AI feeds four Gemini inputs. From the Gemini model node, four workload paths fan out to four Ironwood TPU processors. Products remain separate services.2008MANAGED RUNTIMEApp Engine2011MANAGED DATABigQuery2014ORCHESTRATIONKubernetes2021ML LIFECYCLEVertex AI2023MULTIMODAL MODELGemini2025INFERENCE SILICONIronwood
Capability map after 2021
Before
Join separate tools for building, training, deploying, and monitoring models.
After
Manage more of the model lifecycle through one platform.

Training, deployment, and monitoring moved into one managed workflow.

Read source

Introduced, December 2023

05 / 06

Models

Gemini

Access Gemini through Vertex AI

Google Cloud capability relay through ModelsA chronological flow chart showing how managed work changes shape. App Engine runtime instances converge on BigQuery. Application workloads running as Kubernetes Pods can exchange data with BigQuery while the Kubernetes control plane schedules and reconciles those Pods. Kubernetes fans work into four Vertex AI lifecycle nodes, and Vertex AI feeds four Gemini inputs. From the Gemini model node, four workload paths fan out to four Ironwood TPU processors. Products remain separate services.2008MANAGED RUNTIMEApp Engine2011MANAGED DATABigQuery2014ORCHESTRATIONKubernetes2021ML LIFECYCLEVertex AI2023MULTIMODAL MODELGemini2025INFERENCE SILICONIronwood
Capability map after 2023
Before
Use separate processing paths for text, code, images, audio, and video.
After
Use one natively multimodal model family across those inputs.

The model handled more of the work between input types.

Read source

Announced April, generally available November 2025

06 / 06

Silicon

Ironwood

Co-design compute, memory and networking for AI scale

Google Cloud capability relay through SiliconA chronological flow chart showing how managed work changes shape. App Engine runtime instances converge on BigQuery. Application workloads running as Kubernetes Pods can exchange data with BigQuery while the Kubernetes control plane schedules and reconciles those Pods. Kubernetes fans work into four Vertex AI lifecycle nodes, and Vertex AI feeds four Gemini inputs. From the Gemini model node, four workload paths fan out to four Ironwood TPU processors. Products remain separate services.2008MANAGED RUNTIMEApp Engine2011MANAGED DATABigQuery2014ORCHESTRATIONKubernetes2021ML LIFECYCLEVertex AI2023MULTIMODAL MODELGemini2025INFERENCE SILICONIronwood
Capability map after 2025
GPU cluster
Scale general-purpose parallel processors across an accelerator and network stack.
Ironwood pod
Co-design TPU matrix compute, HBM, ICI networking and XLA as one scale-up system.

Lower serving cost. Faster inference. Built to scale.

Read source