Back to feed
News Story
钛媒体AGI
4 sources

Alphabet Embeds Gemini Directly into Silicon

Alphabet has integrated its Gemini AI model directly into its custom TPU chips, achieving 6 to 10 times improvement in inference efficiency over existing TPUs. This hardware innovation aims to accelerate AI computing performance and strengthen Alphabet's leadership in AI infrastructure.

SynthePulse Insight · AI deep reading

Google Frozen v2: Embedding Gemini Architecture into Silicon, AI Inference Efficiency Could Improve 6-10x

Version 1 · 1 source

Google is developing a new server chip codenamed Frozen v2 that embeds the Gemini model architecture directly into hardware, rather than just freezing weights. It is said to achieve 6 to 10 times the AI inference efficiency of current TPUs, with deployment planned for 2028, but in small volumes primarily to address internal compute bottlenecks.

  • Frozen v2 hard-codes the Gemini model architecture (not weights) into the chip, allowing new weights to be loaded, but the fixed architecture limits it to that model family.
  • According to The Information citing sources, the chip's AI inference efficiency is 6 to 10 times higher than current TPUs, with deployment planned for 2028.
  • Unlike TPUs that can adapt to various models, Frozen v2 is designed exclusively for Gemini architecture, not for external sale, mainly to ease Google's internal AI compute pressure.
  • If the efficiency gains materialize, Google could significantly reduce inference costs, gaining a price advantage over OpenAI and Anthropic.
  • The initial concept (Frozen v1) was proposed by Jeff Dean, planning to freeze weights, but was abandoned because it only worked with a single Gemini version and became obsolete quickly.
  • The degree of architecture hardening in Frozen v2 is not yet finalized, and its production volume is smaller than the TPU series, making it an experiment in specialized chips.
Open section navigationFrom Freezing Weights to Freezing Architecture: The Design Evolution of Frozen v2

From Freezing Weights to Freezing Architecture: The Design Evolution of Frozen v2

According to a report by The Information, Google is developing a server chip with the internal codename Frozen v2, whose core innovation is embedding the architecture of the Gemini AI model directly into silicon. Unlike Google's existing TPUs—which can adapt to various models—Frozen v2 permanently hard-codes part of the model's structure into the chip, thereby reducing computational steps and speeding up response times.

This idea originated from Google DeepMind's chief scientist Jeff Dean. The initial Frozen design (Frozen v1) planned to embed model weights (the specific settings that determine how an AI model responds to queries) directly into the chip. However, that approach was abandoned because the chip could only work with a single Gemini version and became obsolete too quickly. Frozen v2 instead hard-codes the model architecture (the underlying blueprint), not the weights, so new weights can still be loaded. However, according to The Information, exactly how much of the architecture will be hard-coded has not yet been finalized.

Efficiency Promise and Deployment Timeline

According to The Information citing sources, Frozen v2 can achieve 6 to 10 times higher AI inference efficiency than Google's current TPU chips. Google plans to start deploying the chip in 2028. Since Frozen v2 is only compatible with the Gemini architecture and its production volume is smaller than the TPU series, Google views it as an experiment in specialized chips rather than a product for external customers.

Google currently offers TPUs to external cloud customers through programs like TPU@Premises and has set an internal goal of capturing 10% of Nvidia's annual revenue. In contrast, Frozen v2 is primarily aimed at alleviating Google's internal AI compute bottlenecks.

Competitive Impact: Inference Cost Advantage Could Reshape the Market

In the AI business, the degree of inference cost optimization increasingly determines profit margins. If Frozen v2 delivers on its efficiency promise, Google could run powerful models at lower costs, thereby capturing market share from OpenAI and Anthropic. The chip has the potential to significantly reduce Google's AI inference costs, giving the company a price advantage.

However, the degree of architecture hardening in Frozen v2 is not yet determined, and its success depends on Google's long-term commitment to the Gemini architecture. If future model architectures undergo major changes, the chip could quickly become obsolete.

Credibility boundary

The core information in this article originates from a report by The Information, which was republished by THE DECODER. All details regarding efficiency improvements, deployment timelines, and design specifics come from anonymous sources and are at the source_claim level, not yet confirmed by Google.

Insight takeaway

Google's Frozen v2, by embedding the Gemini architecture into a chip, promises a 6-10x improvement in AI inference efficiency, but the chip is low-volume, not for external sale, and the degree of architecture hardening is undetermined. If successful, it would significantly lower Google's inference costs, enhancing its competitiveness against OpenAI and Anthropic.