中文
← Back to news
IndustryJul 21, 2026

Google Reportedly Developing Frozen v2 Chip: Hardcoding Gemini Architecture into Silicon, Energy Efficiency Could Surpass TPU by Tenfold

According to an exclusive report by The Information, Google is developing a server AI chip codenamed "Frozen v2" that aims to permanently etch parts of the underlying architecture of the Gemini model into silicon, rather than serving as a general-purpose computing platform like traditional chips. The project is led by Jeff Dean, Chief Scientist at Google DeepMind, and is an iteration of the earlier Frozen project.

Core Design: From "Running Models" to "Growing Models"

The core idea of Frozen v2 is to "harden the architecture while preserving weights." Unlike general-purpose chips such as NVIDIA GPUs or Google TPUs, Frozen v2 is custom-built solely for the Gemini model, embedding critical computation paths directly into circuits, eliminating redundant steps like real-time scheduling and data movement. Google employees estimate that, measured by tokens processed per watt, its energy efficiency could be 6 to 10 times higher than the latest generation of TPUs.

  • Earlier Frozen project: Attempted to burn model weights directly into the chip, but was shelved because the hardened weights could not adapt to model updates.
  • Frozen v2 compromise: Only the underlying computational blueprint is hardened; weights can still be updated. As long as the Gemini architecture does not undergo major changes, the chip can be used continuously.
  • Engineering trade-off: Google is still debating the degree of hardening—more locking yields higher efficiency but less flexibility.

Compute Crisis: Why Google Bets on Specialized Chips

Google faces a severe compute shortage, which has already forced Google Cloud to reject orders from external customers. In June 2024, Google signed a contract with SpaceX, paying $920 million per month to lease 110,000 NVIDIA GPUs, with the contract lasting until 2029.

  • Industry trend: The entire industry is betting on inference chips, including startups like SambaNova and d-Matrix, as well as giants like OpenAI, Microsoft (Maia), and Amazon (Trainium/Inferentia). NVIDIA also acquired Groq's technology license for $20 billion in December 2023.
  • Extreme route: Canadian startup Taalas has already pursued the "model hardwiring" approach, raising over $200 million.

Risks and Bets: Architecture Convergence or Rigidity?

The cost of Frozen v2 is loss of flexibility: only if subsequent Gemini models use the same underlying architecture can the chip continue to work. Google is betting that model architectures will not undergo fundamental changes in the coming years.

  • Model progress: Gemini 3.5 Pro (codenamed Cappuccino) has been delayed three times due to subpar coding capabilities, originally scheduled for June 2024 but now postponed by several months.
  • Industry context: A Wharton professor pointed out the "disappointment trap of next-generation giant models," where the returns from piling up data and compute are diminishing, and the pace of architectural iteration is slowing.
  • Deployment timeline: The chip is expected to be deployed as early as 2028, with production volume not reaching TPU scale, more like a limited experiment.

Implications and Insights

Frozen v2 marks the evolution of AI chips from general-purpose to extreme specialization, similar to the shift from CPU to ASIC in Bitcoin mining. However, this could hinder the development of new architectures (e.g., non-Transformer models) due to the lack of compatible hardware. Google's bet reflects the industry's expectation of architectural convergence, but if a paradigm shift occurs, Frozen v2's high efficiency could instantly become a technical liability.

Also available in 中文.