Google embeds Gemini into custom silicon for 2028 inference efficiency

Google is engineering custom silicon that embeds Gemini's model architecture directly into hardware, targeting a 6 to 10 fold efficiency leap over current TPUs. The Frozen v2 chip, slated for 2028 deployment, represents a strategic shift toward vertical integration of model and silicon design, potentially reshaping inference economics across the industry. If realized, the cost advantage could reshape competitive positioning between Google, OpenAI, and Anthropic in both cloud and on-device inference markets, signaling a broader trend of AI leaders moving beyond general-purpose accelerators.
Modelwire context
Analyst takeThe detail worth sitting with is the 2028 deployment timeline. That's a long runway, and it means any inference cost advantage is priced into a competitive landscape that will look substantially different from today's, with rivals having two-plus years to respond through their own hardware partnerships or architectural changes.
Modelwire has no prior coverage to anchor this to directly, so the honest framing is that this belongs to a broader thread the site hasn't yet built out: the race among frontier labs to own their compute stack rather than rent it. Google's move mirrors the logic that drove Apple's M-series transition, where architectural co-design between model and chip compounds over time in ways that commodity hardware cannot match. The competitive pressure on OpenAI and Anthropic is real but asymmetric: OpenAI has a deep Microsoft Azure dependency, while Anthropic leans on AWS and Google itself, which makes a Google-owned efficiency advantage structurally awkward for both.
Watch whether Anthropic renegotiates or diversifies its cloud infrastructure agreements in the next 12 to 18 months. A move toward dedicated silicon partnerships would signal that internal modeling teams view this threat as credible well before 2028.
This analysis is generated by Modelwire’s editorial layer from our archive and the summary above. It is not a substitute for the original reporting. How we write it.
MentionsGoogle · Gemini · Frozen v2 · OpenAI · Anthropic · TPU
Modelwire Editorial
This synthesis and analysis was prepared by the Modelwire editorial team. We use advanced language models to read, ground, and connect the day’s most significant AI developments, providing original strategic context that helps practitioners and leaders stay ahead of the frontier.
Modelwire summarizes, we don’t republish. The Decoder originally reported this story as “Google's "Frozen v2" chip reportedly bakes Gemini's architecture directly into silicon for efficiency gains”. The full content lives on the-decoder.com. If you’re a publisher and want a different summarization policy for your work, see our takedown page.