Mountain View: Google is reportedly developing a new server chip designed to make its Gemini artificial intelligence models significantly more efficient by embedding parts of the AI model directly into the hardware. The move is aimed at reducing computing costs, improving inference speed and easing the growing demand for AI infrastructure as the company competes with rivals including Microsoft, OpenAI and Anthropic.
According to a report, the project is internally known as “Frozen v2” and represents a major shift in how Google intends to run large language models in its data centres. Rather than relying solely on software running atop general-purpose AI accelerators, the new chip would hardwire elements of Gemini’s architecture into the silicon itself.
How the new chip will work
Large language models typically require enormous computing power to process user prompts and generate responses. This stage, known as inference, involves constant movement of data between memory and processors.
Google’s proposed chip aims to reduce that overhead by directly integrating the blueprints of Gemini into the processor.
Embedding portions of the model within the hardware is expected to reduce data movement, minimise the number of computational decisions required during inference and improve overall efficiency. The result could be faster AI responses while consuming less power and computing resources.
Tackling AI infrastructure constraints
The reported project also reflects Google’s efforts to address growing pressure on its AI infrastructure.
Demand for AI computing capacity has surged as businesses increasingly deploy generative AI applications. Reports suggest the shortage has, at times, forced Google Cloud to limit or decline certain external customer requests while prioritising its own AI services.
By developing more specialised hardware, Google hopes to improve the utilisation of its data centres while lowering the cost of serving billions of AI queries across products such as Gemini, Search and Workspace.
Strengthening Google’s AI ecosystem
Google has invested heavily in designing its own AI chips through its Tensor Processing Unit (TPU) programme for several years.
The reported “Frozen v2” project would further deepen the integration between Google’s hardware and software, allowing Gemini models to run more efficiently on infrastructure specifically optimised for them.
This approach mirrors a broader industry trend in which major technology companies are increasingly building custom silicon rather than relying entirely on third-party processors for AI workloads.
Competing in the AI chip race
The development comes amid intense competition across the artificial intelligence industry.
Companies are racing not only to build more capable AI models but also to reduce the enormous costs associated with training and serving them. More efficient chips could allow Google to process more AI requests while lowering operational expenses.
Industry analysts believe specialised hardware will become an increasingly important competitive advantage as AI adoption continues to accelerate worldwide.
Deployment could take time
While the project is progressing, reports indicate that Google’s engineers are still finalising the chip’s design, including determining how much of Gemini’s architecture should be permanently embedded into the hardware.
The chip is reportedly expected to enter deployment around 2028, although development timelines could change as the project evolves.
If successful, the new processor could represent one of Google’s most significant advances in AI infrastructure, enabling Gemini to deliver faster responses, lower operating costs and better scalability as demand for generative AI continues to grow.
