Google Faces Internal Doubts Over Gemini 4 as AI Race Intensifies

date
21:54 01/10/2026
avatar
GMT Eight
Google is preparing to launch Gemini 4 amid reported internal debate over whether its next flagship AI model can match rivals in real-world performance, particularly coding. While the model has performed strongly on industry benchmarks, some employees reportedly believe those results do not fully translate into practical tasks. The stakes are significant for Google, which is integrating Gemini across Search, Gmail, Chrome and other products while competing with rapidly advancing models and AI applications from OpenAI, Anthropic and other technology companies.

Gemini 4 has produced strong benchmark results, but some employees with access to the model have raised concerns about its performance on practical coding tasks, according to people familiar with its development. Front-end development, including the design and user experience of websites and applications, has reportedly been one area of uneven performance. Google disputes the characterization that Gemini 4 is underperforming in coding and says internal testing shows the model remains at the frontier.

The debate follows Google’s decision to abandon Gemini 3.5 Pro, which had originally been announced at its I/O conference in May and scheduled for release in June. The company instead shifted its attention toward Gemini 4. Developing frontier models requires substantial resources, with individual training runs potentially costing hundreds of millions of dollars.

Google now faces pressure to demonstrate that the next generation represents a meaningful improvement. Gemini technology already powers features across Search, Maps, Gmail and Chrome, giving the company access to products used by billions of people. Google also said its consumer Gemini app and AI Mode in Search have each surpassed 1 billion users, while its enterprise AI business continues to grow.

Competition, however, is expanding beyond the underlying models themselves. OpenAI and Anthropic are increasingly developing applications and autonomous agents, including tools designed specifically for software development. Some Google employees reportedly believe competing models from the two companies are improving faster than Gemini, while others argue Gemini 4 has closed the performance gap.

One concern centers on what the AI industry calls “benchmaxxing” — optimizing models to perform well on standardized evaluations without producing equivalent improvements in real-world usefulness. Strong benchmark scores have become important marketing tools for AI companies, but critics argue they may not accurately measure qualities such as usability, software design or performance on complex everyday tasks.

Gemini 4 reportedly has strengths beyond coding. People familiar with the model pointed to its ability to process information beyond text, including extracting metadata from video, as well as improvements in safety, cybersecurity and natural communication. However, the model is also reportedly very large, potentially making it more expensive to operate and adding pressure to the economics of deploying AI at massive scale.

The development challenges come amid broader changes within Google’s AI organization. Several prominent researchers have departed, while leadership at DeepMind has shifted, with Demis Hassabis moving into a chairman role and Koray Kavukcuoglu taking responsibility for day-to-day operations. Researchers have also reportedly raised concerns about bureaucracy and shifting priorities as Google attempts to integrate AI across a wide range of products.

Gemini 4 therefore represents more than another model release for Google. Its performance could influence the company’s ability to defend its position across search, software and enterprise technology as competitors increasingly build complete AI products around their own models. Google retains a major distribution advantage through its existing ecosystem, but converting that reach into leadership in the next phase of AI will depend heavily on how Gemini performs outside benchmark tests.