← All posts
Opinion·5 min read

Gemini 3.5 Pro Missed Three Deadlines Over Coding. That Beats a Benchmark Chart.

Google has delayed its next flagship three times, most recently citing hallucinations and coding regressions. A delay disclosure that names a specific failure mode is worth more than most launch-day benchmark charts, and it previews exactly what we'll run once the model reaches the arena.


Google's next flagship has now missed its ship date three times, and every version of the excuse points at the same weak spot: coding. Bloomberg reported on July 16 that the rebuilt Gemini 3.5 Pro fell short of Google's own internal quality bar on hallucination rates and real-world reliability, after a late-June training data update meant to fix coding performance produced what insiders called disappointing results. As of today it still hasn't shipped. That is a more useful piece of information than most benchmark charts we see on launch day, because it names the failure instead of hiding it behind a cherry-picked score.

Three deadlines, one recurring problem

Strip out the speculation and the pattern is plain. Each delay traces back to the same category of failure, and each fix attempt exposed a deeper one underneath it.

  • Delay one (June): early enterprise testing flagged weak token efficiency, coding performance, and multi-step reasoning; Google pushed the date to refine quality.
  • Delay two: Google concluded the original model had structural failures in recursive tool-calling and SVG generation that fine-tuning couldn't patch, and ordered a ground-up pretraining rebuild rather than a touch-up.
  • Delay three (reported July 16): the rebuilt model cleared its early bars but missed on hallucination rates and real-world reliability; a late-June training data update aimed at coding came back disappointing.

Google's public line hasn't changed shape: "We're shipping quickly across a wide range of models while keeping them highly cost-effective for customers." It's a fair statement about the portfolio (GPT-5.6's three-tier launch and Meta's Muse Spark 1.1 both landed on schedule this cycle), but it sits awkwardly next to a flagship that is now testing alongside an upgraded Flash model with partners instead of shipping to the public.

We're shipping quickly across a wide range of models while keeping them highly cost-effective for customers. — Google statement on the Gemini 3.5 Pro delay, reported July 16, 2026

The leak economy already wrote the spec sheet

None of that has stopped the internet from pricing and specifying a model nobody outside Google has run. Third-party reporting circulating since early July puts Gemini 3.5 Pro at a 2 million token context window, a Deep Think reasoning layer for multi-step problems, and API pricing somewhere around $15 per million input tokens and $60 per million output tokens, roughly ten times Flash's rate. Every one of those numbers traces back to unnamed sources or Vertex AI documentation leaks, not an official Google page.

That's the trap worth naming: by the time a delayed flagship actually ships, the narrative around it is already fixed. Readers walk in expecting the leaked context window and the leaked price, and then judge the real model against a rumor instead of against what it actually does on a prompt. We've made the same point about launch-day benchmarks before, most recently in our look at GPT-5.6, Grok 4.5, and Muse Spark's efficiency claims landing in the same week with no shared yardstick between them. A leak is just a benchmark with worse sourcing.

What a delay tells you that a chart doesn't

A shipped benchmark is optimized for the day it's published; a delay disclosure that names recursive tool-calling failures, SVG generation bugs, and coding regressions after a training update is closer to an admission against interest. It cost Google's stock a visible dip on the report. Companies don't volunteer that kind of specificity unless the alternative, silence followed by a weak public launch, is worse. Read plainly, Google just told everyone exactly which three things to stress-test the day 3.5 Pro actually lands: agent chains with nested tool calls, long structured outputs like SVG or diagrams, and whether the 2 million token context holds up on a real multi-file diff rather than a synthetic needle-in-haystack test.

What we'll run once it lands

The same one-shot coding and agentic prompts we run on every other model in the arena, no advance notice, no leaked spec sheet as a scoreboard. If recursive tool-calling and long-context coding are genuinely fixed, blind votes will show it. If they aren't, no context-window number will save it.

In the meantime, the Gemini flagship you can actually test today is Gemini 3.2 Pro, already in rotation and already voted on blind in our coding arena against the current GPT-5.6 tiers, Claude Fable 5, and Grok 4.5. That comparison is real because it's running now, not because a spec sheet says it should win.

European and Polish-speaking readers tracking the same story as it develops can follow it in Polish at nowosci.ai, our sister site covering this beat.

The lesson isn't that Gemini 3.5 Pro will be bad. It's that a lab willing to blow three deadlines over coding reliability is telling you, for free, exactly where to point the arena the day it shows up. We'll take that over a launch-day chart every time.

Don’t take the post’s word for it

The arena runs every model’s real output live. Pick a challenge, go blind, and cast a vote that counts in the public tally.

Open the arena