On the eve of Alphabet’s Q2 earnings, Google DeepMind pushed out three new Gemini models, but the one many developers were waiting for—Gemini 3.5 Pro—didn’t make the cut. Instead, the company doubled down on cost efficiency and niche security with Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and a limited-access version called Flash Cyber. The timing, just as OpenAI and Anthropic turn up the heat with their own cybersecurity offerings, feels less like a celebration and more like a holding action. The elephant in the room is Gemini 3.5 Pro. Originally slated for a June 2026 public launch, the model is still stuck in partner testing, with no firm date in sight. Internal frustration is boiling over, according to reports, with some researchers leaving for Anthropic and other labs. “We’re overthinking shipping,” one engineer told colleagues, pointing to endless cross-team coordination and internal turf wars between Cloud, DeepMind, and Android. Meanwhile, Google has quietly begun pretraining Gemini 4, a signal that the company is already hedging its bets on the next generation while the current flagship remains in limbo.
Token Savings, Same Brains: What’s New in 3.6 Flash
Gemini 3.6 Flash is the workhorse update, and it’s all about doing more with less. Google claims the model uses 17% fewer output tokens on average compared to its predecessor, and in complex engineering tasks like Datacurve’s DeepSWE benchmark, the savings can hit 65%. That translates to faster, cheaper API calls—average task time dropped from 2.7 minutes to 1.3 minutes, and per-task cost fell from $0.59 to $0.50, per Artificial Analysis. The new pricing is $1.50 per million input tokens and $7.50 per million output tokens, a slight cut from 3.5 Flash’s $9.00. But here’s the catch: those savings come without a meaningful intelligence boost. The Artificial Analysis Intelligence Index score remains stuck at 50, the same as 3.5 Flash. On specific tasks, the picture is mixed—coding benchmarks like MLE-Bench jumped 14 percentage points, but HLE dipped slightly. “If you look at intelligence vs time per task, this is a very fast model,” one Hacker News user argued. Overall it feels much more natural, and token output speed is fast.” Yet the lack of an IQ upgrade stings. As one commenter put it, “Saved tokens, but lost IQ.” Early enterprise feedback is cautiously positive. Harvey AI, the legal tech company, says document review and drafting tasks are about 12% faster. Figma’s engineering team reports faster design iteration loops with no quality degradation. GitHub Copilot is rolling out 3.6 Flash across all tiers, targeting web and app development workloads. But developers on X (formerly Twitter) shared early tests showing the model struggling with frontend UI generation, leading some to label it the “worst ever” for spatial reasoning.
Flash-Lite Gets Faster, but the Price Sting Is Real
Gemini 3.5 Flash-Lite is Google’s bet on raw speed and high throughput. It spits out up to 350 tokens per second, nearly double its predecessor, and slashes time per task from 1.0 to 0.6 minutes. Intelligence ticked up from 25 to 36 on the Artificial Analysis index, with big leaps in agentic benchmarks like Terminal-Bench. The price, however, is a sore spot. While the per-token cost looks cheap—$0.30 per million input and $2.50 per million output—it’s actually more expensive per task than the older 3.1 Flash-Lite. Average cost per task rose from $0.04 to $0.09, chiefly because the base rates nearly doubled. A developer on Google’s own forum complained that 3.5 Flash runs were costing 16 times more than Flash-Lite used to. “Actively penalizing developers who write good, efficient prompts” was the title of one thread. Despite the grumbling, Google is positioning Flash-Lite as the engine behind agentic search and document processing, gradually feeding it into Google Search to handle ever more complex queries.
The Cyber Bet: A Direct Shot at Anthropic
The most strategically charged release is Gemini 3.5 Flash Cyber, a fine-tuned model designed to find, verify, and patch software vulnerabilities. It’s baked into Google’s CodeMender agent framework and explicitly marketed as a cost-efficient alternative to Anthropic’s Claude Mythos—a model that has been racking up high-profile enterprise wins through Project Glasswing, including ICE/NYSE, Samsung, and Trend Micro. Google’s benchmarks show Flash Cyber significantly outperforming both baseline 3.5 Flash and Claude Opus 4.6 on vulnerability discovery in the V8 JavaScript engine (55 unique confirmed issues vs. 36 for Opus). In one stress test, it even produced a 100% reliable remote-code execution exploit that bypassed ASLR and W^X protections. Given the dual-use risk, access is heavily gated. Only governments and trusted partners can use it through CodeMinder, and the model won’t be generally available anytime soon. “This will give frontline defenders a head start in finding and fixing critical vulnerabilities before they can be exploited, while mitigating against broader misuse,” said Google DeepMind’s Raluca Ada Popa and Four Flynn. The controlled rollout reignited an old debate on Hacker News about who should hold automated exploit-finding tools, with some arguing it’s a necessary shield and others warning it’s a ticking time bomb. The cybersecurity AI race is heating up. OpenAI already has GPT-5.5-Cyber and the Daybreak partner program with IBM and Check Point, while Anthropic’s Mythos has discovered zero-days in major operating systems. Google’s move feels less like leadership and more like an attempt to not get left behind. As CNBC noted, Flash Cyber is Google’s “clearest answer yet to Anthropic’s lead in cybersecurity.” What does this all mean for Google’s bottom line? Alphabet reports Q2 earnings after the bell July 22, with analysts expecting cloud revenue around $225 billion—a 65% jump year over year, driven heavily by AI workloads. But the company’s capital spending is ballooning toward $190 billion this year, and investors are starting to ask whether the returns justify the outlay. A cheaper, faster Flash lineup might help win over cost-conscious enterprises, but the intelligence plateau and the Pro delay could push developers toward OpenAI’s cheaper Luna or Anthropic’s increasingly dominant API business. On Reddit, the mood is a mix of shrugs and irritation. “If you’ve been using Gemini since last year, it’s like watching your brother slowly develop Alzheimer’s,” one user joked. Others see the practicality: “It’s very good at frontend (much better than GPT-5.5) and it’s fast, so it’s a great tool for iteration.” Perhaps that’s the real story: in a world where every token and millisecond matters, Google is betting that efficiency beats brilliance for most people, most of the time. Whether that bet pays off may depend on how long developers are willing to wait for Pro.