
Google Launches Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber — Flash-Lite Hits 350 tok/s at $2.50/M Output
TLDR
- Google dropped three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
- 3.6 Flash is the new workhorse — 17% fewer output tokens than 3.5 Flash, priced at $1.50 input / $7.50 output per million tokens
- 3.5 Flash-Lite runs at 350 output tokens/second and starts at just $0.30 input / $2.50 output per million tokens
- 3.5 Flash Cyber is locked behind a restricted pilot for governments and trusted partners via Google’s CodeMender agent
- All three are live today in the Gemini API, Google AI Studio, and the Gemini app; Gemini 3.5 Pro is still in partner testing
Google refreshed its mid-tier Gemini stack on 21 July 2026 with three new models aimed at developers shipping production AI agents. The headline release is Gemini 3.6 Flash, positioned as the default workhorse for coding, knowledge work, and multimodal tasks. According to Google’s benchmarks, 3.6 Flash uses 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index, and up to 65% fewer on certain DeepSWE coding workflows from Datacurve.

The price has dropped too. 3.6 Flash now costs $1.50 per million input tokens and $7.50 per million output tokens, undercutting 3.5 Flash on a per-token basis. On OSWorld-Verified, the computer-use benchmark, 3.6 Flash climbs from 78.4% to 83.0%; on MLE Bench it jumps from 49.7% to 63.9%. Customers Hebbia and Harvey have already integrated it for document parsing, chart analysis, and report drafting.
Flash-Lite Targets High-Throughput Agentic Workloads
Alongside 3.6 Flash, Google shipped 3.5 Flash-Lite as its cheapest and fastest 3.5-class option. Artificial Analysis clocks it at 350 output tokens per second, with pricing at $0.30 per million input tokens and $2.50 per million output tokens — roughly five times cheaper than 3.6 Flash on output. That makes it designed for volume-driven tasks like agentic search, document processing pipelines, and batch classification.
On Terminal-Bench 2.1, Flash-Lite hits 54% versus 31% for 3.1 Flash-Lite, and on long-context evals like GDM-MRCR v2 it scores 72.2% versus 60.1%. Flash-Lite also beats the larger 3 Flash on SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%), making it a strong default for cost-sensitive developers who previously had to choose between price and capability.
Flash Cyber Is Restricted, by Design
The third release is the most politically sensitive. Gemini 3.5 Flash Cyber is a specialised model fine-tuned for finding and patching software vulnerabilities, shipped inside Google’s CodeMender agent. Built on 3.5 Flash but tuned for cybersecurity workflows, it lets multiple agents collaborate on a single combined report. On the CyberGym benchmark, Google says 3.5 Flash Cyber hits competitive frontier performance.
Given the dual-use nature of an AI that can both find and fix vulnerabilities, Google is keeping 3.5 Flash Cyber off the open market. Access will be limited to governments and trusted partners through a CodeMender pilot, with Google framing it as giving “frontline defenders a head start in finding and fixing critical vulnerabilities before they can be exploited.” This mirrors the responsible-release pattern Anthropic and OpenAI have used for cyber-capable models.
What’s Next: Gemini 3.5 Pro and Gemini 4
Google also confirmed that Gemini 3.5 Pro is in private partner testing and will roll out broadly when ready. More significantly, the company revealed it has started its most ambitious pre-training run yet — for Gemini 4 — though no release window was shared. All three models are available today via the Gemini API in Google AI Studio, Android Studio, and the Gemini app, with enterprise access through the Gemini Enterprise Agent Platform. 3.5 Flash-Lite is also rolling out inside Google Search itself.
Our Take
Three models in one drop is aggressive even by Google’s 2026 cadence, and the pricing tells the real story. 3.6 Flash undercuts 3.5 Flash on output tokens while improving capability — a price-war move aimed at OpenAI’s GPT-5 family and Anthropic’s Claude Sonnet 5. Flash-Lite at $2.50 per million output tokens is the number that should worry competitors most, putting capable agentic inference within reach of high-volume use cases that were economically unviable six months ago.
For Malaysian developers and enterprises, the practical takeaway is straightforward: 3.6 Flash should be the default for new Gemini API integrations, Flash-Lite is the right pick for any high-volume pipeline like batch document processing for legal tech or customer support triage, and the Flash Cyber gating means cybersecurity startups won’t get to fine-tune on it directly. The bigger signal is Gemini 4 — Google wouldn’t mention a 4-series pre-training run alongside a 3.6 Flash refresh unless it wanted to remind the market that the frontier isn’t standing still.






