Back to feed
News Story
APriority82
Google AI (X)
12 sources

Google Introduces Gemini 3.6 Flash and 3.5 Flash-Lite Models

Google has announced two new AI models: Gemini 3.6 Flash, which improves efficiency and quality over its predecessor, and Gemini 3.5 Flash-Lite, the fastest and most cost-effective model in the 3.5 class, designed for agentic workflows. Both models are available via the Gemini API and the Gemini App.

SynthePulse Insight · AI deep reading

Google Releases Three New Models Including Gemini 3.6 Flash: Efficiency Gains and Enhanced Safety

Version 1 · 1 source

On July 21, 2026, Google DeepMind launched three new models: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, focusing on higher token efficiency, lower latency, and greater reliability to scale AI agents. Meanwhile, Gemini 3.5 Pro is being tested with partners, and pre-training for Gemini 4 has begun.

  • Gemini 3.6 Flash reduces output token usage by 17% compared to 3.5 Flash, achieves up to 65% efficiency gains on the DeepSWE benchmark, and is priced lower (input $1.50/million tokens, output $7.50/million tokens).
  • 3.6 Flash shows performance improvements on multiple benchmarks: DeepSWE from 37% to 49%, MLE Bench from 49.7% to 63.9%, OSWorld-Verified from 78.4% to 83.0%, and GDPval-AA v2 from 1349 to 1421.
  • 3.5 Flash-Lite is the fastest and lowest-cost model in the 3.5 family, reaching up to 350 output tokens per second according to the Artificial Analysis Index, significantly outperforming its predecessor in agentic workflows.
  • 3.5 Flash Cyber is combined with the CodeMender code security agent, designed for cybersecurity applications, offering cutting-edge competitiveness.
  • 3.6 Flash enhances frontier safety in CBRN (chemical, biological, radiological, nuclear) and cyberattack misuse areas, while reducing refusals for beneficial uses.
  • Gemini 3.5 Pro is being tested with partners, and Gemini 4 has started its most ambitious pre-training run.
Open section navigationCore Model: Gemini 3.6 Flash

Core Model: Gemini 3.6 Flash

Gemini 3.6 Flash is the flagship model of this release, built on feedback from developers and customers regarding 3.5 Flash. According to the Artificial Analysis Index, it reduces output token usage by 17% compared to 3.5 Flash and achieves up to 65% efficiency gains on Datacurve's DeepSWE benchmark. Pricing is set at $1.50/million input tokens and $7.50/million output tokens, lower than 3.5 Flash.

In terms of performance, 3.6 Flash achieves 49% on DeepSWE (3.5 Flash: 37%), 63.9% on MLE Bench (3.5 Flash: 49.7%), 83.0% on OSWorld-Verified (3.5 Flash: 78.4%), and 1421 on GDPval-AA v2 (3.5 Flash: 1349). Customers such as Hebbia and Harvey report outstanding performance in multimodal tasks (document parsing, chart analysis, report drafting).

3.6 Flash also includes a built-in computer use client tool, available via the Gemini API and Gemini Enterprise.

Lightweight and Specialized Models: 3.5 Flash-Lite and 3.5 Flash Cyber

Gemini 3.5 Flash-Lite is the fastest and lowest-cost model in the 3.5 family, reaching up to 350 output tokens per second according to the Artificial Analysis Index, significantly outperforming the previous Flash-Lite in agentic workflows.

Gemini 3.5 Flash Cyber is combined with the CodeMender code security agent, designed for cybersecurity applications, offering cutting-edge competitiveness. This model requires careful orchestration with agent infrastructure.

Safety and Future Roadmap

3.6 Flash enhances frontier safety in CBRN and cyberattack misuse areas, making the model more resistant to jailbreak attacks while training to reduce refusals for beneficial uses. More information is available in the 3.6 Flash model card.

Gemini 3.5 Pro is currently being tested with partners and is planned for broad availability once ready. Meanwhile, the team has started the pre-training run for Gemini 4, the most ambitious pre-training project to date.

Credibility boundary

This article's information primarily comes from the official Google DeepMind blog, a first-party source. Benchmark results are provided by Google, with some data cited from third parties such as the Artificial Analysis Index and Datacurve, but independent verification is not provided. Customer feedback (e.g., Hebbia, Harvey) is relayed by Google without original citations.

Insight takeaway

Google balances efficiency and performance with Gemini 3.6 Flash, while also introducing a lightweight version and a cybersecurity-specific model, and clearly outlining the next-generation model roadmap. For developers building AI agents, 3.6 Flash offers lower costs and higher token efficiency, while 3.5 Flash-Lite is suitable for scenarios requiring extremely high speed.