Back to feed
News Story
少数派
3 sources

Google Releases Gemini 3.6 Flash and Two Other Models

On July 21, Google released three new models: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber, focusing on improving AI agent efficiency, latency, and reliability. Gemini 3.6 Flash outperforms its predecessor on coding and computer operation benchmarks while reducing API prices. The Flash-Lite model targets high-throughput tasks, and the Cyber model is optimized for cybersecurity.

SynthePulse Insight · AI deep reading

Google Launches Three Flash Models: Cost Reduction and Efficiency Gains, but Flagship Missing

Version 1 · 3 sources

With flagship model Gemini 3.5 Pro still delayed and Gemini 4 in pre-training, Google released three Flash series models on July 21, 2026, focusing on lower cost and higher efficiency, but the iteration speed raises market concerns.

  • Google released Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber, all optimized for agent scenarios with improved efficiency and cost.
  • Gemini 3.6 Flash reduces output token usage by an average of 17%, with API output price dropping from $9 to $7.50 per million tokens.
  • Gemini 3.5 Flash-Lite achieves a speed of 350 tokens per second, with input price at only $0.30 per million tokens, suitable for high-throughput tasks.
  • Gemini 3.5 Flash Cyber specializes in software vulnerability discovery and remediation, available for limited testing via CodeMender to government agencies and trusted partners.
  • Gemini 3.5 Pro is still in testing, and Gemini 4 has begun pre-training, but Google's model iteration speed is perceived as slower than competitors.
Open section navigationThree Flash Models Released: Efficiency and Cost Reduced

Three Flash Models Released: Efficiency and Cost Reduced

On July 21, 2026, Google released three models: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber, all part of the Flash series emphasizing speed and low cost. Among them, Gemini 3.6 Flash is the successor to 3.5 Flash, targeting programming, knowledge work, and multimodal tasks. It reduces output token usage by an average of 17% and decreases the number of reasoning steps and tool calls needed for multi-step tasks. API input price is $1.50 per million tokens, and output price drops from $9 to $7.50 per million tokens.

Gemini 3.5 Flash-Lite is designed for high-throughput tasks such as agent search and document processing, achieving a speed of 350 tokens per second. API input and output prices are $0.30 and $2.50 per million tokens, respectively. Gemini 3.5 Flash Cyber is optimized for software vulnerability discovery, verification, and remediation, and will first be available for limited testing via CodeMender to government agencies and trusted partners.

Limited Performance Gains, Flagship Model Absent

Google's test results show that Gemini 3.6 Flash achieved a score of 49% on the programming benchmark DeepSWE, up from 37% for Gemini 3.5 Flash; on the OSWorld-Verified computer operation test, the score improved from 78.4% to 83.0%. However, third-party observers noted that the new model 'improved a bit on each metric, but not by much.'

Notably, the previously rumored flagship model Gemini 3.5 Pro did not debut with this release. Google stated that Gemini 3.5 Pro is still being tested with partners, while the next-generation Gemini 4 has begun pre-training, described as 'the largest training run ever.' However, some commentators believe that Google's model iteration speed 'feels really slow compared to other models that update every two to three months, or even monthly.'

Market Reaction and Strategic Intent

On the day before Alphabet's earnings report, Google chose to release three Flash models, but the market reaction was lukewarm. Reports indicate that Google's market value evaporated by $200 billion overnight, though this figure may be influenced by multiple factors. By focusing on the Flash series, Google aims to consolidate its position in agent application scenarios through cost reduction and efficiency gains, but the absence of a flagship model and the lagging iteration pace put it under competitive pressure.

Credibility boundary

This article is based on reports from Shaoshupai, AI Qianxian, and X platform user 歸藏 (guizang.ai). Model performance data comes from official Google announcements, market value evaporation data from AI Qianxian reports, and iteration speed comments from third-party observations, all marked as source claims or inferences.

Insight takeaway

Google responded to market demands for efficiency and cost with three Flash models, but the absence of a flagship model and lagging iteration speed create uncertainty in the AI race.