Launch of Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Launch of Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber

Artificial intelligence is entering the phase of operational maturity. Today, the measure of success for enterprise AI implementations is no longer just the quality of the generated text, but above all cost efficiency, low latency, reliability, and the ability to autonomously execute multi-step tasks. Google’s answer to these challenges is the latest generation of models in the Gemini family: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and the specialized Gemini 3.5 Flash Cyber model. We take a closer look at what these changes mean in implementation practice and how they allow for the optimization of cloud budgets.

Why do agentic systems require a new AI model architecture?

Until now, the development of large language models (LLMs) has focused mainly on increasing parameter scale. However, from the perspective of software engineering and cloud architectures, running massive models for every single micro-task generates immense costs and unacceptable latency. As organizations transition from simple chatbots to complex agentic systems (Agentic Workflows)—where a single process involves dozens of tool calls, code analyses, or API callbacks—an architecture optimized for token efficiency becomes crucial.

Developers building Enterprise-grade solutions need models that:

  • Consume fewer output tokens to achieve the same goal.
  • Execute tasks in fewer inference steps.
  • Deliver drastically higher throughput while maintaining an affordable price.

Google addresses these needs with a precisely designed portfolio of Flash models, proving that speed and high quality can go hand in hand.

Gemini 3.6 Flash: higher quality, better coding, and up to 65% fewer tokens

At the heart of Google's new offering is Gemini 3.6 Flash. The model was designed directly based on feedback from developers and customers of the previous version.

Key parameters and token optimization

The most significant innovation in Gemini 3.6 Flash is a major improvement in response conciseness and precision. According to independent tests by the Artificial Analysis Index, this model consumes an average of 17% fewer output tokens compared to Gemini 3.5 Flash. In specialized software engineering tasks (such as the DeepSWE benchmark developed by Datacurve), the drop in token consumption reaches up to 65%.

Fewer tokens directly translate to a lower cloud bill and significantly faster application response times. At rates around $1.50 per 1 million input tokens and $7.50 per 1 million output tokens, Gemini 3.6 Flash drastically lowers the Total Cost of Ownership (TCO) for complex multi-agent workflows.

Performance gains in production benchmarks

Although Gemini 3.6 Flash is clearly more economical, higher efficiency goes hand in hand with a real leap in quality compared to the 3.5 Flash release. This is visible across all key business applications:

  • Programming and Code Engineering: Gemini 3.6 Flash makes fewer errors and performs fewer unnecessary operations. In DeepSWE tests, its effectiveness jumped from 37% to 49%. Meanwhile, on the MLE Bench benchmark, which evaluates advanced Machine Learning tasks, the model achieved 63.9% (compared to the previous 49.7%).
  • Interface Control (Computer Use): Operating system integration has been significantly improved. Success rates in OSWorld-Verified tests rose to 83.0% (from 78.4%), and the feature itself is now natively available in Gemini Enterprise and via API.
  • Knowledge and Document Work: The model clearly outperforms its predecessor in GDPval-AA v2 analytical tests (1,421 vs. 1,349 points). Tech companies like Harvey and Hebbia highlight its exceptional proficiency in multimodal processing—ranging from advanced data and chart analysis to the automated generation of syntheses and reports.

Thanks to its multimodal capabilities, the model excels at the parallel analysis of financial documents, charts, system architectures, and business reports.

Gemini 3.5 Flash-Lite: lightning-fast speed for scaling sub-agents

Where ultra-high throughput and the lowest cost per operation matter most, Gemini 3.5 Flash-Lite comes into play.

Throughput of 350 tokens per second

Gemini 3.5 Flash-Lite was designed for large-scale tasks, such as bulk document processing, agentic search, and real-time data classification. The model generates 350 output tokens per second (measured by Artificial Analysis), costing just $0.30 per 1 million input tokens and $2.50 per 1 million output tokens.

Flexible thinking levels

Engineers can dynamically configure Gemini 3.5 Flash-Lite’s operational mode depending on the nature of the workload:

  • Low thinking level: Dedicated to simple, high-volume automation tasks with minimal latency.
  • Higher thinking level: Activated when the model operates as a sub-agent executing complex, multi-step task breakdowns.

In many benchmarks related to coding and autonomous agents, Gemini 3.5 Flash-Lite outperforms the previous Gemini 3 Flash model—e.g., reaching 54.2% on SWE-bench Pro (compared to 49.6%), and scoring 54% vs. 31% on Terminal-Bench 2.1.

Gemini 3.5 Flash Cyber and CodeMender: A new standard for cybersecurity

As software complexity has grown, the pace of AI tools detecting vulnerabilities has outpaced the ability of IT teams to patch them manually. The race against time in cybersecurity required the creation of a specialized remediation mechanism.

Dedicated code security agent

Google introduced the Gemini 3.5 Flash Cyber model, optimized for precise searching, verification, and automated patch generation for security vulnerabilities in source code.

This model collaborates with the CodeMender agent, which utilizes multi-agent orchestration of multiple Gemini 3.5 Flash Cyber models running in parallel. This allows for comprehensive code analysis and the generation of a consolidated report accompanied by ready-to-deploy patches.

Due to the advanced nature of this technology, Gemini 3.5 Flash Cyber is being rolled out as part of a controlled pilot program for governments and trusted partners.

How to deploy the new Gemini models in enterprise architecture?

The new models are available to developers and enterprises across a variety of channels:

  • For Developers: Access via the Gemini API in Google AI Studio, Android Studio, and Google Antigravity.
  • For Enterprises: Integration within the Gemini Enterprise Agent platform and Gemini Enterprise application.
  • For End Users: Availability in the Gemini app and integration of Gemini 3.5 Flash-Lite into Google Search.

Summary for IT architects and digital leaders

The introduction of Gemini 3.6 Flash and 3.5 Flash-Lite is a game-changer in designing cloud-based AI systems. Thanks to a significant reduction in required tokens, ultra-fast response times, and specialized agentic capabilities, companies can build scalable and secure solutions without compromising on quality.

Frequently Asked Questions

1. How does Gemini 3.6 Flash differ from Gemini 3.5 Flash? Gemini 3.6 Flash generates up to 17% fewer output tokens (and up to 65% fewer in coding tasks) while offering higher accuracy in knowledge work, an improved task-solving rate for programming, and native support for computer interface control.

2. What are the best use cases for Gemini 3.5 Flash-Lite? The Flash-Lite model is ideal for high-frequency call tasks with strict latency requirements. It excels as a sub-agent in complex orchestration systems and for bulk document data extraction.

3. What is the CodeMender agent with the Gemini 3.5 Flash Cyber model? CodeMender is a specialized automated security remediation system that uses Gemini 3.5 Flash Cyber models to identify, verify, and automatically fix vulnerabilities in source code.

4. How does Gemini 3.5 Flash-Lite differ from standard Gemini Flash? Flash-Lite is optimized for maximum throughput (350 tokens/s) and the lowest price. It was created for large-scale tasks and to serve as a sub-agent supporting larger models.

5. Where can you try the new Gemini models? Developers can access them via the Gemini API in Google AI Studio or Google Antigravity, while business customers can access them within the Gemini Enterprise platform.

All entries All from category: News