Research

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber for Scalable AI Agents

AI
AI Hub Feed
July 21, 20264 min read

Google is accelerating the development of AI agents with the introduction of its latest Gemini Flash models: Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and a specialized Gemini 3.5 Flash Cyber variant. These new models are engineered to meet the critical demands of developers and customers building production-ready AI agents, focusing on higher token efficiency, lower latency, and improved reliability to enable scaling agentic workflows.

Gemini 3.6 Flash: Enhanced Efficiency and Performance

Building directly on feedback from the Gemini 3.5 Flash model, Gemini 3.6 Flash represents a significant step forward. It not only delivers enhanced performance in coding and knowledge work but also achieves this with substantially improved token efficiency. According to the Artificial Analysis Index, 3.6 Flash consumes 17% fewer output tokens than 3.5 Flash, and in specific benchmarks like DeepSWE by Datacurve, reductions of up to 65% have been observed. This increased efficiency translates to fewer reasoning steps and tool calls for multi-step workflows, ultimately lowering the cost per output token. Priced at $1.50/1M input tokens and $7.50/1M output tokens, 3.6 Flash aims to make agentic tasks more cost-effective. Beyond efficiency, 3.6 Flash shows performance gains across various use cases, including higher precision in code edits and reduced execution loops in DeepSWE (49% vs. 37%), and significant improvements in ML Research benchmarks like MLE Bench (63.9% vs. 49.7%). Its computer use capabilities have also been enhanced, as seen in OSWorld-Verified benchmarks (83.0% vs. 78.4%), with computer use now integrated as a client-side tool via the Gemini API and Gemini Enterprise.

Gemini 3.5 Flash-Lite: Speed and Scalability for Agentic Workflows

Complementing 3.6 Flash, Google is also releasing Gemini 3.5 Flash-Lite. This model is specifically designed for low-latency tasks and scenarios where high throughput is paramount, such as agentic search and document processing. As the fastest model in the 3.5 series, 3.5 Flash-Lite achieves an impressive 350 output tokens per second, according to Artificial Analysis. With a price point of $0.3/1M input tokens and $2.5/1M output tokens, it offers a compelling price-to-performance ratio for high-volume production traffic. 3.5 Flash-Lite significantly outperforms prior Flash-Lite generations in agentic workflows and coding tasks, even surpassing some 3 Flash models in benchmarks like SWE-Bench Pro (54.2% vs. 49.6%) and OSWorld-Verified (74.0% vs. 65.1%). Developers can configure the model to prioritize low-latency execution for high-volume tasks or engage higher thinking levels for multi-step subagent workloads. The model now also includes computer use as a built-in tool, enhancing its reliability for agentic tasks across various platforms.

Gemini 3.5 Flash Cyber: Securing the Digital Frontier

Addressing the growing threat landscape in cybersecurity, Google has introduced Gemini 3.5 Flash Cyber. This specialized model is built upon 3.5 Flash and fine-tuned for the specific task of identifying and rectifying cybersecurity vulnerabilities with exceptional efficiency and at a lower cost per token than larger, general-purpose models. Paired with Google's CodeMender code security agent, 3.5 Flash Cyber agents work collaboratively to produce comprehensive reports, achieving competitive performance at the frontier on the CyberGym benchmark. Recognizing the dual-use nature of this technology, Google is adopting a controlled deployment strategy. 3.5 Flash Cyber will be exclusively available to governments and trusted partners through CodeMender as part of a limited-access pilot program. This approach aims to equip frontline defenders with advanced tools to proactively find and fix critical vulnerabilities while mitigating the risks of broader misuse.

Broader Gemini Ecosystem and Future Developments

These new Flash models are integrated into the broader Gemini ecosystem, making them accessible to developers and enterprises. 3.6 Flash and 3.5 Flash-Lite are available starting today via the Gemini API through Google AI Studio and Android Studio, with 3.6 Flash also accessible in Google Antigravity. For enterprises, they are available through the Gemini Enterprise Agent Platform and the Gemini Enterprise app. Consumers can access 3.6 Flash through the Gemini app, and 3.5 Flash-Lite is rolling out in Google Search. Google also provided an update on its future roadmap, noting that Gemini 3.5 Pro is currently in partner testing and planned for broader availability soon. Furthermore, the company has initiated its most ambitious pre-training run yet for Gemini 4, signaling continued investment in pushing the boundaries of AI capabilities.

Safety and Responsible Deployment

Safety remains a core consideration in the development of these new models. Gemini 3.6 Flash is shipping with enhanced Frontier Safety safeguards, particularly in Chemical, Biological, Radiological, and Nuclear (CBRN) domains and against cyber offense misuses. These safeguards are designed to make the model substantially more resistant to jailbreaking attempts while minimizing unnecessary refusals for beneficial uses. The controlled release of Gemini 3.5 Flash Cyber further underscores Google's commitment to responsible AI deployment, ensuring that powerful security tools are used ethically and effectively.

Related Articles

Google AI Blog