
Google has unveiled three new Gemini models—Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber—with a clear focus on helping developers build faster, more efficient, and more affordable AI agents. Instead of chasing larger frontier models, Google is doubling down on performance, latency, and cost optimization for production AI workloads.
Gemini 3.6 Flash: Google’s New Workhorse Model
Google describes Gemini 3.6 Flash as its new workhorse model, delivering improvements in coding, multimodal understanding, reasoning, and knowledge work while reducing overall inference costs.
According to Google, the model:
Uses 17% fewer output tokens than Gemini 3.5 Flash
Requires fewer reasoning steps and tool calls
Delivers stronger coding and agent performance
Costs $1.50 per million input tokens and $7.50 per million output tokens, reducing the price of agentic workloads.
The company says these improvements make Gemini 3.6 Flash particularly well suited for long-running AI agents that need to complete complex, multi-step tasks efficiently.
Gemini 3.5 Flash-Lite Prioritizes Speed and Cost
Google also introduced Gemini 3.5 Flash-Lite, its fastest and most affordable Flash model to date.
Designed for high-volume workloads, Flash-Lite can generate approximately 350 output tokens per second and is optimized for applications such as:
Document processing
Translation
Summarization
Large-scale data extraction
Agent sub-tasks
The model offers adjustable reasoning levels, allowing developers to balance latency, cost, and intelligence depending on the workload.
Gemini 3.5 Flash Cyber Targets Cybersecurity
The third release, Gemini 3.5 Flash Cyber, is a specialized model built on Gemini 3.5 Flash and fine-tuned specifically for cybersecurity.
Integrated into CodeMender, Google’s AI-powered security platform, Flash Cyber is designed to:
Detect software vulnerabilities
Validate security issues
Recommend patches
Automate secure code reviews
Because of its dual-use nature, Google is initially making Flash Cyber available only to government agencies and trusted partners through a limited-access pilot program.
Built for Production AI Agents
Rather than focusing solely on benchmark leadership, Google’s latest releases emphasize what many enterprise customers care about most:
Lower inference costs
Better token efficiency
Faster response times
Reliable agent performance
Scalable production deployments
Google says the new models are intended to power AI agents that can reason, use tools, and execute long-running workflows without dramatically increasing infrastructure costs.
Available Today
Google confirmed that Gemini 3.6 Flash and Gemini 3.5 Flash-Lite are available immediately through:
Google AI Studio
Gemini API
Android Studio
Gemini Enterprise Agent Platform
Gemini 3.6 Flash is also available in Google Antigravity and the Gemini Enterprise app. Meanwhile, Google said Gemini 3.5 Pro remains in testing and will be released once it meets the company’s quality expectations.

Join the conversation! Your thoughts help the community grow.