🔴 🤖 Models Published: · 3 min read ·

Google: Gemini 3.6 Flash and Gemini 3.5 Flash-Lite reach general availability with 17% fewer tokens and a new Cyber model

Illustration of the Gemini 3.6 Flash and 3.5 Flash-Lite models with speed and per-token pricing icons

Google has declared Gemini 3.6 Flash and Gemini 3.5 Flash-Lite generally available. Gemini 3.6 Flash uses 17% fewer output tokens than its predecessor at $1.50/$7.50 per million tokens, while Flash-Lite runs at 350 tokens per second for $0.30/$2.50. A specialized Gemini 3.5 Flash Cyber model was also introduced for governments and partners.

🤖

This article was generated using artificial intelligence from primary sources.

What is Gemini 3.6 Flash and how does it differ from its predecessor?

Gemini 3.6 Flash is Google’s fast, lower-cost model within the Gemini family — the Flash tier targets tasks that need lower latency and cost than the “Pro” models, with only a small quality trade-off. Google has declared the model General Availability (GA), meaning it has left the beta phase and is available to all developers for production use via the API. Pricing is $1.50 per million input tokens and $7.50 per million output tokens. The key change versus Gemini 3.5 Flash is efficiency: the new model uses 17% fewer output tokens for comparable tasks, directly lowering the cost per query. For companies building production AI products, General Availability status also means contractual guarantees of API stability — no sudden changes in model behavior of the kind possible during the preview phase.

Benchmarks: big jumps on DeepSWE, MLE-Bench and OSWorld

On the DeepSWE benchmark, which measures the ability to solve software tasks autonomously, Gemini 3.6 Flash scores 49% versus 37% for Gemini 3.5 Flash. On MLE-Bench, a benchmark for machine learning tasks, the score rises from 49.7% to 63.9%. On OSWorld-Verified, a test of computer interface control, the model scores 83.0% versus the earlier 78.4%. All three jumps show a consistent improvement in agentic capability — not just response speed, but also accuracy in multi-step tasks such as writing code and operating an interface. Together, DeepSWE, MLE-Bench and OSWorld-Verified cover three distinct agentic tasks — writing code, machine-learning engineering, and navigating a computer’s graphical interface — so a consistent score increase across all three suggests the improvement is not narrowly specialized to one task type, but stems from generally better model reasoning.

Gemini 3.5 Flash-Lite: the cheapest tier with a big jump on Terminal-Bench

Alongside the larger Flash model, Gemini 3.5 Flash-Lite — the cheapest and fastest tier in the Gemini family — also reaches general availability: $0.30 per million input tokens, $2.50 per million output tokens, at a speed of 350 tokens per second. On Terminal-Bench 2.1, a test of executing terminal commands, Flash-Lite scores 54% versus 31% for its predecessor — nearly double the improvement. On SWE-Bench Pro, the score rises from 49.6% to 54.2%. For developers choosing between the two tiers, the trade-off is clear: Flash-Lite sacrifices some raw capability for lower cost and faster generation, while the full Gemini 3.6 Flash retains greater precision at a higher cost per token.

Cyber model for governments only, and the announced Gemini 3.5 Pro

Within the CodeMender tool, Google also announced Gemini 3.5 Flash Cyber, a specialized model for automatically fixing security vulnerabilities in code, which will soon be available through a limited-access pilot program for governments and partner organizations, not the wider developer community. Google simultaneously announced Gemini 3.5 Pro, currently in the partner-testing phase ahead of a broader launch. That means the Gemini family gets two General Availability announcements in a single day (3.6 Flash and 3.5 Flash-Lite), one announced cybersecurity model, and one Pro model in preparation — a move that spans the range from the cheapest tier to a specialized, high-security tier.

Frequently Asked Questions

What does General Availability mean for Gemini models?
General Availability means the model has left the beta/preview phase and is available to all developers via the Google API for production use, without access restrictions.
How much faster and cheaper is Gemini 3.6 Flash than its predecessor?
Gemini 3.6 Flash uses 17% fewer output tokens than Gemini 3.5 Flash at $1.50 per million input tokens and $7.50 per million output tokens, and scores 49% versus 37% for its predecessor on the DeepSWE benchmark.
What is Gemini 3.5 Flash Cyber?
Gemini 3.5 Flash Cyber is a specialized model within the CodeMender tool for automatically fixing security vulnerabilities in code, which will soon be available through a limited-access pilot program for governments and partners.

📬 AI news in your inbox

A daily digest built your way — pick topics, sources and cadence. One-click unsubscribe.