>Gemini Flash is becoming a three-part agent stack
ALTIOR AI ADVANTAGEWhat to remember
Gemini Flash, divided by job

Three models, three bottlenecks

Google has turned Flash into a workhorse, a volume engine and a cyber specialist behind a permission gate.

A builder routes work through harder-agent, high-volume and locked-cyber lanes in a futuristic control room.

Google’s release is easier to understand as a team than a version list. Gemini 3.6 Flash takes the harder coding, knowledge and multimodal work; Gemini 3.5 Flash-Lite is built for large, repetitive queues; and Gemini 3.5 Flash Cyber specialises in vulnerability work inside CodeMender.

That division gives us a clearer routing decision, but not three equally available options. The first two models are available now across Google-listed surfaces. The cyber branch is reserved for governments and trusted partners in a forthcoming limited pilot.

Why it matters

The routing choice now matters

Each branch trades on a different combination of task difficulty, volume and access.

Two open workload lanes contrast with a permission-gated cyber-specialist route.
An explanatory view of the release: 3.6 Flash and Flash-Lite occupy the open work lanes, while Flash Cyber remains behind CodeMender’s controlled-access boundary.

For us, the practical question is not which model wins in the abstract. It is whether harder reasoning, repetitive throughput and sensitive security work should share the same route at all.

Google’s benchmark and efficiency figures can help us form a testable hypothesis. They cannot replace testing on our own task mix, especially when output quality, token use and latency can shift together.

The model map

What each Flash model is for

The family now spans harder agent work, high-volume processing and controlled cyber defence.

Three-card comparison of Gemini Flash branches by workload role and current or controlled access.
Gemini 3.6 Flash is the harder-work lane at $1.50/1M input tokens and $7.50/1M output tokens. Flash-Lite is the volume lane at $0.30/1M input tokens and $2.50/1M output tokens. Flash Cyber is the gated vulnerability-work lane. Prices and role descriptions are Google-published.
Provider image: google-deepmind-x-post-2079581707267665921.jpg
Google reports that Gemini 3.6 Flash used 17% fewer output tokens than 3.5 Flash on the Artificial Analysis Index. The comparison has not been independently reproduced here.

Google positions 3.6 Flash as its workhorse for coding, knowledge, multimodal and computer-use tasks. Its published price is $1.50 per million input tokens and $7.50 per million output tokens; Google also links its reported token efficiency to a lower cost per agentic task.

Flash-Lite is aimed at agentic search, document processing and other high-throughput work. Google cites an Artificial Analysis result of 350 output tokens per second and lists prices of $0.30 per million input tokens and $2.50 per million output tokens. That output-rate measure is not the same as end-to-end application latency.

The launch receipt

Google’s three-model release

The official announcement establishes the workhorse, volume and controlled cyber branches.

Provider image: google-blog-02.webp
Google’s launch material names Gemini 3.6 Flash, Gemini 3.5 Flash-Lite and Gemini 3.5 Flash Cyber, while separating the models available now from the forthcoming limited pilot.

The announcement supports the shape of the stack and its stated availability. Gemini 3.6 Flash and 3.5 Flash-Lite are rolling out now, including in the Gemini app, while developer and enterprise access varies by Google surface.

The performance, efficiency and reliability language remains Google’s account of the release. The evidence tells us what Google has announced; it does not independently establish how the models will perform in our systems.

Our newest Gemini models deliver the efficiency, latency, and reliability to build AI agents at scale.

Google Blog
The catch

The cyber model is gated

Flash Cyber’s specialist capability arrives with an explicit access boundary.

Provider image: google-blog-03.webp
Google says Flash Cyber will be available exclusively to governments and trusted partners through CodeMender, soon, as part of a limited-access pilot.

Google describes Flash Cyber as a specialised model for finding and fixing vulnerabilities, paired with its CodeMender security agent. It says multiple Flash Cyber agents can work together inside CodeMender to produce a combined report.

The access restriction is part of the story, not a footnote. Google cites the technology’s dual-use nature for its controlled deployment, so we cannot treat Flash Cyber as a public API option or a model available for general testing.

The selection pattern

Route work by its bottleneck

Difficulty, repetition and sensitivity point towards different branches of the stack.

Three-branch explanatory flow routing harder reasoning, high-volume tasks and vulnerability work to distinct model lanes.
A practical selection pattern—not Google’s published deployment architecture—routes harder reasoning to 3.6 Flash, repetitive volume to Flash-Lite and vulnerability work to the gated CodeMender branch.

We can treat 3.6 Flash as the colleague given the knottier case: multi-step coding, research, multimodal interpretation or tool-heavy work. Flash-Lite takes the long queue of repeatable jobs where throughput and unit cost matter.

Security work forms a separate branch. Flash Cyber may fit vulnerability discovery and remediation, but its route ends at an access check: unless we are participating as an approved government or trusted partner, that branch is not available to us.

The practical path

What we can use—and what we cannot

Start with access, match the workload, then test the published claims under real conditions.

Five-step journey covering model choice, access, cost review, workload testing and the restricted cyber branch.
Our path begins with 3.6 Flash or Flash-Lite on an available Google surface, moves through a workload-specific quality, token and speed test, and stops at the controlled gate for Flash Cyber.

We can use 3.6 Flash and 3.5 Flash-Lite now through the Google-listed developer, enterprise and consumer surfaces, with some surface-specific differences. We should choose an initial lane from the workload, not from the headline number alone.

Then we can measure quality, output-token use, throughput and total task cost on representative work. Flash Cyber remains outside that path unless Google grants access through the limited CodeMender pilot. Gemini 3.5 Pro is only in partner testing, while Gemini 4 is in pre-training.

The bigger shift

Flash is becoming a stack

The release makes model routing more important than chasing one universal default.

The useful unit is no longer one model. It is the route from each job to the right balance of capability, volume, cost and control.

Altior analysis, grounded in Google’s release

The stronger interpretation is also the more cautious one: Google is offering three answers to three operating constraints, not proof that one family will dominate every agent workflow.

Our advantage comes from making the split explicit and testing it. If Flash-Lite can clear routine volume while 3.6 Flash handles the difficult tail, the combination may matter more than either model’s headline result. That remains a workload hypothesis until our own evidence supports it.

Route a real agent workload

Act as an AI systems architect evaluating Gemini 3.6 Flash and Gemini 3.5 Flash-Lite for a customer-support agent. The system must classify and summarise 10,000 tickets each day, extract order details, draft routine replies and investigate the 5% of cases involving conflicting records or multi-step reasoning. Propose which model should handle each task, then design a controlled comparison using 200 representative tickets. Return: a routing table; quality, output-token, throughput and cost measures; pass/fail thresholds; an escalation rule; and a short decision note. Treat Google's published prices and performance figures as claims to test, not guaranteed results. Exclude Gemini 3.5 Flash Cyber because it is restricted to a limited CodeMender pilot.
Ready to copy
ALTIOR AI ADVANTAGE
The next move

Test the route, not the claim

Choose one representative queue, separate routine volume from harder cases, and compare the available Flash models against quality, token, speed and cost thresholds before changing production routing.

Try the prompt

What could change the decision

  • Flash Cyber pilot access
  • Gemini 3.5 Pro
  • Gemini 4
  • Our workload results