Gemini 3.8 Flash 2026: Flash Cyber, Fairwind, and the evals
Google’s third Flash in six weeks: a public workhorse and a gated Cyber SKU. Split the SKUs before reading CyberGym or CWE-Bench.
Google shipped two Gemini 3.8 SKUs on 2 September 2026: a public Flash, and Flash Cyber behind Fairwind. It is the third Flash in six weeks, not a new flagship family.
The query is Gemini 3.8 Flash and Gemini 3.8 Flash Cyber. What follows matches the blog.google post and Digital Applied’s AA relay. Vendor benches are not independent replications. Prior note: 3.7 Flash.
What the launch post prints
Authors: product lead Tulsee Doshi and DeepMind security lead Raluca Ada Popa. Both SKUs share one foundation and long-running agentic loops. Flash is pitched for long-horizon coding and agents; named benches include DeepSWE v1.1, Vals Finance Agent V2, Harvey’s Legal Agent Benchmark, and 54.9% on HLE-Verified. On hard tasks the model takes extra reasoning steps and tool calls; higher effort can spend more tokens.
How to read the two rows
- 1
Name the public SKU before the gated one
Developers get 3.8 Flash on the Gemini API, AI Studio, Android Studio, Antigravity, and Stitch; enterprises via Gemini Enterprise; consumers in the Gemini app and Sheets. Cyber needs a Fairwind application — not the same API key. Control: Guidelight.
- 2
A near-frontier patch score is not a swap
CWE-Bench: Cyber at 47.2% pass@1 versus 47.8% for an unnamed frontier model; Google stresses a cheaper Pareto point. Wiz’s internal pentest bench: +7.5–9.7% recall at 2.3–5.2× lower cost. The Chrome 2.6× patch claim and Cloud Vulnerability Research’s “critical foundational vuln in under two hours” are Google-side cases, not a public reproduce pack.
- 3
Same token price is not the same task bill
Digital Applied lists AA index 59 and about $0.58 per task versus about $0.40 for 3.7 — same list price, higher task cost because it works harder. OfficeChai’s chart puts 3.8 Flash (high) at 59, below Muse Spark 1.3’s xhigh 61; see Muse Spark 1.3.
CyberGym and the 20-language internal bench
Google says Cyber beats 3.5 Flash Cyber and much larger frontier models on CyberGym. An internal bench spanning 20 languages is said to exceed 70% success. Those items are not public.
Patching before exploitation
The post says the investment order is fix-first, not exploit-first. CWE-Bench is run externally by Collinear. That is a defender narrative, not a public attack API.
3.7 is not deprecated
Google writes that 3.7 Flash remains fully supported for efficiency-first loads. Lower effort or staying on 3.7 is the stated path when tokens are the constraint.
| Claim | Checkable source | How to read it |
|---|---|---|
| HLE-Verified 54.9% | blog.google 2026-09-02 | Vendor-run; no independent harness |
| CWE-Bench 47.2% vs 47.8% | Collinear; counterpart unnamed | Near, cheaper — not a crown |
| AA high 59; ~$0.58 / task | Digital Applied September ledger | Third-party index, not Google’s table |
Google: both releases are powered by the same foundational intelligence, further accelerated by long-running agentic loops that recursively evaluate and refine the underlying models.
# Gemini 3.8 · Google blog 2026-09-02
# authors: Tulsee Doshi, Raluca Ada Popa
sku: gemini-3.8-flash # general
sku: gemini-3.8-flash-cyber # Fairwind only
price_intro: $0.75 / $3.75 per 1M (to 2026-12-31)
price_2027: $1.50 / $7.50
hle_verified: 54.9% # vendor-run
cwe_bench_pass1: 47.2% # Collinear
aa_index_high: 59 # Digital Applied / AA
3.7_flash: remains fully supportedBoundaries
- Vendor tables ≠ a third-party index
- DeepSWE, HLE, and the finance/legal agent benches are lab-named. AA 59 is a ledger relay. Do not mash them into one crown.
- The Cyber SKU ≠ a key anyone can buy
- Fairwind is for trusted defenders. The looser cyber mitigations are the access condition, not an open checkpoint.
- This page is not a meeting brief
- Release and eval only. For a short shared board, open wbmeet and create or join a room.
Questions worth checking
Are 3.8 Flash and Flash Cyber the same weights?
Google writes “the same foundational intelligence” in different deployment environments. Public docs do not publish a weight hash or downloadable checkpoint. Treat them as two SKUs with two access policies.
Is 54.9% or 59 the “official” score?
Neither is a regulator’s finding. 54.9% is Google’s HLE-Verified. 59 is Digital Applied’s AA high relay. Name the suite on any lab note.
How long does the intro price last?
Footnote: intro pricing through 31 December 2026. From 1 January 2027, $1.50 / $7.50 per million tokens.
Does this prove Google passed other labs?
No. CWE-Bench still sits a hair under an unnamed counterpart. On AA charts, 3.8 Flash high sits below several tie-band models. This page only unpacks the release.