Gemini 3.8 Flash
Gemini 3.8 Flash is Google's generally available September 2026 workhorse model with the stable API ID gemini-3.8-flash, a 1,048,576-token input limit, 65,536-token output limit, text/image/video/audio/PDF input, and introductory pricing of $0.75 input and $3.75 output per million tokens through December 31, 2026.
Gemini 3.8 Flash combines frontier-adjacent agent performance with Flash pricing, fast output, broad multimodal input, and built-in Google tools. That makes it a practical value candidate for production systems that do not need the highest possible independent benchmark score or a restricted security model.
Google's launch, model page, pricing, and release notes control general availability, model ID, modalities, limits, tools, and promotional pricing. Google's launch benchmarks are vendor-reported. Artificial Analysis independently measured a score of 59 at high reasoning and a low cost per task on September 2; that result remains configuration- and date-specific.
- Run long-horizon coding and tool-using agents at a lower token price.
- Analyze mixed text, document, image, audio, and video inputs in one request.
- Use Google search, Maps, file search, code execution, and structured output in supported workflows.
- Route high-volume enterprise tasks where speed and cost matter alongside capability.
The stable model ID is gemini-3.8-flash. Google documents text, image, video, audio, and PDF input with text output, 1,048,576 input tokens, 65,536 output tokens, and low, medium, or high thinking. Caching, code execution, file search, function calling, structured output, search and Maps grounding, URL context, and preview computer use are supported.
Google lists an introductory rate of $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. Standard pricing from January 1, 2027 is $1.50 input and $7.50 output, so production budgets should record the promotion end date rather than treating the launch price as permanent.
Artificial Analysis measured Gemini 3.8 Flash at 59 on its September 2 Intelligence Index run and placed it on the intelligence-versus-cost Pareto frontier. It was much cheaper and faster than Claude Fable 5.1 in that harness, but Fable scored higher. The practical question is whether the extra quality of a frontier model changes task success enough to justify the added cost and latency.
An earlier GA Flash model that remains supported for efficiency-first workloads.
The restricted trusted-defender sibling built on the same foundational intelligence.
A higher-scoring but much more expensive generally available frontier alternative.
Source confidence
Google AI for Developers
Google AI for Developers
Artificial Analysis
Reddit / r/google_antigravity
Gemini 3.8 Flash FAQ
Common questions about Gemini 3.8 Flash.
Is Gemini 3.8 Flash generally available?+
Yes. Google lists gemini-3.8-flash as a stable GA model in the Gemini API and exposes it through AI Studio and other supported Google products.
How much does Gemini 3.8 Flash cost?+
The introductory rate through December 31, 2026 is $0.75 per million input tokens and $3.75 per million output tokens. Google lists standard pricing of $1.50 input and $7.50 output from January 1, 2027.
Should I choose Gemini 3.8 Flash or Claude Fable 5.1?+
Choose Gemini first when speed, multimodal inputs, Google tools, and cost dominate. Test Fable 5.1 when the hardest coding or knowledge-work tasks justify higher latency and a $10/$50 base rate. Compare both on the same tasks because benchmark configurations do not guarantee production outcomes.