AI KNOWLEDGE DESK

Models · entities · concepts · comparisons · practical tools

GETLLMS.ORG
ModelMultimodal agent models

Gemini 3.8 Flash

Gemini 3.8 Flash is Google's generally available September 2026 workhorse model with the stable API ID gemini-3.8-flash, a 1,048,576-token input limit, 65,536-token output limit, text/image/video/audio/PDF input, and introductory pricing of $0.75 input and $3.75 output per million tokens through December 31, 2026.

Why it matters

Gemini 3.8 Flash combines frontier-adjacent agent performance with Flash pricing, fast output, broad multimodal input, and built-in Google tools. That makes it a practical value candidate for production systems that do not need the highest possible independent benchmark score or a restricted security model.

Source-backed summary

Google's launch, model page, pricing, and release notes control general availability, model ID, modalities, limits, tools, and promotional pricing. Google's launch benchmarks are vendor-reported. Artificial Analysis independently measured a score of 59 at high reasoning and a low cost per task on September 2; that result remains configuration- and date-specific.

Primary use cases
  • Run long-horizon coding and tool-using agents at a lower token price.
  • Analyze mixed text, document, image, audio, and video inputs in one request.
  • Use Google search, Maps, file search, code execution, and structured output in supported workflows.
  • Route high-volume enterprise tasks where speed and cost matter alongside capability.
Stable API and broad multimodality

The stable model ID is gemini-3.8-flash. Google documents text, image, video, audio, and PDF input with text output, 1,048,576 input tokens, 65,536 output tokens, and low, medium, or high thinking. Caching, code execution, file search, function calling, structured output, search and Maps grounding, URL context, and preview computer use are supported.

Promotional and standard pricing

Google lists an introductory rate of $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026. Standard pricing from January 1, 2027 is $1.50 input and $7.50 output, so production budgets should record the promotion end date rather than treating the launch price as permanent.

How the independent result changes the choice

Artificial Analysis measured Gemini 3.8 Flash at 59 on its September 2 Intelligence Index run and placed it on the intelligence-versus-cost Pareto frontier. It was much cheaper and faster than Claude Fable 5.1 in that harness, but Fable scored higher. The practical question is whether the extra quality of a frontier model changes task success enough to justify the added cost and latency.

Gemini 3.8 Flash FAQ

Common questions about Gemini 3.8 Flash.

Is Gemini 3.8 Flash generally available?+

Yes. Google lists gemini-3.8-flash as a stable GA model in the Gemini API and exposes it through AI Studio and other supported Google products.

How much does Gemini 3.8 Flash cost?+

The introductory rate through December 31, 2026 is $0.75 per million input tokens and $3.75 per million output tokens. Google lists standard pricing of $1.50 input and $7.50 output from January 1, 2027.

Should I choose Gemini 3.8 Flash or Claude Fable 5.1?+

Choose Gemini first when speed, multimodal inputs, Google tools, and cost dominate. Test Fable 5.1 when the hardest coding or knowledge-work tasks justify higher latency and a $10/$50 base rate. Compare both on the same tasks because benchmark configurations do not guarantee production outcomes.