Skip to content
All systems operational0 AI providers monitored, polled every 2 minutes
Live status

Gemini 3.6 Flash

Mid-tier

by Google

Gemini 3.6 Flash went generally available on July 21, 2026, and the pitch is efficiency rather than raw intelligence. The model id is gemini-3.6-flash. It launched at $1.50 per million input tokens and $7.50 per million output, a cut from the $9.00 output rate on Gemini 3.5 Flash, and Google has since put it on the same introductory schedule as 3.7 and 3.8 Flash: $0.75 and $3.75 through December 31, 2026, then back to $1.50 and $7.50 from January 1, 2027. It carries a 1,048,576 token input window, up to 65,536 output tokens, and takes text, image, audio, video, and PDF input. Google says it is built directly on Gemini 3.5 Flash and consumes 17 percent fewer output tokens on the Artificial Analysis Index, with up to 65 percent fewer on DeepSWE. Google's published gains over 3.5 Flash are DeepSWE 49 percent against 37, MLE-Bench 63.9 against 49.7, OSWorld-Verified 83.0 against 78.4, and GDPval-AA v2 at 1421 Elo against 1349. The honest caveat is that Artificial Analysis scored the Intelligence Index flat at 50, identical to 3.5 Flash, while measuring average time per task falling from 2.7 minutes to 1.3 and measured cost per task from $0.59 to $0.50. Read it as a per-task economics upgrade for existing Flash workloads, not a capability jump. Computer use is now a built-in client-side tool through the Gemini API, and the knowledge cutoff is March 2026. Day-one availability spans Google AI Studio, the Gemini API, Android Studio, Google Antigravity, the Gemini Enterprise Agent Platform, and the Gemini app.

Input Price

$0.75

per 1M tokens

Output Price

$3.75

per 1M tokens

Context Window

1.0M

tokens

Released

2026-07

API access

Capabilities

textvisionaudiovideotool-usecodereasoning

Key Strengths

  • $0.75/$3.75 through December 31, 2026, then $1.50/$7.50
  • 17 percent fewer output tokens on the Artificial Analysis Index
  • Measured time per task roughly halved, 2.7 minutes to 1.3
  • OSWorld-Verified 83.0 percent with computer use as a built-in tool
  • 1M token input window with 64K output
  • Text, image, audio, video, and PDF input

Best For

  • Long-running agentic workflows where token spend dominates
  • Agentic coding and multi-step tool use
  • Computer-use and GUI automation
  • Document parsing, chart analysis, and report drafting

Benchmark Scores

BenchmarkScoreDescription
SWE-bench79.6Real-world software engineering tasks from GitHub issues (SWE-bench Verified)
MMLU-Pro89.3General knowledge and reasoning across 57 subjects
GPQA Diamond93.4Graduate-level science questions verified by domain experts
OSWorld 2.033.8Computer use across real desktop applications and multi-step GUI tasks
FrontierCode v1.134.4Agentic coding on frontier software engineering tasks (Main split)

Scores sourced from public benchmark datasets. See full benchmark leaderboard for all models.

Pricing Details

Input tokens

$0.75

per 1M tokens

Output tokens

$3.75

per 1M tokens

Estimated cost per 1K requests

$2.62

~1K input + ~500 output tokens avg

Prices are subject to change. Check the official documentation for current pricing. See the cost calculator for detailed estimates.

Related Models

View DocumentationCompare ModelsCost CalculatorFull Pricing Guide