Gemini 3.7 Flash
Gemini 3.7 Flash is Google's workhorse model for coding and agents, with a 1M tokens context window, text, image, audio, and video input, and configurable thinking. Your use is subject to Google's Terms & Privacy Policies.
import { streamText } from 'ai'
const result = streamText({ model: 'google/gemini-3.7-flash', prompt: 'Why is the sky blue?'})Copy link to headingPlayground
Try out Gemini 3.7 Flash by Google. Usage is billed to your team at API rates. Free users (those who haven't made a payment) get $5 of credits every 30 days.
Gemini 3.7 Flash
Copy link to headingProviders
Route requests across multiple providers. Copy a provider slug to set your preference. Visit the docs for more info. Using a provider means you agree to their terms, listed under Legal.
| Provider |
|---|
Copy link to headingThroughput24 hours
P50 throughput on live AI Gateway traffic, in tokens per second (TPS). Visit the docs for more info.
Copy link to headingLatency24 hours
P50 time to first token (TTFT) on live AI Gateway traffic, in milliseconds. View the docs for more info.
Copy link to headingUptime24 hours
Direct request success rate on AI Gateway and per-provider. Visit the docs for more info.
Copy link to headingAbout Gemini 3.7 Flash
Gemini 3.7 Flash was released August 13, 2026 as Google's workhorse model for coding and agents, arriving three weeks after Gemini 3.6 Flash. It refines the reasoning foundation of 3.6 Flash rather than starting from a new pretraining run, and the improvements land mainly in coding and agentic execution.
Gemini 3.7 Flash accepts text, images, audio, and video, returns text, and works within a context window of 1M tokens with up to 65.5K tokens per response. Thinking is configurable per request, so you can spend more tokens on a hard problem and fewer on a routine one. The knowledge cutoff is March 2026.
Coding and agentic benchmarks moved the most between generations. Gemini 3.7 Flash scores 65.3% on DeepSWE v1.1, up from 49.0%, and 43.6% on FrontierCode 1.1 Main, up from 34.4%. Document comprehension on GDP.pdf reaches 34%, and long-context retrieval on GDM-MRCR v2 at 128k reaches 97.0%, which matters if you intend to actually fill the context window rather than just have it available.
Pricing is introductory and set to rise at the end of 2026. See the pricing panel on this page for current rates.
You can integrate Gemini 3.7 Flash through AI SDK, Chat Completions API, Responses API, Messages API, or other API formats, from TypeScript or Python. Routing rules move traffic from another Gemini model to Gemini 3.7 Flash without changing application code.
Copy link to headingWhat To Consider When Choosing a Provider
- Configuration: Current pricing is introductory and expires at the end of 2026, after which the rate roughly doubles. Anything you size on today's numbers should be re-checked against the pricing panel on this page before it becomes a long-term commitment.
- Configuration: Gemini 3.7 Flash is a refinement of Gemini 3.6 Flash rather than a new foundation, so the gains are concentrated in coding and agentic work. On workloads outside those areas the difference from 3.6 Flash may be small enough not to justify a migration.
- Configuration: The knowledge cutoff is March 2026, so pair Gemini 3.7 Flash with web search or retrieval for anything more recent. Thinking is configurable, which means an unconfigured request can spend more output tokens than you expect on a simple task.
- Zero Data Retention: Zero Data Retention is available for this model. It is offered on a per-provider and model basis. See the documentation for details.
- Authentication: AI Gateway authenticates requests using an API key or OIDC token. You do not need to manage provider credentials directly.
Copy link to headingWhen to Use Gemini 3.7 Flash
Best for
- High-Volume Coding: Workhorse pricing on software work that runs constantly
- Long-Context Retrieval: Accurate recall across a filled 1M tokens window
- Multimodal Input: Text, images, audio, and video in one request
- Configurable Thinking: Token spend tuned per request instead of a fixed budget
- Agentic Execution: The area that improved most over Gemini 3.6 Flash
Consider alternatives when
- Frontier Reasoning: A Pro-tier model handles the hardest problems
- Long-Term Price Stability: The introductory rate expires at the end of 2026
- Non-Coding Workloads: Gains over Gemini 3.6 Flash are smaller outside coding
- Post-Cutoff Facts: March 2026 knowledge needs web search or retrieval
Copy link to headingConclusion
Gemini 3.7 Flash is Google's workhorse for coding and agents, with a 1M tokens multimodal context window and thinking you configure per request. Point google/gemini-3.7-flash at AI Gateway to route requests behind one API key, and re-check the pricing panel before the introductory rate ends in December 2026.