← Back to all models

Gemini 3.5 Flash

GoogleLarge Language ModelProprietary

Google's high-efficiency multimodal model bringing near-Pro level coding and reasoning at Flash-tier cost. Optimized for agentic execution loops, supports text, image, video, audio, and PDF inputs with configurable thinking levels.

Abilities

Text GenerationCode GenerationReasoningVisionFunction CallingStructured OutputLong ContextMultilingual

Use Cases

Coding AssistantChatbotData AnalysisAutomationEnterpriseResearch

Available in Tools

Availability

Vertex AI

How to Use

Use the Gemini API with model ID `gemini-3.5-flash`. Supports thinking levels (minimal, low, medium, high) for fine-grained cost/performance trade-offs. Defaults to medium thinking effort.

Pros

  • Near-Pro level intelligence at Flash pricing — strongest agentic and coding model in the Flash tier
  • 1M token context window — same massive context as flagship models at fraction of the cost
  • Configurable thinking levels — tune reasoning depth vs cost for each request
  • Outperforms Gemini 3.1 Pro on agentic and coding benchmarks at lower cost
  • Native multimodal — text, image, video, audio, and PDF inputs in one model
  • Cost-effective at $1.50/$9 per MTok with 90% cached-token discount

Cons

  • Maximum output limited to 65K tokens vs 128K for Claude models
  • Closed source — no self-hosting option
  • Thinking surcharge may apply for high-effort reasoning requests

What to Use It For

star

Perfect For

Agentic coding at scale

Strongest agentic and coding model in the Flash tier — outperforms Gemini 3.1 Pro at fraction of the cost

Long-horizon agentic task pipelines

Optimized for parallel agentic execution loops — great for multi-agent orchestration at cost

Multimodal document and media processing

Native text+image+video+audio+PDF support with 1M context for analyzing complex mixed-media documents

thumb_up

Good For

Cost-effective enterprise automation

$1.50/1M input with near-Pro intelligence makes it practical for high-volume enterprise pipelines

Research and iterative analysis workflows

Good combination of reasoning depth via thinking levels and speed for iterative tasks

warning

Not Recommended

Tasks requiring maximum reasoning depth

Complex scientific proofs and advanced math — Gemini 3 Pro or Claude Opus 4.8 are stronger

Try instead: Gemini 3 Pro, Claude Opus 4.8

Long output generation over 65K tokens

Max output is 65K — Claude models support up to 128K for very long generations

Try instead: Claude Sonnet 5, Claude Opus 4.8

block

Do Not Use For

Image generation

Language model only — no image generation; use Imagen 3 for that

Try instead: Imagen 3, FLUX.1

Self-hosting

Closed source Google model — cannot run locally

Try instead: Gemma 3, Llama 4 Scout

Technical Details

Pricing$1.50 / 1M input tokens, $9 / 1M output tokens
ParametersUndisclosed
detail.contextWindow1M tokens (65K max output)

Links