Gemini 3.5 Flash
GoogleLarge Language ModelProprietaryGoogle's high-efficiency multimodal model bringing near-Pro level coding and reasoning at Flash-tier cost. Optimized for agentic execution loops, supports text, image, video, audio, and PDF inputs with configurable thinking levels.
Abilities
Use Cases
Available in Tools
Availability
How to Use
Pros
- Near-Pro level intelligence at Flash pricing — strongest agentic and coding model in the Flash tier
- 1M token context window — same massive context as flagship models at fraction of the cost
- Configurable thinking levels — tune reasoning depth vs cost for each request
- Outperforms Gemini 3.1 Pro on agentic and coding benchmarks at lower cost
- Native multimodal — text, image, video, audio, and PDF inputs in one model
- Cost-effective at $1.50/$9 per MTok with 90% cached-token discount
Cons
- Maximum output limited to 65K tokens vs 128K for Claude models
- Closed source — no self-hosting option
- Thinking surcharge may apply for high-effort reasoning requests
What to Use It For
Perfect For
Strongest agentic and coding model in the Flash tier — outperforms Gemini 3.1 Pro at fraction of the cost
Optimized for parallel agentic execution loops — great for multi-agent orchestration at cost
Native text+image+video+audio+PDF support with 1M context for analyzing complex mixed-media documents
Good For
$1.50/1M input with near-Pro intelligence makes it practical for high-volume enterprise pipelines
Good combination of reasoning depth via thinking levels and speed for iterative tasks
Not Recommended
Complex scientific proofs and advanced math — Gemini 3 Pro or Claude Opus 4.8 are stronger
Try instead: Gemini 3 Pro, Claude Opus 4.8
Max output is 65K — Claude models support up to 128K for very long generations
Try instead: Claude Sonnet 5, Claude Opus 4.8
Do Not Use For
Language model only — no image generation; use Imagen 3 for that
Try instead: Imagen 3, FLUX.1
Closed source Google model — cannot run locally
Try instead: Gemma 3, Llama 4 Scout