Google · Free
Gemini 3.1 Flash-Lite (Preview)
Maximum Throughput, Minimum Cost
Context window
1M
tokens
Max output
66K
tokens
Tier
Free
access
Pricing
Input
$0.25
per 1M tokens
Output
$1.50
per 1M tokens
Overview
Overview
Gemini 3.1 Flash-Lite is Google's most cost-effective multimodal model. It is built for the fastest performance on lightweight, high-frequency workloads, making it a strong fit for large-scale agent tasks, simple data extraction, translation, transcription, and ultra-low-latency applications where budget and speed matter most.
Best Use Cases
✅ Recommended For
- High-Volume Agents: Lightweight routing, classification, and orchestration at scale
- Translation: Fast and inexpensive multilingual processing
- Transcription: Audio-to-text workflows without a separate speech pipeline
- Structured Extraction: Entity extraction, classification, and JSON outputs from text or documents
- Document Summaries: Quick summaries over PDFs and mixed multimodal inputs
⚠️ Consider Alternatives For
- Complex Reasoning: Gemini 3.1 Pro is better for deeper multi-step reasoning
- Advanced Coding: Gemini 3 Flash or 3.1 Pro are better choices for heavier coding tasks
Technical Highlights
| Feature | Details |
|---|---|
| Context Window | 1,048,576 input tokens |
| Max Output | 65,536 output tokens |
| Pricing | $0.25 / $1.50 per MTok (text, image, video input / output), audio input at $0.50 per MTok |
| Thinking | ✅ Supported with thinkingLevel (minimal, low, medium, high), dynamic high by default |
| Multimodal Input | ✅ Text, image, video, audio, and PDF |
| Tools | ✅ Function calling, structured output, Google Search grounding, URL context, file search, and code execution |
| Cache & Batch | ✅ Context caching and Batch API supported |
| Streaming | ✅ Full support |
| Not Supported | No Live API, computer use, Google Maps grounding, image generation, or audio generation |
Capabilities
- Cost Effective
- Fastest
- Agents
Recommended for
- Simple Tasks
- Quick Tasks
- High Volume
- Translation