GoogleVerified Active: September 18, 2026

Gemini 3.1 Flash-LiteAPI Pricing & Specifications

Ultra-low-latency multimodal model built for high-throughput mobile and streaming applications.

📊 Core Specifications & Pricing Rates

Standard Input
$0.25 / 1M tokens
Standard Output
$1.50 / 1M tokens
Cached Input
$0.0625 / 1M tokens
Context Window
1M (1,048,576 tokens)
Max Output Tokens
8,192 tokens
Supported Capabilities:✓ Vision✓ Function Calling / Tools✓ JSON ModeStandard Inference
📦 Asynchronous Batch API Supported: Official provider documentation offers a 50% discount on non-streaming 24h batch requests.

Live Workload Cost Estimator: Gemini 3.1 Flash-Lite

Adjust token volumes to simulate your exact monthly API bill.

Interactive Simulation
Estimated 60% repeated context
Estimated Total Monthly Cost
$20.16/ month
$0.00081 per API request(Saves $4.22/mo with caching)
Compare All Models in Calculator

💡 Estimated Monthly Workload Scenarios

Mobile Audio Transcription & Translation

$47.50/mo

50,000 audio processing requests (2,000 input, 300 output)

50,000 requests • 2000 in / 300 out

⚖️ Head-to-Head Competitor Comparisons

Gemini 3.1 Flash-Lite vs. GPT-5.1 Mini

Gemini 3.1 Flash-Lite provides 2.5x larger context (1M vs 400K) and 25% cheaper output ($1.50 vs $2.00).

1.33x cheaper outputView GPT-5.1 Mini

📜 Historical Price Changes & Updates

2026-07-20
Gemini 3.1 Flash-Lite introduced for ultra-low latency mobile use cases.

🔗 Verified Primary Sources

All pricing, token limits, and capabilities are cross-checked against official provider documentation and live API endpoints.

Google AI Developer Official Pricing