Google Gemini API Cost Calculator
Project your Google Gemini API costs and evaluate context caching and free-tier values.
Gemini context caching saves 75% on inputs if your prompt is larger than 32,768 tokens (Flash) or 32,768 tokens (Pro).
What is the Google Gemini API Cost Calculator?
The Google Gemini API Cost Calculator helps developers and startup teams estimate monthly API expenses for Gemini 1.5 Flash and Gemini 1.5 Pro models. Gemini is famous for its massive 1M+ token context window, which makes managing input costs critical. This calculator models Gemini's tiered pricing structure (discounted rates for prompts under 128k tokens vs standard rates above 128k), context caching savings (which drop cached input costs by 75%), and incorporates free-tier allotments to project your true monthly bill.
How Does the Google Gemini API Cost Calculator Work?
1. Select Gemini Model â Choose Gemini 1.5 Flash (high speed) or Gemini 1.5 Pro (complex reasoning).
2. Set Monthly Requests â Define the average number of API requests your application runs per month.
3. Define Token Counts â Input average prompt (input) and response (output) sizes.
4. Configure Context Caching â Set the percentage of inputs read from cache. Caching reduces input costs by 75%.
5. Toggle Free Tier â Choose whether to apply Google AI Studio's monthly free allotments to offset your bill.
6. Analyze Costs â Review your total bill, caching savings, and free-tier valuations.
Formula & Calculation Method
Tiered Pricing Brackets (per million tokens):
- Gemini 1.5 Flash (<128k): Input: $0.075 | Output: $0.30
- Gemini 1.5 Flash (>128k): Input: $0.15 | Output: $0.60
- Gemini 1.5 Pro (<128k): Input: $1.25 | Output: $5.00
- Gemini 1.5 Pro (>128k): Input: $2.50 | Output: $10.00
- Context Cache Read Rate: 25% of standard input rate (75% savings).
Equations:
- Free Tier Allotment: Flash gets 15,000 free requests/mo. Pro gets 1,500 free requests/mo.
- Billable Requests: Total Requests - Free Tier Allotment
- Input Cost (Cached): ((Uncached Tokens à Input Rate) + (Cached Tokens à Cache Read Rate)) / 1,000,000 à Billable Requests
Example Calculation
Example: App running 20,000 requests/mo to Gemini 1.5 Flash, 15,000 input tokens (50% cached), 1,000 output tokens. Free tier active.
- Billable Requests: 20,000 total - 15,000 free limit = 5,000 billable requests
- Standard Cost (No Caching on billable portion):
- Input Cost = (15,000 / 1M) Ã $0.075 Ã 5,000 = $5.625
- Output Cost = (1,000 / 1M) Ã $0.30 Ã 5,000 = $7.50
- Total standard bill: $13.13 / month
- With Context Caching (50% cached):
- Cached Tokens = 7,500 | Uncached = 7,500
- Uncached Input = (7,500 / 1M) Ã $0.075 Ã 5,000 = $2.81
- Cached Input = (7,500 / 1M) Ã ($0.075 Ã 25%) Ã 5,000 = $0.70
- Output Cost = $7.50
- Total Optimized Bill: $11.01 / month (plus $39.38 saved by the Free Tier!).
Frequently Asked Questions
**Tiered Pricing Brackets (per million tokens):** - **Gemini 1.5 Flash (<128k):** Input: $0.075 | Output: $0.30 - **Gemini 1.5 Flash (>128k):** Input: $0.15 | Output: $0.60 - **Gemini 1.5 Pro (<128k):** Input: $1.25 | Output: $5.00 - **Gemini 1.5 Pro (>128k):** Input: $2.50 | Output: $10.00 - **Context Cache Read Rate:** 25% of standard input rate (75% savings). **Equations:** - **Free Tier Allotment:** Flash gets 15,000 free requests/mo. Pro gets 1,500 free requests/mo. - **Billable Requests:** Total Requests - Free Tier Allotment - **Input Cost (Cached):** ((Uncached Tokens à Input Rate) + (Cached Tokens à Cache Read Rate)) / 1,000,000 à Billable Requests
Disclaimer: This tool is provided for informational and calculation purposes. Output values are estimates based on standard user inputs.