Artificial Intelligence
AI API Pricing 2026: What Every Model Actually Costs
Every current model's token pricing in one table, including the promotional rates that expire and the long-context rates that catch people out.


Quick verdict
Gemini 3.7 Flash is the cheapest credible model at $0.75 input and $3.75 output per million tokens, but that rate is promotional through 31 December 2026. For pricing you can plan around, GPT-5.6 Terra at $2 and $12 and Claude Sonnet 5 at $2 and $10 are the stable mid-tier options - and Sonnet 5's rate was made permanent on 10 August 2026.
Best default
Cheapest tiers
Alternative
Flagship tiers
India check
Billing and data policy
Review cycle
Recheck before paying
Decision shortcut
Choose faster with these rules
Compare on output price, not input. Output is typically five to six times the input rate, so it dominates the bill on anything generative.
Check whether a rate is promotional. Gemini 3.7 Flash and the Batch and Flex discounts all expire on 31 December 2026.
Check long-context rates separately. GPT-5.6 doubles above 272K tokens, which can matter more than the headline rate on document work.
Prices moved twice in six weeks over July and August 2026. Verify current rates against the provider's own page before committing to a budget.
Deep comparison
What changes in real daily use
Quick answer
Gemini 3.7 Flash is the cheapest credible model at $0.75 input and $3.75 output per million tokens, but that rate is promotional through 31 December 2026. For pricing you can plan around, GPT-5.6 Terra at $2 and $12 and Claude Sonnet 5 at $2 and $10 are the stable mid-tier options - and Sonnet 5's rate was made permanent on 10 August 2026.
- Compare on output price, not input. Output is typically five to six times the input rate, so it dominates the bill on anything generative.
- Check whether a rate is promotional. Gemini 3.7 Flash and the Batch and Flex discounts all expire on 31 December 2026.
- Check long-context rates separately. GPT-5.6 doubles above 272K tokens, which can matter more than the headline rate on document work.
India buying checks
Test access and total cost before committing.
- Check the final card charge and applicable tax.
- Confirm employer or client data policy.
- Run the same real task in both products before choosing.
Pros
- Clear workflow fit for Cheapest tiers
- Clear workflow fit for Flagship tiers
- India-focused subscription and privacy checks
Cons
- Features and limits change frequently
- Both products still require human verification
Best for whom
Match the choice to the buyer
Bulk classification and tagging
GPT-5.6 Luna
At $0.20 and $1.20 after the 80% July cut, it changes what is economic to automate.
General production work
Claude Sonnet 5 or GPT-5.6 Terra
Both around $2 input, and Sonnet 5's rate is now permanent rather than introductory.
Cost-sensitive at any scale
Gemini 3.7 Flash
$0.75 and $3.75 undercuts everything credible - with a December expiry to plan around.
Hard reasoning
Gemini 3.1 Pro
$2 and $12 for benchmark-leading reasoning is better value than either flagship.
Alternatives
Other options worth considering
Batch and Flex processing
Where latency does not matter, Gemini Batch and Flex run at $0.375 and $1.875 through 31 December 2026 - roughly half the standard Flash-Lite rate.
Open-weight models
Worth evaluating if volume is high and data cannot leave your infrastructure. The cost moves from tokens to hosting and operations.
FAQ
What is the cheapest AI model in 2026?
GPT-5.6 Luna at $0.20 input and $1.20 output per million tokens, after OpenAI cut it 80% on 30 July 2026. For general-purpose work rather than bulk processing, Gemini 3.7 Flash at $0.75 and $3.75 is the cheapest credible option.
How do I estimate my monthly AI cost in rupees?
Estimate tokens rather than requests - roughly 750 words per thousand tokens - then multiply by the output rate, since output dominates. Convert at the rate your card issuer uses, not the mid-market rate, and add applicable tax. Most Indian buyers are billed in dollars with a forex markup on top.
Why did prices fall so sharply in 2026?
Competition at the low end. OpenAI cut Terra 20% and Luna 80% on 30 July, Anthropic made Sonnet 5's introductory rate permanent on 10 August rather than stepping it up, and Google launched 3.7 Flash on 14 August at half the previous Flash rate. The flagship tiers did not move.

