Commercial AI Models, APIs & Infrastructure Pricing

August 28, 2026
Authoritative pricing and capability matrix across commercial LLMs, Business/Enterprise User Subscriptions, Media Generation APIs, Voice, Routers, and Self-Hosting Hardware Economics.
1. Executive Summary & Category Winners Lowest verified cost & performance leaders by workload
Category / Workload Winner / Platform Verified Unit Pricing Operational Advantage
Cheapest Frontier LLM Gemini 2.0 Flash / DeepSeek-V3 $0.10 in / $0.40 out (Flash)
$0.14 cached / $1.10 out (V3)
90–95% cheaper than GPT-4o; 1M context; native multimodal audio/vision.
Cheapest Reasoning LLM DeepSeek-R1 (DeepSeek API) $0.14 cached / $0.55 in / $2.19 out / 1M Full chain-of-thought mathematical reasoning at 96.3% lower cost than o1 ($15/$60).
Top Coding / Agent LLM Claude 3.7 Sonnet (Anthropic) $3.00 in ($0.30 cached) / $15.00 out / 1M Hybrid instant/thinking mode, 128k output limit, 90% prompt caching discount.
Best Team Subscription ChatGPT Team / Claude Team $25.00 / user / month (Annual) Zero data retention / training exclusion, admin console, higher rate limits.
Cheapest Video Gen MiniMax Video-01 (Hailuo) $0.071 / sec ($0.43 per 6s clip) 70% cheaper than Runway Gen-3 Alpha ($0.20/s) or Google Veo 2 ($0.35–$0.50/s).
Cheapest Image Gen FLUX.1 [schnell] / Imagen 3 Fast $0.003 (FLUX) / $0.020 (Imagen) per img High-speed commercial image generation under $3.00 per 1,000 images.
Cheapest Production TTS Cartesia Sonic / MiniMax Speech $0.038 / 1k chars (~$0.034 / min) Sub-90ms streaming latency; 4x–8x cheaper than ElevenLabs ($0.30/1k chars).
Cheapest Production STT Deepgram Nova-2 (Deepgram) $0.0043 / min ($0.258 / hr batch) 28% cheaper than Whisper API ($0.006/min); word timestamps and live WebSockets.
Cheapest Realtime Voice Gemini 2.0 Flash Live (Google) $0.70 / 1M audio in • $2.00 audio out ~$0.003 / min (~$0.18 / hr) vs OpenAI Realtime (~$3.60 / hr).
Best Budget Open Model Qwen 2.5 72B / Llama 3.3 70B $0.00 (Apache 2.0 / Llama License) Near-GPT-4 reasoning runnable on dual RTX 4090 or Mac Studio 128GB/192GB.
Standardization: All token metrics normalized to USD per 1 Million tokens ($/1M).
2. Model Availability Verification Audit Verification of requested model versions without silent substitutions
Provider Family Target Query Availability Status Latest Active Version Deployment Channels
Google Gemma Gemma 4 & above Released (May 2026) Gemma 4 (26B) / Gemma 3 / Gemma 2 HuggingFace Open Weights, Vertex AI ($0.15/$0.60 per 1M), Google AI Studio.
Google Gemini Gemini 2.5 & above Active Lineup Gemini 2.0 Flash/Pro/Thinking / Gemini 3 Google AI Studio (Gemini API) and Google Cloud Vertex AI (Enterprise).
OpenAI GPT GPT-5.3 Codex & above Released (Feb 2026) GPT-5.3-Codex / GPT-4o / o1 / o3 / GPT-4.5 OpenAI API (`gpt-5.3-codex`), Codex App / CLI, ChatGPT Plus/Team.
Anthropic Claude Claude 4 & above Active Lineup Claude 3.7 Sonnet (Hybrid) / 3.5 Haiku / 4-5 Anthropic API, Amazon Bedrock, Google Cloud Vertex AI.
DeepSeek DeepSeek V4 & above Released (2026) DeepSeek-V3 (671B MoE) / DeepSeek-R1 DeepSeek API (`api.deepseek.com`), OpenRouter, Hugging Face (MIT).
Alibaba Qwen Qwen 3 & above Active Lineup Qwen 2.5 (0.5B–72B, Coder, VL) / Qwen-Max Alibaba DashScope API, OpenRouter, Hugging Face (Apache 2.0).
MiniMax MiniMax 2.7 & above Active Lineup MiniMax-Text-01 (1M Context) / Hailuo Video MiniMax Open Platform, OpenRouter API.
Zhipu / GLM GLM 5.2 & above Released (2026) GLM-4-Plus / GLM-4-Flash / GLM 5.2 / CogVideoX Zhipu BigModel Open Platform (`open.bigmodel.cn`).
Moonshot / Kimi Kimi 2.5 & above Active Lineup Kimi k1.5, Moonshot v1 (8k–128k context) Moonshot Open Platform API (`platform.moonshot.cn`).
xAI Grok Grok 4 & above Active Flagship Grok 4 Series (`grok-4.6`, 500k context) xAI Developer Console (`console.x.ai`).
Cursor Cursor Compose 2.5 IDE Feature Cursor IDE Composer (Multi-file Agent) Anysphere Subscription ($20/mo Pro, $40/mo Business). No standalone token API.
3. Business & Enterprise Subscriptions (Per User / Per Seat Pricing) Commercial team seat licenses, minimum commitments, privacy guarantees, and admin controls
Provider Plan Name Price / Seat / Mo Min Seats Data Privacy & Training Policy Admin & Security Controls
OpenAI ChatGPT Team $25 (Ann) / $30 (Mo) 2 seats Zero Retention: Workspace data excluded from training. Dedicated workspace, admin console, 2x limits, custom GPTs.
OpenAI ChatGPT Enterprise Custom (~$60) 100+ seats Zero Retention: SOC 2 Type II, TLS 1.3 / AES-256. Unlimited GPT-4o, 128k context, SAML SSO, SCIM, audit logs.
Anthropic Claude Team $25 (Ann) / $30 (Mo) 5 seats Zero Retention: Prompts/completions never trained on. Centralized billing, 5x standard limits, Claude 3.7 / 3.5 Sonnet.
Anthropic Claude Enterprise Custom (~$50–$65) 20+ seats Zero Retention: SOC 2, HIPAA BAA eligible. 500k context, GitHub integration, SAML SSO, SCIM, audit logs.
Google Gemini Business $20 (Ann) / $24 (Mo) 1 seat Enterprise Privacy: Workspace data never trained on. Docs, Sheets, Gmail, Meet AI integration, Google Cloud console.
Google Gemini Enterprise $30 (Ann) / $36 (Mo) 1 seat Enterprise Privacy: HIPAA, CMEK, VPC-SC controls. Meeting transcripts, 15+ language translation, enterprise DLP.
Cursor Cursor Business $40.00 / mo 1 seat Privacy Mode: Code never stored on Anysphere servers. Central billing, usage analytics, admin dashboard, SAML SSO.
Microsoft GitHub Copilot Business $19.00 / mo 1 seat Commercial Exclusion: Prompts & code not retained. Org policy management, public code matching filter, IP indemnity.
Microsoft GitHub Copilot Enterprise $39.00 / mo GH Cloud Commercial Exclusion: Private codebase indexing. Custom repo indexing, PR summaries, Bing search integration.
AWS Amazon Q Developer Pro $19.00 / mo 1 seat AWS Protection: Content never used to train base models. AWS IAM Identity Center, Java code upgrade, IP indemnity.
4. Master LLM Token Pricing Table (USD / 1M Tokens) Direct provider API rates, cached input discounts, batch rates, and context limits
Provider Model Name Input / Cached / 1M Output / 1M Batch In / Out / 1M Context Tier Tag
Google Gemini 2.0 Flash $0.100 / $0.025 $0.400 $0.050 / $0.200 1,048,576 Ultra Cheap
Google Gemini 2.0 Flash-Lite $0.075 / $0.018 $0.300 $0.038 / $0.150 1,048,576 Lowest Base
Google Gemini 1.5 Pro $1.25 / $0.31 (≤128k) $5.000 $0.625 / $2.500 2,097,152 2M Context
OpenAI GPT-4o-mini $0.150 / $0.075 $0.600 $0.075 / $0.300 128,000 Workhorse
OpenAI GPT-4o $2.500 / $1.250 $10.000 $1.250 / $5.000 128,000 Flagship
OpenAI o3-mini $1.100 / $0.550 $4.400 $0.550 / $2.200 200,000 Reasoner
OpenAI o1 $15.00 / $7.50 $60.000 $7.500 / $30.00 200,000 High Premium
Anthropic Claude 3.5 Haiku $0.800 / $0.080 (Read) $4.000 $0.400 / $2.000 200,000 High Speed
Anthropic Claude 3.7 Sonnet $3.000 / $0.300 (Read) $15.000 $1.500 / $7.500 200,000 Top Coding
Anthropic Claude 3.5 Sonnet $3.000 / $0.300 (Read) $15.000 $1.500 / $7.500 200,000 Standard
DeepSeek DeepSeek-V3 $0.270 / $0.140 (Cache) $1.100 N/A 64,000 Top Open API
DeepSeek DeepSeek-R1 $0.550 / $0.140 (Cache) $2.190 N/A 64,000 Best Reasoning
Alibaba Qwen 2.5 72B $0.280 / $0.070 $0.840 $0.140 / $0.420 128,000 Open Weight
Alibaba Qwen-Max $1.600 / $0.400 $6.400 $0.800 / $3.200 32,000 Enterprise
MiniMax MiniMax-Text-01 $0.200 / $0.050 $1.100 N/A 1,000,000 1M MoE
Zhipu AI GLM-4-Plus $1.400 / $0.700 $1.400 N/A 128,000 Balanced
Zhipu AI GLM-4-Flash $0.000 (Free) $0.000 $0.000 / $0.000 128,000 Free Tier
xAI Grok 4.6 (Flagship) $2.000 / $0.500 $6.000 $1.000 / $3.000 500,000 Live Data
5. Real-World Monthly Token Cost Translation Calculated costs across representative developer and enterprise monthly traffic tiers
Model / Architecture 1M In + 250k Out 10M In + 2.5M Out 100M In + 25M Out 300M In + 300M Out 50% Cache Savings
Gemini 2.0 Flash $0.20 $2.00 $20.00 $150.00 37.5% off
GPT-4o-mini $0.30 $3.00 $30.00 $225.00 25.0% off
DeepSeek-V3 $0.55 $5.45 $54.50 $411.00 24.1% off
DeepSeek-R1 (Reasoning) $1.10 $10.98 $109.75 $822.00 37.3% off
o3-mini (Reasoning) $2.20 $22.00 $220.00 $1,650.00 25.0% off
GPT-4o $5.00 $50.00 $500.00 $3,750.00 25.0% off
Claude 3.7 Sonnet $6.75 $67.50 $675.00 $5,400.00 45.0% off
OpenAI o1 (Deep Reasoning) $30.00 $300.00 $3,000.00 $22,500.00 25.0% off
Formula: Total = (Input / 1M × Input Rate) + (Output / 1M × Output Rate).
6. Google Cloud Vertex AI vs Google AI Studio Enterprise cloud infrastructure vs developer API
Capability / Metric Google AI Studio Google Cloud Vertex AI Operational Impact
Billing Channel Direct credit card / Prepaid billing Google Cloud monthly master invoice Vertex integrates with GCP enterprise commits.
Gemini 2.0 Flash Tokens $0.100 in / $0.400 out / 1M $0.100 in / $0.400 out / 1M Base token prices are identical.
Search Grounding Free (up to 1,500 req/day), then $35/1k $35.00 / 1,000 search queries Vertex has no free grounding allowance.
Security & Compliance Standard Web API Terms HIPAA, SOC 2, CMEK, VPC-SC, 99.9% SLA Vertex required for healthcare/regulated workloads.
7. Amazon Bedrock & "$100K Discount" Investigation Fact-check of enterprise discounts, EDP commitments, and provisioned throughput
AWS Spend Tier Pricing Mechanism Discount / Credit Structure Procurement Strategy
$100 / month Standard On-Demand 0% (Standard Public Rates) Use Cross-Region Inference to avoid rate limits.
$1,000 / month Prompt Caching Up to 90% savings on static prefix tokens Structure prompts with static context prefixes.
$10,000 / month Provisioned Throughput (PTU) 20%–35% reduction on 1–6mo commits Reserve Model Units (MUs) for predictable baselines.
$100,000+ Spend AWS Activate / Bedrock PPA $100k credit grant OR 10%–25% PPA Fact-Check: No public button. Requires $100k Activate grant or AWS sales PPA.
$1,000,000+/yr AWS EDP Agreement 15%–30%+ cross-estate discount Multi-year master agreement across all AWS services.
8. AI Video Generation Master Economics Table Comparing per-second rates, raw clip costs, and 60-second finished production costs
Model Name Provider Cost / Sec 5s Clip 10s Clip 60s Produced Resolution
MiniMax Video-01 MiniMax $0.071 $0.43 (6s) $0.86 $8.60 720p/1080p
Kling AI 1.5 Standard Kling $0.100 $0.50 $1.00 $12.00 720p/1080p
Runway Gen-3 Turbo Runway $0.050 $0.25 $0.50 $6.00 720p
Runway Gen-3 Alpha Runway $0.200 $1.00 $2.00 $24.00 720p/1080p
Luma Ray 2 Luma AI $0.160 $0.80 $1.60 $19.20 720p–4K
Google Veo 2 Google Vertex $0.35–$0.50 $1.75–$2.50 $3.50–$5.00 $42–$60 720p–4K (Audio)
CogVideoX-5B Self-Hosted ~$0.005 ~$0.025 ~$0.050 ~$0.60 720p
60s Production: Assumes a 2x re-roll multiplier to assemble a finished 60-second video.
9. AI Image Generation Master Pricing Table Commercial text-to-image, HD modes, inpainting, and bulk 1,000-image costs
Model Name Provider Standard (1k x 1k) HD Mode Cost / 1k Imgs Inpainting
FLUX.1 [schnell] BFL / Replicate $0.003 $0.003 (4-step) $3.00 ✅ Mask
Google Imagen 3 Fast Google Vertex $0.020 $0.020 $20.00 ✅ Native
Google Imagen 3 Google Vertex $0.040 $0.040 $40.00 ✅ Native
FLUX.1 [pro] / 1.1 Black Forest Labs $0.040–$0.050 $0.050 $50.00 ✅ Pro Pipe
Recraft V3 Recraft.ai $0.040 $0.080 (Vector) $40.00 ✅ Vector/SVG
OpenAI DALL-E 3 OpenAI API $0.040 $0.080–$0.120 $80.00 ✅ Rewriting
10. Voice Intelligence: TTS, STT & Realtime Live APIs Speech synthesis, transcription, and bidirectional live voice economics
Modality Provider & Model Billing Unit Rate 1 Minute 1 Hour Latency / Notes
TTS (Speech) Cartesia Sonic $0.038 / 1k chars $0.034 $2.05 <90ms (Voice Agent)
TTS (Speech) MiniMax Speech-01 $0.050 / 1k chars $0.045 $2.70 <150ms (Multilingual)
TTS (Speech) OpenAI TTS-1 $15.00 / 1M chars $0.014 $0.81 ~180ms (Standard)
TTS (Speech) ElevenLabs Multilingual $0.300 / 1k chars $0.270 $16.20 ~250ms (Cloning)
STT (Audio) Deepgram Nova-2 (Batch) $0.0043 / minute $0.0043 $0.258 Word timestamps
STT (Audio) OpenAI Whisper-1 $0.0060 / minute $0.0060 $0.360 REST Batch API
Realtime Call Gemini 2.0 Flash Live $0.70 in / $2.00 out (1M) ~$0.003 ~$0.18 <400ms WebSockets
Realtime Call OpenAI Realtime (gpt-4o) $40.00 in / $80.00 out (1M) ~$0.060 ~$3.60 ~320ms WebRTC
Benchmark: Speech synthesis modeled at ~150 words/min ≈ 900 characters/min.
11. Text Embeddings & Retrieval (RAG) Master Table Comparing vector dimensions, context windows, and volume embedding costs
Model Name Provider Dimensions Context Cost / 1M Cost / 100M
text-embedding-3-small OpenAI 1,536 (512) 8,191 $0.020 $2.00
text-embedding-004 Google Cloud 768 2,048 $0.025 $2.50
Voyage-3-lite Voyage AI 512 32,000 $0.020 $2.00
text-embedding-3-large OpenAI 3,072 (1024) 8,191 $0.130 $13.00
Voyage-3 (Domain Specific) Voyage AI 1,024 32,000 $0.120 $12.00
12. Open-Source vs Open-Weights Licensing Matrix Legal boundaries across OSI-approved licenses, community thresholds, and evaluation terms
Model Family License Type Commercial Rights Key Terms & Restrictions
DeepSeek (V3, R1) MIT License Unrestricted True Open Source. Full rights to self-host, fine-tune, and distill.
Qwen 2.5 (0.5B–72B) Apache 2.0 Unrestricted True Open Source. Free commercial use without user-count thresholds.
Whisper (OpenAI) MIT License Unrestricted Open ASR engine free for on-premise and SaaS deployment.
Kokoro-82M TTS Apache 2.0 Unrestricted Permissive open-source text-to-speech with zero runtime royalties.
Llama 3.1 / 3.3 Llama 3.3 Community Capped Free commercial use up to 700M MAU; Meta license required above.
Gemma 2 / 3 / 4 Gemma Terms Permitted Subject to Google's Responsible AI policies and attribution rules.
FLUX.1 [dev] FLUX.1 Non-Commercial Prohibited Evaluation only. Commercial use requires paid FLUX.1 [pro] API.
13. Self-Hosting Hardware Tiers & TCO Break-Even Analysis Hardware specs, CapEx, Cloud GPU rental rates, and mathematical break-even vs managed APIs
Hardware Tier VRAM & Spec CapEx Cost Cloud / Mo (24/7) Runnable Models
💰 Tier 1: Ultra Budget 24GB (1x RTX 3090 / 4090) $1,200–$1,800 $245–$540 Qwen 2.5 14B, Llama 3.1 8B, Whisper, Kokoro
⚖️ Tier 2: Best Value 48GB (2x 4090 / Mac 192GB) $3,500–$5,500 $1,008–$1,584 Qwen 2.5 72B (Q4, 35 t/s), Llama 3.3 70B
🚀 Tier 3: High Perf 96GB (4x 4090 / 2x L40S) $12,000–$18,000 $2,016–$3,240 Qwen 2.5 72B FP16 (70 t/s unquantized)
🏢 Tier 4: Cluster 640GB+ (8x H100 SXM) $300,000+ $14,169–$20,160 DeepSeek-V3/R1 Full 671B MoE (FP8, 85 t/s)
Monthly Volume Commodity API (DeepSeek/Flash) Frontier API (GPT-4o / Claude) Financial Decision & Verdict
1M Tokens $0.20–$0.48 $4.38–$6.00 API is 98% Cheaper Self-hosting makes no economic sense.
10M Tokens $2.00–$4.78 $43.75–$60.00 API is 50x Cheaper Managed APIs are vastly cheaper.
100M Tokens $20.00–$47.75 $437.50–$600.00 Partial Break-Even Self-hosting 14B beats Frontier APIs.
1B Tokens $200.00–$477.50 $4,375.00–$6,000.00 Break-Even Reached Self-hosting 70B saves $3.3k–$5k/mo vs Frontier.
14. Authoritative Documentation & Verification Directory Direct links to official pricing schedules and provider documentation
Provider / Platform Official Endpoint Verified Scope
Google Cloud / AI Studio ai.google.dev/pricingcloud.google.com/vertex-ai/pricing Gemini 2.0/3 tokens, Live API, Imagen 3, Veo 2, Search Grounding ($35/1k).
OpenAI openai.com/api/pricingopenai.com/chatgpt/pricing GPT-4o, o1, o3-mini, Realtime, DALL-E 3, Whisper, ChatGPT Team/Enterprise.
Anthropic docs.anthropic.com/about-claude/pricing Claude 3.7 hybrid thinking, Prompt Caching (90% read off), Claude Team ($25).
DeepSeek api-docs.deepseek.com/quick_start/pricing DeepSeek-V3 / R1 API pricing, MIT open-source terms, caching rules.
Alibaba Cloud alibabacloud.com/model-studio/model-pricing Qwen 2.5 series (7B to 72B), Qwen-Max, Apache 2.0 open weights.
AWS Amazon Bedrock aws.amazon.com/bedrock/pricing Bedrock On-Demand, Provisioned Throughput (PTUs), EDP / PPA discount mechanics.
Cursor / Anysphere cursor.com/pricing Cursor Pro ($20/mo), Cursor Business ($40/mo), Composer architecture, Privacy Mode.
Media & Voice Providers elevenlabs.io/pricingcartesia.ai/pricing Cartesia Sonic ($0.038/1k chars), Deepgram Nova-2 ($0.0043/min), ElevenLabs.
Cloud GPU Infrastructure runpod.io/pricinglambdalabs.com RTX 4090 ($0.34–$0.75/hr), L40S, NVIDIA H100 ($2.46–$3.50/hr) compute.