1. Executive Summary & Category Winners
Lowest verified cost & performance leaders by workload
| Category / Workload |
Winner / Platform |
Verified Unit Pricing |
Operational Advantage |
| Cheapest Frontier LLM |
Gemini 2.0 Flash / DeepSeek-V3 |
$0.10 in / $0.40 out (Flash) $0.14 cached / $1.10 out (V3) |
90–95% cheaper than GPT-4o; 1M context; native multimodal audio/vision. |
| Cheapest Reasoning LLM |
DeepSeek-R1 (DeepSeek API) |
$0.14 cached / $0.55 in / $2.19 out / 1M |
Full chain-of-thought mathematical reasoning at 96.3% lower cost than o1 ($15/$60). |
| Top Coding / Agent LLM |
Claude 3.7 Sonnet (Anthropic) |
$3.00 in ($0.30 cached) / $15.00 out / 1M |
Hybrid instant/thinking mode, 128k output limit, 90% prompt caching discount. |
| Best Team Subscription |
ChatGPT Team / Claude Team |
$25.00 / user / month (Annual) |
Zero data retention / training exclusion, admin console, higher rate limits. |
| Cheapest Video Gen |
MiniMax Video-01 (Hailuo) |
$0.071 / sec ($0.43 per 6s clip) |
70% cheaper than Runway Gen-3 Alpha ($0.20/s) or Google Veo 2 ($0.35–$0.50/s). |
| Cheapest Image Gen |
FLUX.1 [schnell] / Imagen 3 Fast |
$0.003 (FLUX) / $0.020 (Imagen) per img |
High-speed commercial image generation under $3.00 per 1,000 images. |
| Cheapest Production TTS |
Cartesia Sonic / MiniMax Speech |
$0.038 / 1k chars (~$0.034 / min) |
Sub-90ms streaming latency; 4x–8x cheaper than ElevenLabs ($0.30/1k chars). |
| Cheapest Production STT |
Deepgram Nova-2 (Deepgram) |
$0.0043 / min ($0.258 / hr batch) |
28% cheaper than Whisper API ($0.006/min); word timestamps and live WebSockets. |
| Cheapest Realtime Voice |
Gemini 2.0 Flash Live (Google) |
$0.70 / 1M audio in • $2.00 audio out |
~$0.003 / min (~$0.18 / hr) vs OpenAI Realtime (~$3.60 / hr). |
| Best Budget Open Model |
Qwen 2.5 72B / Llama 3.3 70B |
$0.00 (Apache 2.0 / Llama License) |
Near-GPT-4 reasoning runnable on dual RTX 4090 or Mac Studio 128GB/192GB. |
Standardization: All token metrics normalized to USD per 1 Million tokens ($/1M).
2. Model Availability Verification Audit
Verification of requested model versions without silent substitutions
| Provider Family |
Target Query |
Availability Status |
Latest Active Version |
Deployment Channels |
| Google Gemma |
Gemma 4 & above |
Released (May 2026) |
Gemma 4 (26B) / Gemma 3 / Gemma 2 |
HuggingFace Open Weights, Vertex AI ($0.15/$0.60 per 1M), Google AI Studio. |
| Google Gemini |
Gemini 2.5 & above |
Active Lineup |
Gemini 2.0 Flash/Pro/Thinking / Gemini 3 |
Google AI Studio (Gemini API) and Google Cloud Vertex AI (Enterprise). |
| OpenAI GPT |
GPT-5.3 Codex & above |
Released (Feb 2026) |
GPT-5.3-Codex / GPT-4o / o1 / o3 / GPT-4.5 |
OpenAI API (`gpt-5.3-codex`), Codex App / CLI, ChatGPT Plus/Team. |
| Anthropic Claude |
Claude 4 & above |
Active Lineup |
Claude 3.7 Sonnet (Hybrid) / 3.5 Haiku / 4-5 |
Anthropic API, Amazon Bedrock, Google Cloud Vertex AI. |
| DeepSeek |
DeepSeek V4 & above |
Released (2026) |
DeepSeek-V3 (671B MoE) / DeepSeek-R1 |
DeepSeek API (`api.deepseek.com`), OpenRouter, Hugging Face (MIT). |
| Alibaba Qwen |
Qwen 3 & above |
Active Lineup |
Qwen 2.5 (0.5B–72B, Coder, VL) / Qwen-Max |
Alibaba DashScope API, OpenRouter, Hugging Face (Apache 2.0). |
| MiniMax |
MiniMax 2.7 & above |
Active Lineup |
MiniMax-Text-01 (1M Context) / Hailuo Video |
MiniMax Open Platform, OpenRouter API. |
| Zhipu / GLM |
GLM 5.2 & above |
Released (2026) |
GLM-4-Plus / GLM-4-Flash / GLM 5.2 / CogVideoX |
Zhipu BigModel Open Platform (`open.bigmodel.cn`). |
| Moonshot / Kimi |
Kimi 2.5 & above |
Active Lineup |
Kimi k1.5, Moonshot v1 (8k–128k context) |
Moonshot Open Platform API (`platform.moonshot.cn`). |
| xAI Grok |
Grok 4 & above |
Active Flagship |
Grok 4 Series (`grok-4.6`, 500k context) |
xAI Developer Console (`console.x.ai`). |
| Cursor |
Cursor Compose 2.5 |
IDE Feature |
Cursor IDE Composer (Multi-file Agent) |
Anysphere Subscription ($20/mo Pro, $40/mo Business). No standalone token API. |
3. Business & Enterprise Subscriptions (Per User / Per Seat Pricing)
Commercial team seat licenses, minimum commitments, privacy guarantees, and admin controls
| Provider |
Plan Name |
Price / Seat / Mo |
Min Seats |
Data Privacy & Training Policy |
Admin & Security Controls |
| OpenAI |
ChatGPT Team |
$25 (Ann) / $30 (Mo) |
2 seats |
Zero Retention: Workspace data excluded from training. |
Dedicated workspace, admin console, 2x limits, custom GPTs. |
| OpenAI |
ChatGPT Enterprise |
Custom (~$60) |
100+ seats |
Zero Retention: SOC 2 Type II, TLS 1.3 / AES-256. |
Unlimited GPT-4o, 128k context, SAML SSO, SCIM, audit logs. |
| Anthropic |
Claude Team |
$25 (Ann) / $30 (Mo) |
5 seats |
Zero Retention: Prompts/completions never trained on. |
Centralized billing, 5x standard limits, Claude 3.7 / 3.5 Sonnet. |
| Anthropic |
Claude Enterprise |
Custom (~$50–$65) |
20+ seats |
Zero Retention: SOC 2, HIPAA BAA eligible. |
500k context, GitHub integration, SAML SSO, SCIM, audit logs. |
| Google |
Gemini Business |
$20 (Ann) / $24 (Mo) |
1 seat |
Enterprise Privacy: Workspace data never trained on. |
Docs, Sheets, Gmail, Meet AI integration, Google Cloud console. |
| Google |
Gemini Enterprise |
$30 (Ann) / $36 (Mo) |
1 seat |
Enterprise Privacy: HIPAA, CMEK, VPC-SC controls. |
Meeting transcripts, 15+ language translation, enterprise DLP. |
| Cursor |
Cursor Business |
$40.00 / mo |
1 seat |
Privacy Mode: Code never stored on Anysphere servers. |
Central billing, usage analytics, admin dashboard, SAML SSO. |
| Microsoft |
GitHub Copilot Business |
$19.00 / mo |
1 seat |
Commercial Exclusion: Prompts & code not retained. |
Org policy management, public code matching filter, IP indemnity. |
| Microsoft |
GitHub Copilot Enterprise |
$39.00 / mo |
GH Cloud |
Commercial Exclusion: Private codebase indexing. |
Custom repo indexing, PR summaries, Bing search integration. |
| AWS |
Amazon Q Developer Pro |
$19.00 / mo |
1 seat |
AWS Protection: Content never used to train base models. |
AWS IAM Identity Center, Java code upgrade, IP indemnity. |
4. Master LLM Token Pricing Table (USD / 1M Tokens)
Direct provider API rates, cached input discounts, batch rates, and context limits
| Provider |
Model Name |
Input / Cached / 1M |
Output / 1M |
Batch In / Out / 1M |
Context |
Tier Tag |
| Google |
Gemini 2.0 Flash |
$0.100 / $0.025 |
$0.400 |
$0.050 / $0.200 |
1,048,576 |
Ultra Cheap |
| Google |
Gemini 2.0 Flash-Lite |
$0.075 / $0.018 |
$0.300 |
$0.038 / $0.150 |
1,048,576 |
Lowest Base |
| Google |
Gemini 1.5 Pro |
$1.25 / $0.31 (≤128k) |
$5.000 |
$0.625 / $2.500 |
2,097,152 |
2M Context |
| OpenAI |
GPT-4o-mini |
$0.150 / $0.075 |
$0.600 |
$0.075 / $0.300 |
128,000 |
Workhorse |
| OpenAI |
GPT-4o |
$2.500 / $1.250 |
$10.000 |
$1.250 / $5.000 |
128,000 |
Flagship |
| OpenAI |
o3-mini |
$1.100 / $0.550 |
$4.400 |
$0.550 / $2.200 |
200,000 |
Reasoner |
| OpenAI |
o1 |
$15.00 / $7.50 |
$60.000 |
$7.500 / $30.00 |
200,000 |
High Premium |
| Anthropic |
Claude 3.5 Haiku |
$0.800 / $0.080 (Read) |
$4.000 |
$0.400 / $2.000 |
200,000 |
High Speed |
| Anthropic |
Claude 3.7 Sonnet |
$3.000 / $0.300 (Read) |
$15.000 |
$1.500 / $7.500 |
200,000 |
Top Coding |
| Anthropic |
Claude 3.5 Sonnet |
$3.000 / $0.300 (Read) |
$15.000 |
$1.500 / $7.500 |
200,000 |
Standard |
| DeepSeek |
DeepSeek-V3 |
$0.270 / $0.140 (Cache) |
$1.100 |
N/A |
64,000 |
Top Open API |
| DeepSeek |
DeepSeek-R1 |
$0.550 / $0.140 (Cache) |
$2.190 |
N/A |
64,000 |
Best Reasoning |
| Alibaba |
Qwen 2.5 72B |
$0.280 / $0.070 |
$0.840 |
$0.140 / $0.420 |
128,000 |
Open Weight |
| Alibaba |
Qwen-Max |
$1.600 / $0.400 |
$6.400 |
$0.800 / $3.200 |
32,000 |
Enterprise |
| MiniMax |
MiniMax-Text-01 |
$0.200 / $0.050 |
$1.100 |
N/A |
1,000,000 |
1M MoE |
| Zhipu AI |
GLM-4-Plus |
$1.400 / $0.700 |
$1.400 |
N/A |
128,000 |
Balanced |
| Zhipu AI |
GLM-4-Flash |
$0.000 (Free) |
$0.000 |
$0.000 / $0.000 |
128,000 |
Free Tier |
| xAI |
Grok 4.6 (Flagship) |
$2.000 / $0.500 |
$6.000 |
$1.000 / $3.000 |
500,000 |
Live Data |
5. Real-World Monthly Token Cost Translation
Calculated costs across representative developer and enterprise monthly traffic tiers
| Model / Architecture |
1M In + 250k Out |
10M In + 2.5M Out |
100M In + 25M Out |
300M In + 300M Out |
50% Cache Savings |
| Gemini 2.0 Flash |
$0.20 |
$2.00 |
$20.00 |
$150.00 |
37.5% off |
| GPT-4o-mini |
$0.30 |
$3.00 |
$30.00 |
$225.00 |
25.0% off |
| DeepSeek-V3 |
$0.55 |
$5.45 |
$54.50 |
$411.00 |
24.1% off |
| DeepSeek-R1 (Reasoning) |
$1.10 |
$10.98 |
$109.75 |
$822.00 |
37.3% off |
| o3-mini (Reasoning) |
$2.20 |
$22.00 |
$220.00 |
$1,650.00 |
25.0% off |
| GPT-4o |
$5.00 |
$50.00 |
$500.00 |
$3,750.00 |
25.0% off |
| Claude 3.7 Sonnet |
$6.75 |
$67.50 |
$675.00 |
$5,400.00 |
45.0% off |
| OpenAI o1 (Deep Reasoning) |
$30.00 |
$300.00 |
$3,000.00 |
$22,500.00 |
25.0% off |
Formula: Total = (Input / 1M × Input Rate) + (Output / 1M × Output Rate).
6. Google Cloud Vertex AI vs Google AI Studio
Enterprise cloud infrastructure vs developer API
| Capability / Metric |
Google AI Studio |
Google Cloud Vertex AI |
Operational Impact |
| Billing Channel |
Direct credit card / Prepaid billing |
Google Cloud monthly master invoice |
Vertex integrates with GCP enterprise commits. |
| Gemini 2.0 Flash Tokens |
$0.100 in / $0.400 out / 1M |
$0.100 in / $0.400 out / 1M |
Base token prices are identical. |
| Search Grounding |
Free (up to 1,500 req/day), then $35/1k |
$35.00 / 1,000 search queries |
Vertex has no free grounding allowance. |
| Security & Compliance |
Standard Web API Terms |
HIPAA, SOC 2, CMEK, VPC-SC, 99.9% SLA |
Vertex required for healthcare/regulated workloads. |
7. Amazon Bedrock & "$100K Discount" Investigation
Fact-check of enterprise discounts, EDP commitments, and provisioned throughput
| AWS Spend Tier |
Pricing Mechanism |
Discount / Credit Structure |
Procurement Strategy |
| $100 / month |
Standard On-Demand |
0% (Standard Public Rates) |
Use Cross-Region Inference to avoid rate limits. |
| $1,000 / month |
Prompt Caching |
Up to 90% savings on static prefix tokens |
Structure prompts with static context prefixes. |
| $10,000 / month |
Provisioned Throughput (PTU) |
20%–35% reduction on 1–6mo commits |
Reserve Model Units (MUs) for predictable baselines. |
| $100,000+ Spend |
AWS Activate / Bedrock PPA |
$100k credit grant OR 10%–25% PPA |
Fact-Check: No public button. Requires $100k Activate grant or AWS sales PPA. |
| $1,000,000+/yr |
AWS EDP Agreement |
15%–30%+ cross-estate discount |
Multi-year master agreement across all AWS services. |
8. AI Video Generation Master Economics Table
Comparing per-second rates, raw clip costs, and 60-second finished production costs
| Model Name |
Provider |
Cost / Sec |
5s Clip |
10s Clip |
60s Produced |
Resolution |
| MiniMax Video-01 |
MiniMax |
$0.071 |
$0.43 (6s) |
$0.86 |
$8.60 |
720p/1080p |
| Kling AI 1.5 Standard |
Kling |
$0.100 |
$0.50 |
$1.00 |
$12.00 |
720p/1080p |
| Runway Gen-3 Turbo |
Runway |
$0.050 |
$0.25 |
$0.50 |
$6.00 |
720p |
| Runway Gen-3 Alpha |
Runway |
$0.200 |
$1.00 |
$2.00 |
$24.00 |
720p/1080p |
| Luma Ray 2 |
Luma AI |
$0.160 |
$0.80 |
$1.60 |
$19.20 |
720p–4K |
| Google Veo 2 |
Google Vertex |
$0.35–$0.50 |
$1.75–$2.50 |
$3.50–$5.00 |
$42–$60 |
720p–4K (Audio) |
| CogVideoX-5B |
Self-Hosted |
~$0.005 |
~$0.025 |
~$0.050 |
~$0.60 |
720p |
60s Production: Assumes a 2x re-roll multiplier to assemble a finished 60-second video.
9. AI Image Generation Master Pricing Table
Commercial text-to-image, HD modes, inpainting, and bulk 1,000-image costs
| Model Name |
Provider |
Standard (1k x 1k) |
HD Mode |
Cost / 1k Imgs |
Inpainting |
| FLUX.1 [schnell] |
BFL / Replicate |
$0.003 |
$0.003 (4-step) |
$3.00 |
✅ Mask |
| Google Imagen 3 Fast |
Google Vertex |
$0.020 |
$0.020 |
$20.00 |
✅ Native |
| Google Imagen 3 |
Google Vertex |
$0.040 |
$0.040 |
$40.00 |
✅ Native |
| FLUX.1 [pro] / 1.1 |
Black Forest Labs |
$0.040–$0.050 |
$0.050 |
$50.00 |
✅ Pro Pipe |
| Recraft V3 |
Recraft.ai |
$0.040 |
$0.080 (Vector) |
$40.00 |
✅ Vector/SVG |
| OpenAI DALL-E 3 |
OpenAI API |
$0.040 |
$0.080–$0.120 |
$80.00 |
✅ Rewriting |
10. Voice Intelligence: TTS, STT & Realtime Live APIs
Speech synthesis, transcription, and bidirectional live voice economics
| Modality |
Provider & Model |
Billing Unit Rate |
1 Minute |
1 Hour |
Latency / Notes |
| TTS (Speech) |
Cartesia Sonic |
$0.038 / 1k chars |
$0.034 |
$2.05 |
<90ms (Voice Agent) |
| TTS (Speech) |
MiniMax Speech-01 |
$0.050 / 1k chars |
$0.045 |
$2.70 |
<150ms (Multilingual) |
| TTS (Speech) |
OpenAI TTS-1 |
$15.00 / 1M chars |
$0.014 |
$0.81 |
~180ms (Standard) |
| TTS (Speech) |
ElevenLabs Multilingual |
$0.300 / 1k chars |
$0.270 |
$16.20 |
~250ms (Cloning) |
| STT (Audio) |
Deepgram Nova-2 (Batch) |
$0.0043 / minute |
$0.0043 |
$0.258 |
Word timestamps |
| STT (Audio) |
OpenAI Whisper-1 |
$0.0060 / minute |
$0.0060 |
$0.360 |
REST Batch API |
| Realtime Call |
Gemini 2.0 Flash Live |
$0.70 in / $2.00 out (1M) |
~$0.003 |
~$0.18 |
<400ms WebSockets |
| Realtime Call |
OpenAI Realtime (gpt-4o) |
$40.00 in / $80.00 out (1M) |
~$0.060 |
~$3.60 |
~320ms WebRTC |
Benchmark: Speech synthesis modeled at ~150 words/min ≈ 900 characters/min.
11. Text Embeddings & Retrieval (RAG) Master Table
Comparing vector dimensions, context windows, and volume embedding costs
| Model Name |
Provider |
Dimensions |
Context |
Cost / 1M |
Cost / 100M |
| text-embedding-3-small |
OpenAI |
1,536 (512) |
8,191 |
$0.020 |
$2.00 |
| text-embedding-004 |
Google Cloud |
768 |
2,048 |
$0.025 |
$2.50 |
| Voyage-3-lite |
Voyage AI |
512 |
32,000 |
$0.020 |
$2.00 |
| text-embedding-3-large |
OpenAI |
3,072 (1024) |
8,191 |
$0.130 |
$13.00 |
| Voyage-3 (Domain Specific) |
Voyage AI |
1,024 |
32,000 |
$0.120 |
$12.00 |
12. Open-Source vs Open-Weights Licensing Matrix
Legal boundaries across OSI-approved licenses, community thresholds, and evaluation terms
| Model Family |
License Type |
Commercial Rights |
Key Terms & Restrictions |
| DeepSeek (V3, R1) |
MIT License |
Unrestricted |
True Open Source. Full rights to self-host, fine-tune, and distill. |
| Qwen 2.5 (0.5B–72B) |
Apache 2.0 |
Unrestricted |
True Open Source. Free commercial use without user-count thresholds. |
| Whisper (OpenAI) |
MIT License |
Unrestricted |
Open ASR engine free for on-premise and SaaS deployment. |
| Kokoro-82M TTS |
Apache 2.0 |
Unrestricted |
Permissive open-source text-to-speech with zero runtime royalties. |
| Llama 3.1 / 3.3 |
Llama 3.3 Community |
Capped |
Free commercial use up to 700M MAU; Meta license required above. |
| Gemma 2 / 3 / 4 |
Gemma Terms |
Permitted |
Subject to Google's Responsible AI policies and attribution rules. |
| FLUX.1 [dev] |
FLUX.1 Non-Commercial |
Prohibited |
Evaluation only. Commercial use requires paid FLUX.1 [pro] API. |
13. Self-Hosting Hardware Tiers & TCO Break-Even Analysis
Hardware specs, CapEx, Cloud GPU rental rates, and mathematical break-even vs managed APIs
| Hardware Tier |
VRAM & Spec |
CapEx Cost |
Cloud / Mo (24/7) |
Runnable Models |
| 💰 Tier 1: Ultra Budget |
24GB (1x RTX 3090 / 4090) |
$1,200–$1,800 |
$245–$540 |
Qwen 2.5 14B, Llama 3.1 8B, Whisper, Kokoro |
| ⚖️ Tier 2: Best Value |
48GB (2x 4090 / Mac 192GB) |
$3,500–$5,500 |
$1,008–$1,584 |
Qwen 2.5 72B (Q4, 35 t/s), Llama 3.3 70B |
| 🚀 Tier 3: High Perf |
96GB (4x 4090 / 2x L40S) |
$12,000–$18,000 |
$2,016–$3,240 |
Qwen 2.5 72B FP16 (70 t/s unquantized) |
| 🏢 Tier 4: Cluster |
640GB+ (8x H100 SXM) |
$300,000+ |
$14,169–$20,160 |
DeepSeek-V3/R1 Full 671B MoE (FP8, 85 t/s) |
| Monthly Volume |
Commodity API (DeepSeek/Flash) |
Frontier API (GPT-4o / Claude) |
Financial Decision & Verdict |
| 1M Tokens |
$0.20–$0.48 |
$4.38–$6.00 |
API is 98% Cheaper Self-hosting makes no economic sense. |
| 10M Tokens |
$2.00–$4.78 |
$43.75–$60.00 |
API is 50x Cheaper Managed APIs are vastly cheaper. |
| 100M Tokens |
$20.00–$47.75 |
$437.50–$600.00 |
Partial Break-Even Self-hosting 14B beats Frontier APIs. |
| 1B Tokens |
$200.00–$477.50 |
$4,375.00–$6,000.00 |
Break-Even Reached Self-hosting 70B saves $3.3k–$5k/mo vs Frontier. |
14. Authoritative Documentation & Verification Directory
Direct links to official pricing schedules and provider documentation
| Provider / Platform |
Official Endpoint |
Verified Scope |
| Google Cloud / AI Studio |
ai.google.dev/pricing • cloud.google.com/vertex-ai/pricing |
Gemini 2.0/3 tokens, Live API, Imagen 3, Veo 2, Search Grounding ($35/1k). |
| OpenAI |
openai.com/api/pricing • openai.com/chatgpt/pricing |
GPT-4o, o1, o3-mini, Realtime, DALL-E 3, Whisper, ChatGPT Team/Enterprise. |
| Anthropic |
docs.anthropic.com/about-claude/pricing |
Claude 3.7 hybrid thinking, Prompt Caching (90% read off), Claude Team ($25). |
| DeepSeek |
api-docs.deepseek.com/quick_start/pricing |
DeepSeek-V3 / R1 API pricing, MIT open-source terms, caching rules. |
| Alibaba Cloud |
alibabacloud.com/model-studio/model-pricing |
Qwen 2.5 series (7B to 72B), Qwen-Max, Apache 2.0 open weights. |
| AWS Amazon Bedrock |
aws.amazon.com/bedrock/pricing |
Bedrock On-Demand, Provisioned Throughput (PTUs), EDP / PPA discount mechanics. |
| Cursor / Anysphere |
cursor.com/pricing |
Cursor Pro ($20/mo), Cursor Business ($40/mo), Composer architecture, Privacy Mode. |
| Media & Voice Providers |
elevenlabs.io/pricing • cartesia.ai/pricing |
Cartesia Sonic ($0.038/1k chars), Deepgram Nova-2 ($0.0043/min), ElevenLabs. |
| Cloud GPU Infrastructure |
runpod.io/pricing • lambdalabs.com |
RTX 4090 ($0.34–$0.75/hr), L40S, NVIDIA H100 ($2.46–$3.50/hr) compute. |