A clean, neutral benchmark of foundation model token costs, prompt caching rates, and context limits. Normalized per 1 Million tokens.
| Model | Provider | Context | Input / 1M | Output / 1M | Cache Read / 1M | Cache Write / 1M |
|---|---|---|---|---|---|---|
| GPT-4o | OpenAI | 128k | 2.5000 | 10.0000 | 1.2500 | 2.5000 |
| GPT-4o mini | OpenAI | 128k | 0.1500 | 0.6000 | 0.0750 | 0.1500 |
| o1Reasoning | OpenAI | 200k | 15.0000 | 60.0000 | 7.5000 | 15.0000 |
| o1-miniReasoning | OpenAI | 128k | 1.1000 | 4.4000 | 0.5500 | 1.1000 |
| o3-miniReasoning | OpenAI | 200k | 1.1000 | 4.4000 | 0.5500 | 1.1000 |
| GPT-4 Turbo | OpenAI | 128k | 10.0000 | 30.0000 | - | - |
| GPT-3.5 Turbo | OpenAI | 16k | 0.5000 | 1.5000 | - | - |
| Claude 3.7 SonnetReasoning | Anthropic | 200k | 3.0000 | 15.0000 | 0.3000 | 3.7500 |
| Claude 3.5 Sonnet | Anthropic | 200k | 3.0000 | 15.0000 | 0.3000 | 3.7500 |
| Claude 3.5 Haiku | Anthropic | 200k | 0.8000 | 4.0000 | 0.0800 | 1.0000 |
| Claude 3 Opus | Anthropic | 200k | 15.0000 | 75.0000 | 1.5000 | 18.7500 |
| Gemini 2.5 ProReasoning | 2000k | 1.2500 | 5.0000 | 0.3125 | 1.2500 | |
| Gemini 2.5 Flash | 1000k | 0.0750 | 0.3000 | 0.0187 | 0.0750 | |
| Gemini 1.5 Pro | 2000k | 1.2500 | 5.0000 | 0.3125 | 1.2500 | |
| Gemini 1.5 Flash | 1000k | 0.0750 | 0.3000 | 0.0187 | 0.0750 | |
| Gemini 1.5 Flash-8B | 1000k | 0.0375 | 0.1500 | 0.0100 | 0.0375 | |
| DeepSeek-V3 | DeepSeek | 64k | 0.1400 | 0.2800 | 0.0140 | 0.1400 |
| DeepSeek-R1Reasoning | DeepSeek | 64k | 0.5500 | 2.1900 | 0.1400 | 0.5500 |
| Llama 3.3 70B | Meta | 128k | 0.2000 | 0.4000 | - | - |
| Llama 3.1 405B | Meta | 128k | 1.2500 | 2.5000 | - | - |
| Llama 3.1 70B | Meta | 128k | 0.2500 | 0.5000 | - | - |
| Llama 3.1 8B | Meta | 128k | 0.0500 | 0.0800 | - | - |
| Mistral Large 2 | Mistral | 128k | 2.0000 | 6.0000 | - | - |
| Codestral 25.01 | Mistral | 256k | 0.3000 | 0.9000 | - | - |
| Mistral Small | Mistral | 128k | 0.1000 | 0.3000 | - | - |
| Grok-2 | xAI | 131k | 2.0000 | 10.0000 | - | - |
| Grok-2 mini | xAI | 131k | 0.2000 | 1.0000 | - | - |
| Command R+ | Cohere | 128k | 2.5000 | 10.0000 | - | - |
| Command R | Cohere | 128k | 0.1500 | 0.6000 | - | - |
| Amazon Nova Pro | Amazon | 300k | 0.8000 | 3.2000 | 0.2000 | 0.8000 |
| Amazon Nova Lite | Amazon | 300k | 0.0600 | 0.2400 | 0.0150 | 0.0600 |
| Amazon Nova Micro | Amazon | 128k | 0.0350 | 0.1400 | - | - |
Download the entire real-time model catalog as structured JSON.
Download models.json →Machine-readable index optimized for LLMs and answer engines.
View /llms.txt →Explicit crawler allowlist for search bots, GPTBot, ClaudeBot, and Perplexity.
View /robots.txt →