Every model, one place

One account, every major provider. Switch models mid-conversation in the app, or call any of them by ID through the OpenAI-compatible API — point your client at https://api.privateer.pro/v1 and set model. Chat, image, video, music, sound effects, voices and 3D meshes — every one of them with its ID and what it costs.

Chat billed per token Media billed per unit of output Prices at the developer-API rate
545 models 49+ chat providers 423 chat · 57 image · 30 video · 25 audio · 10 3D 21 confidential 54 on sale · up to 90% off

On sale now 54 models up to 90% off · in / out, USD per 1M tokens · provider promotions, billed through at the discounted rate

DeepSeek V4 Flash 0731−90%
deepseek/deepseek-v4-flash-0731
In$0.02
Outfrom $0.18list $1.79
DeepSeek V4 Pro 0423−88%
deepseek/deepseek-v4-pro
In$0.29
Out$0.58
Mercury 2.5−80%
inception/mercury-2.5
In$0.06
Out$0.21
Qwen3 235B A22B Instruct 2507−75%
qwen/qwen3-235b-a22b-2507
In$0.12
Out$0.49
Ling 3.0 Flash VL−72%
inclusionai/ling-3.0-flash-vl
In$0.03
Out$0.09
GLM 5.3−70%
z-ai/glm-5.3
Infrom $0.04list $1.96
Outfrom $1.85list $6.16
Solar Pro 4−70%
upstage/solar-pro4
In$0.13
Out$0.50
DeepSeek V4 Flash 0423−70%
deepseek/deepseek-v4-flash
In$0.04
Outfrom $0.12list $1.79
Ling 3.0 Flash−65%
inclusionai/ling-3.0-flash
In$0.03
Out$0.09
LongCat 2.0−60%
meituan/longcat-2.0
In$0.42
Out$1.68
Qwen3 30B A3B Instruct 2507−55%
qwen/qwen3-30b-a3b-instruct-2507
In$0.07
Out$0.27
GLM 5.2−54%
z-ai/glm-5.2
In$0.03
Outfrom $2.52list $22.40
DeepSeek V4.1 Flash−53%
deepseek/deepseek-v4.1-flash
In<$0.01
Outfrom $0.25list $3.36
DeepSeek V4 Flash Vision Exp−51%
deepseek/deepseek-v4-flash-vision-exp
In$0.30
Out$0.91
Solar Mini 4−50%
upstage/solar-mini4
In$0.07
Out$0.28
Gemini 3.8 Flash−50%
google/gemini-3.8-flash
Infrom $0.52list $1.05
Outfrom $2.63list $5.25
Gemini 3.8 Flash (batch)−50%
google/gemini-3.8-flash:batch
In$0.52
Out$2.63
GLM 5.3 Flash−50%
z-ai/glm-5.3-flash
Infrom $0.05list $0.21
Outfrom $0.35list $0.70
GLM 5.3 Flash (batch)−50%
z-ai/glm-5.3-flash:batch
In$0.08
Out$0.28
Gemini 3.7 Flash−50%
google/gemini-3.7-flash
Infrom $0.52list $1.05
Outfrom $2.63list $5.25
Gemini 3.7 Flash (batch)−50%
google/gemini-3.7-flash:batch
In$0.52
Out$2.63
DeepSeek V4 Pro 0813−50%
deepseek/deepseek-v4-pro-0813
Infrom $0.27list $0.77
Outfrom $2.77list $7.00
GPT-5.6 Sol Pro−50%
openai/gpt-5.6-sol-pro
Infrom $1.40list $2.80
Outfrom $7.00list $14.00
GPT-5.6 Sol Pro (batch)−50%
openai/gpt-5.6-sol-pro:batch
In$1.40
Out$7.00
GPT-5.6 Sol−50%
openai/gpt-5.6-sol
Infrom $1.40list $2.80
Outfrom $7.00list $14.00
GPT-5.6 Sol (batch)−50%
openai/gpt-5.6-sol:batch
In$1.40
Out$7.00
MiniMax M3−50%
minimax/minimax-m3
Infrom $0.32list $0.42
Outfrom $1.34list $1.68
Ling 3.0 Flash Fin−44%
inclusionai/ling-3.0-flash-fin
In$0.06
Out$0.17
Laguna XS 2.1−40%
poolside/laguna-xs-2.1
In$0.08
Out$0.17
GLM 5−40%
z-ai/glm-5
In$0.84
Out$2.69
Qwen3 Coder Next−40%
qwen/qwen3-coder-next
In$0.17
Out$1.12
DeepSeek V3.1 Terminus−40%
deepseek/deepseek-v3.1-terminus
In$0.38
Out$1.40
GLM 5.3 (batch)−38%
z-ai/glm-5.3:batch
In$0.63
Out$2.80
Kimi K2.6−37%
moonshotai/kimi-k2.6
Infrom $0.65list $1.33
Outfrom $3.36list $5.60
Kimi K3−35%
moonshotai/kimi-k3
Infrom $0.82list $0.94
Outfrom $13.65list $19.60
GLM 5.1−31%
z-ai/glm-5.1
Infrom $1.35list $1.96
Outfrom $4.25list $6.16
DeepSeek V4.1 Flash (batch)−30%
deepseek/deepseek-v4.1-flash:batch
In$0.16
Out$0.47
MiMo-V2.5-Pro−30%
xiaomi/mimo-v2.5-pro
Infrom $0.43list $0.61
Outfrom $0.85list $1.22
MiniMax M2.7−30%
minimax/minimax-m2.7
In$0.29
Out$1.18
DeepSeek V3.2−28%
deepseek/deepseek-v3.2
Infrom $0.29list $0.39
Outfrom $0.43list $0.59
GLM 4.7−27%
z-ai/glm-4.7
Infrom $0.56list $0.84
Outfrom $2.45list $3.08
Mistral Nemo−27%
mistralai/mistral-nemo
In$0.03
Out$0.04
Qwen3.8 27B−25%
qwen/qwen3.8-27b
Infrom $0.03list $0.59
Outfrom $2.09list $3.57
Kimi K2.7 Code−25%
moonshotai/kimi-k2.7-code
In$0.94
Outfrom $4.20list $4.69
Gemma 4 26B A4B −25%
google/gemma-4-26b-a4b-it
Infrom $0.06list $0.09
Out$0.31
Llama 3.3 70B Instruct−25%
meta-llama/llama-3.3-70b-instruct
Infrom $0.14list $0.31
Outfrom $0.45list $0.70
Qwen3.8 2.4T A95B−20%
qwen/qwen3.8-2.4t-a95b
In$2.80
Out$8.40
Nemotron 3.5 Lightning−15%
nvidia/nemotron-3.5-lightning
Infrom $0.05list $0.08
Out$0.22
MiMo-V2.5−15%
xiaomi/mimo-v2.5
Infrom $0.17list $0.20
Outfrom $0.33list $0.39
MiniMax M2−15%
minimax/minimax-m2
Infrom $0.36list $0.42
Outfrom $1.43list $1.68
Laguna S 2.1−10%
poolside/laguna-s-2.1
In$0.13
Out$0.25
MiniMax M2.5−10%
minimax/minimax-m2.5
In$0.38
Outfrom $1.33list $1.51
DeepSeek V3−10%
deepseek/deepseek-chat
In$0.36
Outfrom $1.25list $1.44
Kimi K2.5−5%
moonshotai/kimi-k2.5
In$0.63
Out$3.15

Chat & reasoning 423 models 49 providers · search, filter or sort the whole list

Confidential 21 of these models run inside an attested hardware enclave — Intel TDX, NVIDIA Confidential Computing or AMD SEV-SNP, served by Tinfoil, Phala and NEAR AI. The prompt is decrypted only in there, and the app verifies the hardware attestation on every response; “verified” is claimed only where your own device does the attesting. Context and feature tags are not published per host, so those cells read “—”. What confidential compute means →

GLM 5.3 FlashConfidential
near/z-ai/glm-5.3-flash

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear…

NEAR AIIntel TDX + NVIDIA confidential GPU
Context
—
Input /1M
$0.21
Output /1M
$0.70
Qwen3-VL-30B-A3B-InstructConfidential
near/Qwen/Qwen3-VL-30B-A3B-Instruct

Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct…

NEAR AIIntel TDX + NVIDIA confidential GPU
Context
—
Input /1M
$0.21
Output /1M
$0.77
Qwen 3.6 35B A3B FP8Confidential
near/Qwen/Qwen3.6-35B-A3B-FP8
NEAR AIIntel TDX + NVIDIA confidential GPU
Context
—
Input /1M
$0.24
Output /1M
$1.54
Qwen 3.8 27BConfidential
near/Qwen/Qwen3.8-27B

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal…

NEAR AIIntel TDX + NVIDIA confidential GPU
Context
—
Input /1M
$0.62
Output /1M
$4.62
Kimi K2.6Confidential
near/moonshotai/kimi-k2.6

Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent…

NEAR AIIntel TDX + NVIDIA confidential GPU
Context
—
Input /1M
$1.13
Output /1M
$5.39
Kimi K3Confidential
near/moonshotai/kimi-k3

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and…

NEAR AIIntel TDX + NVIDIA confidential GPU
Context
—
Input /1M
$4.62
Output /1M
$23.10
Nemotron 3.5 LightningConfidential
phala/nvidia/nemotron-3.5-lightning

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for…

PhalaIntel TDX + NVIDIA confidential GPU
Context
—
Input /1M
$0.10
Output /1M
$0.28
Qwen2.5 7B InstructConfidential
phala/qwen/qwen-2.5-7b-instruct

Qwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge…

PhalaIntel TDX + NVIDIA confidential GPU
Context
—
Input /1M
$0.14
Output /1M
$0.28
Gemma-4 26B-A4B Uncensored (Heretic)Confidential
phala/phala/gemma-4-26b-a4b-uncensored
PhalaIntel TDX + NVIDIA confidential GPU
Context
—
Input /1M
$0.21
Output /1M
$0.98
Qwen3.8 27BConfidential
phala/qwen/qwen3.8-27b

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal…

PhalaIntel TDX + NVIDIA confidential GPU
Context
—
Input /1M
$0.28
Output /1M
$3.50
Muse Glimmer 30BConfidential
phala/meta/muse-glimmer-30b

Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous…

PhalaIntel TDX + NVIDIA confidential GPU
Context
—
Input /1M
$0.42
Output /1M
$1.54
Qwen3.8 27B Uncensored (Aggressive)Confidential
phala/phala/qwen3.8-27b-uncensored
PhalaIntel TDX + NVIDIA confidential GPU
Context
—
Input /1M
$0.42
Output /1M
$2.10
Qwen3.6 27BConfidential
phala/qwen/qwen3.6-27b

Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal…

PhalaIntel TDX + NVIDIA confidential GPU
Context
—
Input /1M
$0.45
Output /1M
$3.78
GLM 5.3Confidential
phala/z-ai/glm-5.3

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and…

PhalaIntel TDX + NVIDIA confidential GPU
Context
—
Input /1M
$1.96
Output /1M
$6.16
GPT-OSS 120BConfidential
tinfoil/gpt-oss-120b

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and…

TinfoilAMD SEV-SNP + NVIDIA confidential GPU
Context
—
Input /1M
$0.21
Output /1M
$0.84
Gemma 4 31BConfidential
tinfoil/gemma4-31b
TinfoilAMD SEV-SNP + NVIDIA confidential GPU
Context
—
Input /1M
$0.56
Output /1M
$1.40
GLM 5.3 FlashConfidential
tinfoil/glm-5-3-flash
TinfoilAMD SEV-SNP + NVIDIA confidential GPU
Context
—
Input /1M
$0.56
Output /1M
$1.75
DeepSeek V4.1 FlashConfidential
tinfoil/deepseek-v4-1-flash
TinfoilAMD SEV-SNP + NVIDIA confidential GPU
Context
—
Input /1M
$0.91
Output /1M
$2.03
Llama 3.3 70BConfidential
tinfoil/llama3-3-70b
TinfoilAMD SEV-SNP + NVIDIA confidential GPU
Context
—
Input /1M
$2.45
Output /1M
$3.85
GLM 5.3Confidential
tinfoil/glm-5-3
TinfoilAMD SEV-SNP + NVIDIA confidential GPU
Context
—
Input /1M
$2.52
Output /1M
$8.05
Kimi K3Confidential
tinfoil/kimi-k3

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and…

TinfoilAMD SEV-SNP + NVIDIA confidential GPU
Context
—
Input /1M
$5.60
Output /1M
$28.00
gpt-oss-20b
openai/gpt-oss-20b

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture…

OpenAITool callingReasoningStructured outputPrompt caching
Context
131K
Input /1M
$0.03
Output /1M
$0.13
gpt-oss-20b (batch)
openai/gpt-oss-20b:batch

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture…

OpenAITool callingReasoningStructured output
Context
131K
Input /1M
$0.03
Output /1M
$0.16
GPT-5 Nano (batch)
openai/gpt-5-nano:batch

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
400K
Input /1M
$0.03
Output /1M
$0.28
gpt-oss-120b (batch)
openai/gpt-oss-120b:batch

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and…

OpenAITool callingReasoningStructured output
Context
131K
Input /1M
$0.04
Output /1M
$0.19
gpt-oss-120b
openai/gpt-oss-120b

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and…

OpenAITool callingReasoningStructured output
Context
131K
Input /1M
$0.05
Output /1M
$0.24
GPT-6 Luna Pro (batch)
openai/gpt-6-luna-pro:batch

GPT-6 Luna Pro is the same underlying model as GPT-6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$0.07
Output /1M
$0.35
GPT-6 Luna (batch)
openai/gpt-6-luna:batch

GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$0.07
Output /1M
$0.35
GPT-5 Nano
openai/gpt-5-nano

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
400K
Input /1M
$0.07
Output /1M
$0.56
GPT-4.1 Nano (batch)
openai/gpt-4.1-nano:batch

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a…

OpenAIImage inFile inTool callingStructured outputPrompt caching
Context
1.0M
Input /1M
$0.07
Output /1M
$0.28
gpt-oss-safeguard-20b
openai/gpt-oss-safeguard-20b

gpt-oss-safeguard-20b is a safety reasoning model from OpenAI built upon gpt-oss-20b. This open-weight, 21B-parameter Mixture-of-Experts (MoE)…

OpenAITool callingReasoningStructured outputPrompt caching
Context
131K
Input /1M
$0.10
Output /1M
$0.42
GPT-4o-mini (batch)
openai/gpt-4o-mini:batch

GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs. As their most advanced small model…

OpenAIImage inFile inTool callingStructured outputPrompt caching
Context
128K
Input /1M
$0.10
Output /1M
$0.42
GPT-6 Luna Pro
openai/gpt-6-luna-pro

GPT-6 Luna Pro is the same underlying model as GPT-6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$0.14
Output /1M
$0.70
GPT-6 Luna
openai/gpt-6-luna

GPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$0.14
Output /1M
$0.70
GPT-5.6 Luna Pro (batch)
openai/gpt-5.6-luna-pro:batch

GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$0.14
Output /1M
$0.84
GPT-5.6 Luna (batch)
openai/gpt-5.6-luna:batch

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$0.14
Output /1M
$0.84
GPT-5.4 Nano (batch)
openai/gpt-5.4-nano:batch

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
400K
Input /1M
$0.14
Output /1M
$0.88
GPT-4.1 Nano
openai/gpt-4.1-nano

For tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a…

OpenAIImage inFile inTool callingStructured outputPrompt caching
Context
1.0M
Input /1M
$0.14
Output /1M
$0.56
GPT-5 Mini (batch)
openai/gpt-5-mini:batch

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
400K
Input /1M
$0.17
Output /1M
$1.40
GPT-4o-mini
openai/gpt-4o-mini

GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs. As their most advanced small model…

OpenAIImage inFile inTool callingStructured outputPrompt caching
Context
128K
Input /1M
$0.21
Output /1M
$0.84
GPT-4o-mini (2024-07-18)
openai/gpt-4o-mini-2024-07-18

GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs. As their most advanced small model…

OpenAIImage inFile inTool callingStructured outputPrompt caching
Context
128K
Input /1M
$0.21
Output /1M
$0.84
GPT-5.6 Luna Pro
openai/gpt-5.6-luna-pro

GPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$0.28
Output /1M
$1.68
GPT-5.6 Luna
openai/gpt-5.6-luna

GPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$0.28
Output /1M
$1.68
GPT-5.4 Nano
openai/gpt-5.4-nano

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
400K
Input /1M
$0.28
Output /1M
$1.75
GPT-4.1 Mini (batch)
openai/gpt-4.1-mini:batch

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million…

OpenAIImage inFile inTool callingStructured outputPrompt caching
Context
1.0M
Input /1M
$0.28
Output /1M
$1.12
GPT-5.1-Codex-Mini
openai/gpt-5.1-codex-mini

GPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex

OpenAIImage inTool callingReasoningStructured outputPrompt caching
Context
400K
Input /1M
$0.35
Output /1M
$2.80
GPT-5 Mini
openai/gpt-5-mini

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
400K
Input /1M
$0.35
Output /1M
$2.80
GPT-3.5 Turbo (batch)
openai/gpt-3.5-turbo:batch

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional…

OpenAITool callingStructured output
Context
16K
Input /1M
$0.35
Output /1M
$1.05
GPT-5.4 Mini (batch)
openai/gpt-5.4-mini:batch

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
400K
Input /1M
$0.52
Output /1M
$3.15
GPT-4.1 Mini
openai/gpt-4.1-mini

GPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million…

OpenAIImage inFile inTool callingStructured outputPrompt caching
Context
1.0M
Input /1M
$0.56
Output /1M
$2.24
GPT-3.5 Turbo
openai/gpt-3.5-turbo

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional…

OpenAITool callingStructured output
Context
16K
Input /1M
$0.70
Output /1M
$2.10
o4 Mini (batch)
openai/o4-mini:batch

OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
200K
Input /1M
$0.77
Output /1M
$3.08
o3 Mini (batch)
openai/o3-mini:batch

OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding…

OpenAIFile inTool callingReasoningStructured outputPrompt caching
Context
200K
Input /1M
$0.77
Output /1M
$3.08
GPT-5.1 (batch)
openai/gpt-5.1:batch

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
400K
Input /1M
$0.88
Output /1M
$7.00
GPT-5 (batch)
openai/gpt-5:batch

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
400K
Input /1M
$0.88
Output /1M
$7.00
GPT-5.4 Mini
openai/gpt-5.4-mini

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
400K
Input /1M
$1.05
Output /1M
$6.30
GPT-5.2 (batch)
openai/gpt-5.2:batch

GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
400K
Input /1M
$1.22
Output /1M
$9.80
GPT-6 Sol Pro (batch)
openai/gpt-6-sol-pro:batch

GPT-6 Sol Pro is the same underlying model as GPT-6 Sol, served with reasoning.mode set to pro for higher-quality responses on complex tasks. Cost…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$1.40
Output /1M
$7.00
GPT-6 Sol (batch)
openai/gpt-6-sol:batch

GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$1.40
Output /1M
$7.00
GPT-5.6 Terra Pro (batch)
openai/gpt-5.6-terra-pro:batch

GPT-5.6 Terra Pro is the same underlying model as GPT-5.6 Terra, served with reasoning.mode set to pro for higher-quality responses on complex…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$1.40
Output /1M
$8.40
GPT-5.6 Terra (batch)
openai/gpt-5.6-terra:batch

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$1.40
Output /1M
$8.40
GPT-5.6 Sol Pro (batch)−50%
openai/gpt-5.6-sol-pro:batch

GPT-5.6 Sol Pro is the same underlying model as GPT-5.6 Sol, served with reasoning.mode set to pro for higher-quality responses on complex tasks…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$1.40
Output /1M
$7.00
GPT-5.6 Sol (batch)−50%
openai/gpt-5.6-sol:batch

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$1.40
Output /1M
$7.00
o3 (batch)
openai/o3:batch

o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
200K
Input /1M
$1.40
Output /1M
$5.60
GPT-4.1 (batch)
openai/gpt-4.1:batch

GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context…

OpenAIImage inFile inTool callingStructured outputPrompt caching
Context
1.0M
Input /1M
$1.40
Output /1M
$5.60
GPT-3.5 Turbo (older v0613)
openai/gpt-3.5-turbo-0613

GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional…

OpenAITool callingStructured output
Context
4K
Input /1M
$1.40
Output /1M
$2.80
o4 Mini High
openai/o4-mini-high

OpenAI o4-mini-high is the same model as o4-mini with reasoningeffort set to high. OpenAI o4-mini is a compact reasoning model in the o-series…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
200K
Input /1M
$1.54
Output /1M
$6.16
o4 Mini
openai/o4-mini

OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
200K
Input /1M
$1.54
Output /1M
$6.16
o3 Mini High
openai/o3-mini-high

OpenAI o3-mini-high is the same model as o3-mini with reasoningeffort set to high. o3-mini is a cost-efficient language model optimized for STEM…

OpenAIFile inTool callingReasoningStructured outputPrompt caching
Context
200K
Input /1M
$1.54
Output /1M
$6.16
o3 Mini
openai/o3-mini

OpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding…

OpenAIFile inTool callingReasoningStructured outputPrompt caching
Context
200K
Input /1M
$1.54
Output /1M
$6.16
GPT-5.4 (batch)
openai/gpt-5.4:batch

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$1.75
Output /1M
$10.50
GPT-5.1-Codex-Max
openai/gpt-5.1-codex-max

GPT-5.1-Codex-Max is OpenAI’s latest agentic coding model, designed for long-running, high-context software development tasks. It is based on an…

OpenAIImage inTool callingReasoningStructured outputPrompt caching
Context
400K
Input /1M
$1.75
Output /1M
$14.00
GPT-5.1
openai/gpt-5.1

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
400K
Input /1M
$1.75
Output /1M
$14.00
GPT-5.1-Codex
openai/gpt-5.1-codex

GPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both interactive…

OpenAIImage inTool callingReasoningStructured outputPrompt caching
Context
400K
Input /1M
$1.75
Output /1M
$14.00
GPT-5
openai/gpt-5

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
400K
Input /1M
$1.75
Output /1M
$14.00
GPT-4o (batch)
openai/gpt-4o:batch

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level…

OpenAIImage inFile inTool callingStructured outputPrompt caching
Context
128K
Input /1M
$1.75
Output /1M
$7.00
GPT-3.5 Turbo Instruct
openai/gpt-3.5-turbo-instruct

This model is a variant of GPT-3.5 Turbo tuned for instructional prompts and omitting chat-related optimizations. Training data: up to Sep 2021.

OpenAIStructured output
Context
4K
Input /1M
$2.10
Output /1M
$2.80
GPT-5.3-Codex
openai/gpt-5.3-codex

GPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
400K
Input /1M
$2.45
Output /1M
$19.60
GPT-5.2-Codex
openai/gpt-5.2-codex

GPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows. It is designed for both interactive…

OpenAIImage inTool callingReasoningStructured outputPrompt caching
Context
400K
Input /1M
$2.45
Output /1M
$19.60
GPT-5.2 Chat
openai/gpt-5.2-chat

GPT-5.2 Chat (AKA Instant) is the fast, lightweight member of the 5.2 family, optimized for low-latency chat while retaining strong general…

OpenAIImage inFile inTool callingStructured outputPrompt caching
Context
128K
Input /1M
$2.45
Output /1M
$19.60
GPT-5.2
openai/gpt-5.2

GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
400K
Input /1M
$2.45
Output /1M
$19.60
GPT-6.1 Sol Pro
openai/gpt-6.1-sol-pro

GPT-6.1 Sol Pro is the same underlying model as GPT-6.1 Sol, served with reasoning.mode set to pro for higher-quality responses on complex tasks…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$2.80
Output /1M
$14.00
GPT-6.1 Sol
openai/gpt-6.1-sol

GPT-6.1 Sol is an upgrade to GPT-6 Sol from OpenAI, positioned below the flagship GPT-6 Astra in the GPT-6 series. It is suited for agentic coding…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$2.80
Output /1M
$14.00
GPT-6 Sol Pro
openai/gpt-6-sol-pro

GPT-6 Sol Pro is the same underlying model as GPT-6 Sol, served with reasoning.mode set to pro for higher-quality responses on complex tasks. Cost…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$2.80
Output /1M
$14.00
GPT-6 Sol
openai/gpt-6-sol

GPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$2.80
Output /1M
$14.00
GPT-5.6 Terra Pro
openai/gpt-5.6-terra-pro

GPT-5.6 Terra Pro is the same underlying model as GPT-5.6 Terra, served with reasoning.mode set to pro for higher-quality responses on complex…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$2.80
Output /1M
$16.80
GPT-5.6 Terra
openai/gpt-5.6-terra

GPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$2.80
Output /1M
$16.80
GPT-5.6 Sol Pro−50%
openai/gpt-5.6-sol-pro

GPT-5.6 Sol Pro is the same underlying model as GPT-5.6 Sol, served with reasoning.mode set to pro for higher-quality responses on complex tasks…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$2.80from $1.40
Output /1M
$14.00from $7.00
GPT-5.6 Sol−50%
openai/gpt-5.6-sol

GPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$2.80from $1.40
Output /1M
$14.00from $7.00
o3
openai/o3

o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
200K
Input /1M
$2.80
Output /1M
$11.20
GPT-4.1
openai/gpt-4.1

GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context…

OpenAIImage inFile inTool callingStructured outputPrompt caching
Context
1.0M
Input /1M
$2.80
Output /1M
$11.20
GPT-5.5 (batch)
openai/gpt-5.5:batch

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$3.50
Output /1M
$21.00
GPT-5.4
openai/gpt-5.4

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$3.50
Output /1M
$21.00
GPT-4o (2024-11-20)
openai/gpt-4o-2024-11-20

The 2024-11-20 version of GPT-4o offers a leveled-up creative writing ability with more natural, engaging, and tailored writing to improve…

OpenAIImage inFile inTool callingStructured outputPrompt caching
Context
128K
Input /1M
$3.50
Output /1M
$14.00
GPT-4o (2024-08-06)
openai/gpt-4o-2024-08-06

The 2024-08-06 version of GPT-4o offers improved performance in structured outputs, with the ability to supply a JSON schema in the responeformat…

OpenAIImage inFile inTool callingStructured outputPrompt caching
Context
128K
Input /1M
$3.50
Output /1M
$14.00
GPT-4o
openai/gpt-4o

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level…

OpenAIImage inFile inTool callingStructured outputPrompt caching
Context
128K
Input /1M
$3.50
Output /1M
$14.00
GPT-3.5 Turbo 16k
openai/gpt-3.5-turbo-16k

This model offers four times the context length of gpt-3.5-turbo, allowing it to support approximately 20 pages of text in a single request at a…

OpenAITool callingStructured output
Context
16K
Input /1M
$4.20
Output /1M
$5.60
GPT-6 Astra (batch)
openai/gpt-6-astra:batch

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$7.00
Output /1M
$35.00
GPT-6 Astra Pro (batch)
openai/gpt-6-astra-pro:batch

GPT-6 Astra Pro is the same underlying model as GPT-6 Astra, served with reasoning.mode set to pro for higher-quality responses on complex tasks…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$7.00
Output /1M
$35.00
GPT Chat Latest
openai/gpt-chat-latest

GPT Chat Latest points to OpenAI's stable API alias chat-latest that always resolves to the latest Instant chat model used in ChatGPT. As OpenAI…

OpenAIImage inFile inTool callingStructured outputPrompt caching
Context
400K
Input /1M
$7.00
Output /1M
$42.00
GPT-5.5
openai/gpt-5.5

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$7.00
Output /1M
$42.00
GPT-4o (2024-05-13)
openai/gpt-4o-2024-05-13

GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level…

OpenAIImage inFile inTool callingStructured output
Context
128K
Input /1M
$7.00
Output /1M
$21.00
GPT-4 Turbo (batch)
openai/gpt-4-turbo:batch

The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December…

OpenAIImage inTool callingStructured output
Context
128K
Input /1M
$7.00
Output /1M
$21.00
GPT-5 Pro (batch)
openai/gpt-5-pro:batch

GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex…

OpenAIImage inFile inTool callingReasoningStructured output
Context
400K
Input /1M
$10.50
Output /1M
$84.00
GPT-6 Astra
openai/gpt-6-astra

GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$14.00
Output /1M
$70.00
GPT-6 Astra Pro
openai/gpt-6-astra-pro

GPT-6 Astra Pro is the same underlying model as GPT-6 Astra, served with reasoning.mode set to pro for higher-quality responses on complex tasks…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$14.00
Output /1M
$70.00
GPT-4 Turbo
openai/gpt-4-turbo

The latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December…

OpenAIImage inTool callingStructured output
Context
128K
Input /1M
$14.00
Output /1M
$42.00
GPT-5.2 Pro (batch)
openai/gpt-5.2-pro:batch

GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is…

OpenAIImage inFile inTool callingReasoningStructured output
Context
400K
Input /1M
$14.70
Output /1M
$117.60
GPT-5.5 Pro (batch)
openai/gpt-5.5-pro:batch

GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token…

OpenAIImage inFile inTool callingReasoningStructured output
Context
1.1M
Input /1M
$21.00
Output /1M
$126.00
GPT-5.4 Pro (batch)
openai/gpt-5.4-pro:batch

GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex…

OpenAIImage inFile inTool callingReasoningStructured output
Context
1.1M
Input /1M
$21.00
Output /1M
$126.00
GPT-5 Pro
openai/gpt-5-pro

GPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex…

OpenAIImage inFile inTool callingReasoningStructured output
Context
400K
Input /1M
$21.00
Output /1M
$168.00
o1
openai/o1

The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trained…

OpenAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
200K
Input /1M
$21.00
Output /1M
$84.00
o3 Pro
openai/o3-pro

The o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model uses…

OpenAIImage inFile inTool callingReasoningStructured output
Context
200K
Input /1M
$28.00
Output /1M
$112.00
GPT-5.2 Pro
openai/gpt-5.2-pro

GPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is…

OpenAIImage inFile inTool callingReasoningStructured output
Context
400K
Input /1M
$29.40
Output /1M
$235.20
GPT-5.5 Pro
openai/gpt-5.5-pro

GPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token…

OpenAIImage inFile inTool callingReasoningStructured output
Context
1.1M
Input /1M
$42.00
Output /1M
$252.00
GPT-5.4 Pro
openai/gpt-5.4-pro

GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex…

OpenAIImage inFile inTool callingReasoningStructured output
Context
1.1M
Input /1M
$42.00
Output /1M
$252.00
GPT-4
openai/gpt-4

OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy than previous…

OpenAITool callingStructured output
Context
8K
Input /1M
$42.00
Output /1M
$84.00
o1-pro
openai/o1-pro

The o1 series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o1-pro model uses…

OpenAIImage inFile inReasoningStructured output
Context
200K
Input /1M
$210.00
Output /1M
$840.00
Claude Haiku 4.5 (batch)
anthropic/claude-haiku-4.5:batch

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
200K
Input /1M
$0.70
Output /1M
$3.50
Claude Sonnet 5.5 (batch)
anthropic/claude-sonnet-5.5:batch

Claude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, succeeding Claude Sonnet 5 as a direct upgrade. It is especially…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$1.40
Output /1M
$7.00
Claude Sonnet 5 (batch)
anthropic/claude-sonnet-5:batch

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$1.40
Output /1M
$7.00
Claude Haiku 4.5
anthropic/claude-haiku-4.5

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
200K
Input /1M
$1.40
Output /1M
$7.00
Claude Sonnet 4.6 (batch)
anthropic/claude-sonnet-4.6:batch

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$2.10
Output /1M
$10.50
Claude Sonnet 4.5 (batch)
anthropic/claude-sonnet-4.5:batch

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$2.10
Output /1M
$10.50
Claude Sonnet 5.5
anthropic/claude-sonnet-5.5

Claude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, succeeding Claude Sonnet 5 as a direct upgrade. It is especially…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$2.80
Output /1M
$14.00
Claude Opus 5.5 (batch)
anthropic/claude-opus-5.5:batch

Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$2.80
Output /1M
$14.00
Claude Sonnet 5
anthropic/claude-sonnet-5

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$2.80
Output /1M
$14.00
Claude Opus 5 (batch)
anthropic/claude-opus-5:batch

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$3.50
Output /1M
$17.50
Claude Opus 4.8 (batch)
anthropic/claude-opus-4.8:batch

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$3.50
Output /1M
$17.50
Claude Opus 4.7 (batch)
anthropic/claude-opus-4.7:batch

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$3.50
Output /1M
$17.50
Claude Opus 4.6 (batch)
anthropic/claude-opus-4.6:batch

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$3.50
Output /1M
$17.50
Claude Opus 4.5 (batch)
anthropic/claude-opus-4.5:batch

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
200K
Input /1M
$3.50
Output /1M
$17.50
Claude Sonnet 4.6
anthropic/claude-sonnet-4.6

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$4.20
Output /1M
$21.00
Claude Sonnet 4.5
anthropic/claude-sonnet-4.5

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$4.20
Output /1M
$21.00
Claude Sonnet 4
anthropic/claude-sonnet-4

Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved…

AnthropicImage inFile inTool callingReasoningPrompt caching
Context
200K
Input /1M
$4.20
Output /1M
$21.00
Claude Opus 5.5
anthropic/claude-opus-5.5

Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$5.60
Output /1M
$28.00
Claude Fable 5.1 (batch)
anthropic/claude-fable-5.1:batch

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$7.00
Output /1M
$35.00
Claude Opus 5
anthropic/claude-opus-5

Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$7.00
Output /1M
$35.00
Claude Fable 5 (batch)
anthropic/claude-fable-5:batch

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$7.00
Output /1M
$35.00
Claude Opus 4.8
anthropic/claude-opus-4.8

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$7.00
Output /1M
$35.00
Claude Opus 4.7
anthropic/claude-opus-4.7

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$7.00
Output /1M
$35.00
Claude Opus 4.6
anthropic/claude-opus-4.6

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$7.00
Output /1M
$35.00
Claude Opus 4.5
anthropic/claude-opus-4.5

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
200K
Input /1M
$7.00
Output /1M
$35.00
Claude Opus 4.1 (batch)
anthropic/claude-opus-4.1:batch

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
200K
Input /1M
$10.50
Output /1M
$52.50
Claude Fable 5.1
anthropic/claude-fable-5.1

Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$14.00
Output /1M
$70.00
Claude Fable 5
anthropic/claude-fable-5

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs…

AnthropicImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$14.00
Output /1M
$70.00
Claude Opus 4.1
anthropic/claude-opus-4.1

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It…

AnthropicImage inFile inTool callingReasoningPrompt caching
Context
200K
Input /1M
$21.00
Output /1M
$105.00
Gemini 2.5 Flash Lite (batch)
google/gemini-2.5-flash-lite:batch

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.07
Output /1M
$0.28
Gemma 3 4B
google/gemma-3-4b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over…

GoogleImage inStructured output
Context
131K
Input /1M
$0.07
Output /1M
$0.14
Gemma 3 12B
google/gemma-3-12b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over…

GoogleImage inTool callingStructured output
Context
131K
Input /1M
$0.07
Output /1M
$0.21
Gemma 4 26B A4B −25%
google/gemma-4-26b-a4b-it

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate…

GoogleImage inTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.09from $0.06
Output /1M
$0.31
Gemma 3 27B
google/gemma-3-27b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over…

GoogleImage inTool callingStructured outputPrompt caching
Context
131K
Input /1M
$0.11
Output /1M
$0.63
Gemma 4 31B
google/gemma-4-31b-it

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token…

GoogleImage inTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.13
Output /1M
$0.48
Gemini 2.5 Flash Lite
google/gemini-2.5-flash-lite

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.14
Output /1M
$0.56
Gemini 3.1 Flash Lite (batch)
google/gemini-3.1-flash-lite:batch

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.17
Output /1M
$1.05
Gemini 3.5 Flash Lite (batch)
google/gemini-3.5-flash-lite:batch

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.21
Output /1M
$1.75
Gemini 2.5 Flash (batch)
google/gemini-2.5-flash:batch

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.21
Output /1M
$1.75
Gemini 3.1 Flash Lite
google/gemini-3.1-flash-lite

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.35
Output /1M
$2.10
Gemini 3.1 Flash Lite Preview
google/gemini-3.1-flash-lite-preview

Gemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.35
Output /1M
$2.10
Gemini 3 Flash Preview (batch)
google/gemini-3-flash-preview:batch

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It…

GoogleImage inFile inAudio inTool callingReasoningStructured output
Context
1.0M
Input /1M
$0.35
Output /1M
$2.10
Gemini 3.5 Flash Lite
google/gemini-3.5-flash-lite

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.42
Output /1M
$3.50
Gemini 2.5 Flash
google/gemini-2.5-flash

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.42
Output /1M
$3.50
Gemini 3.8 Flash (batch)−50%
google/gemini-3.8-flash:batch

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.52
Output /1M
$2.63
Gemini 3.7 Flash (batch)−50%
google/gemini-3.7-flash:batch

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.52
Output /1M
$2.63
Gemini 3.6 Flash (batch)
google/gemini-3.6-flash:batch

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.52
Output /1M
$2.63
Gemini 3 Flash Preview
google/gemini-3-flash-preview

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.70
Output /1M
$4.20
Gemini 2.5 Pro (batch)
google/gemini-2.5-pro:batch

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.88
Output /1M
$7.00
Gemma 2 27B
google/gemma-2-27b-it

Gemma 2 27B by Google is an open model built from the same research and technology used to create the Gemini models. Gemma models are well-suited…

GoogleStructured output
Context
8K
Input /1M
$0.91
Output /1M
$0.91
Gemini 3.8 Flash−50%
google/gemini-3.8-flash

Gemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$1.05from $0.52
Output /1M
$5.25from $2.63
Gemini 3.7 Flash−50%
google/gemini-3.7-flash

Gemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$1.05from $0.52
Output /1M
$5.25from $2.63
Gemini 3.6 Flash
google/gemini-3.6-flash

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$1.05
Output /1M
$5.25
Gemini 3.5 Flash (batch)
google/gemini-3.5-flash:batch

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$1.05
Output /1M
$6.30
Gemini 3.1 Pro Preview (batch)
google/gemini-3.1-pro-preview:batch

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability…

GoogleImage inFile inAudio inTool callingReasoningStructured output
Context
1.0M
Input /1M
$1.40
Output /1M
$8.40
Gemini 2.5 Pro
google/gemini-2.5-pro

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$1.75
Output /1M
$14.00
Gemini 2.5 Pro Preview 06-05
google/gemini-2.5-pro-preview

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$1.75
Output /1M
$14.00
Gemini 3.5 Flash
google/gemini-3.5-flash

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$2.10
Output /1M
$12.60
Gemini 3.1 Pro Preview Custom Tools
google/gemini-3.1-pro-preview-customtools

Gemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a general bash…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$2.80
Output /1M
$16.80
Gemini 3.1 Pro Preview
google/gemini-3.1-pro-preview

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability…

GoogleImage inFile inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$2.80
Output /1M
$16.80
Grok Build 0.1
x-ai/grok-build-0.1

Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs…

xAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
256K
Input /1M
$1.40
Output /1M
$2.80
Grok 4.3 (batch)
x-ai/grok-4.3:batch

Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows…

xAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$1.40
Output /1M
$2.80
Grok 4.3
x-ai/grok-4.3

Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows…

xAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$1.75
Output /1M
$3.50
Grok 4.20 Multi-Agent
x-ai/grok-4.20-multi-agent

Grok 4.20 Multi-Agent is a variant of SpaceXAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel…

xAIImage inFile inReasoningStructured outputPrompt caching
Context
2M
Input /1M
$1.75
Output /1M
$3.50
Grok 4.20
x-ai/grok-4.20

Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest…

xAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
2M
Input /1M
$1.75
Output /1M
$3.50
Grok 4.7
x-ai/grok-4.7

Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running…

xAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
500K
Input /1M
$2.80
Output /1M
$8.40
Grok 4.6
x-ai/grok-4.6

Grok 4.6 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM. It is succeeded by Grok 4.7.

xAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
500K
Input /1M
$2.80
Output /1M
$8.40
Grok 4.5
x-ai/grok-4.5

Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM.

xAIImage inFile inTool callingReasoningStructured outputPrompt caching
Context
500K
Input /1M
$2.80
Output /1M
$8.40
Llama 3.2 1B Instruct
meta-llama/llama-3.2-1b-instruct

Llama 3.2 1B is a 1-billion-parameter language model focused on efficiently performing natural language tasks, such as summarization, dialogue, and…

Meta
Context
60K
Input /1M
$0.04
Output /1M
$0.28
Llama 3.2 3B Instruct
meta-llama/llama-3.2-3b-instruct

Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue…

MetaStructured output
Context
131K
Input /1M
$0.07
Output /1M
$0.46
Llama 3.1 8B Instruct
meta-llama/llama-3.1-8b-instruct

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has…

MetaTool callingStructured outputPrompt caching
Context
131K
Input /1M
$0.07
Output /1M
$0.11
Llama 4 Scout
meta-llama/llama-4-scout

Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of…

MetaImage inTool callingStructured output
Context
1.3M
Input /1M
$0.14
Output /1M
$0.42
Llama Guard 4 12B
meta-llama/llama-guard-4-12b

Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions…

MetaImage in
Context
164K
Input /1M
$0.25
Output /1M
$0.25
Llama 4 Maverick
meta-llama/llama-4-maverick

Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with…

MetaImage inTool callingStructured output
Context
1.0M
Input /1M
$0.26
Output /1M
$0.91
Llama 3.3 70B Instruct−25%
meta-llama/llama-3.3-70b-instruct

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The…

MetaTool callingStructured outputPrompt caching
Context
131K
Input /1M
$0.31from $0.14
Output /1M
$0.70from $0.45
Llama 3.1 70B Instruct
meta-llama/llama-3.1-70b-instruct

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high quality…

MetaTool callingStructured output
Context
131K
Input /1M
$0.56
Output /1M
$0.56
Mistral Nemo−27%
mistralai/mistral-nemo

A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting…

MistralTool callingStructured output
Context
131K
Input /1M
$0.03
Output /1M
$0.04
Mistral Small 3
mistralai/mistral-small-24b-instruct-2501

Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks. Released under the Apache 2.0…

MistralStructured output
Context
33K
Input /1M
$0.07
Output /1M
$0.11
Mistral Small 4 (batch)
mistralai/mistral-small-2603:batch

Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single…

MistralImage inTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.10
Output /1M
$0.42
Ministral 3 8B 2512 (batch)
mistralai/ministral-8b-2512:batch

A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.

MistralImage inTool callingStructured outputPrompt caching
Context
262K
Input /1M
$0.10
Output /1M
$0.10
Mistral Small 3.2 24B
mistralai/mistral-small-3.2-24b-instruct

Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduction, and…

MistralImage inTool callingStructured output
Context
256K
Input /1M
$0.13
Output /1M
$0.35
Ministral 3 3B 2512
mistralai/ministral-3b-2512

The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.

MistralImage inTool callingStructured outputPrompt caching
Context
131K
Input /1M
$0.14
Output /1M
$0.14
Voxtral Small 24B 2507
mistralai/voxtral-small-24b-2507

Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text…

MistralFile inAudio inTool callingStructured outputPrompt caching
Context
33K
Input /1M
$0.14
Output /1M
$0.42
Mistral Small 4
mistralai/mistral-small-2603

Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single…

MistralImage inTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.21
Output /1M
$0.84
Ministral 3 8B 2512
mistralai/ministral-8b-2512

A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.

MistralImage inTool callingStructured outputPrompt caching
Context
262K
Input /1M
$0.21
Output /1M
$0.21
Codestral 2508 (batch)
mistralai/codestral-2508:batch

Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as…

MistralFile inTool callingStructured outputPrompt caching
Context
256K
Input /1M
$0.21
Output /1M
$0.63
Ministral 3 14B 2512
mistralai/ministral-14b-2512

The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small…

MistralImage inTool callingStructured outputPrompt caching
Context
262K
Input /1M
$0.28
Output /1M
$0.28
Mistral Medium 3.1 (batch)
mistralai/mistral-medium-3.1:batch

Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver…

MistralImage inFile inTool callingStructured outputPrompt caching
Context
131K
Input /1M
$0.28
Output /1M
$1.40
Saba
mistralai/mistral-saba

Mistral Saba is a 24B-parameter language model specifically designed for the Middle East and South Asia, delivering accurate and contextually…

MistralFile inTool callingStructured outputPrompt caching
Context
33K
Input /1M
$0.28
Output /1M
$0.84
Mistral Large 3 2512 (batch)
mistralai/mistral-large-2512:batch

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B…

MistralImage inFile inTool callingStructured outputPrompt caching
Context
262K
Input /1M
$0.35
Output /1M
$1.05
Codestral 2508
mistralai/codestral-2508

Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as…

MistralFile inTool callingStructured outputPrompt caching
Context
256K
Input /1M
$0.42
Output /1M
$1.26
Mistral Small 3.1 24B
mistralai/mistral-small-3.1-24b-instruct

Mistral Small 3.1 24B Instruct is an upgraded variant of Mistral Small 3 (2501), featuring 24 billion parameters with advanced multimodal…

MistralImage inTool callingStructured output
Context
128K
Input /1M
$0.49
Output /1M
$0.78
Devstral 2 2512
mistralai/devstral-2512

Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer model…

MistralFile inTool callingStructured outputPrompt caching
Context
262K
Input /1M
$0.56
Output /1M
$2.80
Mistral Medium 3.1
mistralai/mistral-medium-3.1

Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver…

MistralImage inFile inTool callingStructured outputPrompt caching
Context
131K
Input /1M
$0.56
Output /1M
$2.80
Mistral Medium 3
mistralai/mistral-medium-3

Mistral Medium 3 is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced…

MistralImage inFile inTool callingStructured outputPrompt caching
Context
131K
Input /1M
$0.56
Output /1M
$2.80
Mistral Large 3 2512
mistralai/mistral-large-2512

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B…

MistralImage inFile inTool callingStructured outputPrompt caching
Context
262K
Input /1M
$0.70
Output /1M
$2.10
Mistral Medium 3.5 (batch)
mistralai/mistral-medium-3-5:batch

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed…

MistralImage inFile inTool callingReasoningStructured output
Context
262K
Input /1M
$1.05
Output /1M
$5.25
Mistral Medium 3.5
mistralai/mistral-medium-3-5

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed…

MistralImage inFile inTool callingReasoningStructured output
Context
262K
Input /1M
$2.10
Output /1M
$10.50
Mistral Large 2407
mistralai/mistral-large-2407

This is Mistral AI's flagship model, Mistral Large 2 (version mistral-large-2407). It's a proprietary weights-available model and excels at…

MistralFile inTool callingStructured outputPrompt caching
Context
131K
Input /1M
$2.80
Output /1M
$8.40
Mixtral 8x22B Instruct
mistralai/mixtral-8x22b-instruct

Mistral's official instruct fine-tuned version of Mixtral 8x22B. It uses 39B active parameters out of 141B, offering unparalleled cost efficiency…

MistralFile inTool callingStructured outputPrompt caching
Context
66K
Input /1M
$2.80
Output /1M
$8.40
Mistral Large
mistralai/mistral-large

This is Mistral AI's flagship model, Mistral Large 2 (version mistral-large-2407). It's a proprietary weights-available model and excels at…

MistralFile inTool callingStructured outputPrompt caching
Context
128K
Input /1M
$2.80
Output /1M
$8.40
DeepSeek V4.1 Flash−53%
deepseek/deepseek-v4.1-flash

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED)…

DeepSeekImage inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
<$0.01
Output /1M
$3.36from $0.25
DeepSeek V4 Flash 0731−90%
deepseek/deepseek-v4-flash-0731

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained…

DeepSeekTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.02
Output /1M
$1.79from $0.18
DeepSeek V4 Flash 0423−70%
deepseek/deepseek-v4-flash

DeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters…

DeepSeekTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.04
Output /1M
$1.79from $0.12
DeepSeek V4.1 Flash (batch)−30%
deepseek/deepseek-v4.1-flash:batch

DeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED)…

DeepSeekImage inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.16
Output /1M
$0.47
DeepSeek V4 Pro 0423−88%
deepseek/deepseek-v4-pro

DeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a…

DeepSeekTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.29
Output /1M
$0.58
DeepSeek V4 Flash Vision Exp−51%
deepseek/deepseek-v4-flash-vision-exp

DeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of DeepSeek V4 Flash 0731 from DeepSeek, adding image understanding while…

DeepSeekImage inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.30
Output /1M
$0.91
DeepSeek V3.1
deepseek/deepseek-chat-v3.1

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt…

DeepSeekTool callingReasoningStructured outputPrompt caching
Context
164K
Input /1M
$0.35
Output /1M
$1.33
DeepSeek V3−10%
deepseek/deepseek-chat

DeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions…

DeepSeekTool callingStructured output
Context
164K
Input /1M
$0.36
Output /1M
$1.44from $1.25
DeepSeek V3.2 Exp
deepseek/deepseek-v3.2-exp

DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It…

DeepSeekTool callingReasoningStructured output
Context
164K
Input /1M
$0.38
Output /1M
$0.57
DeepSeek V3.1 Terminus−40%
deepseek/deepseek-v3.1-terminus

DeepSeek-V3.1 Terminus is an update to DeepSeek V3.1 that maintains the model's original capabilities while addressing issues reported by users…

DeepSeekTool callingReasoningStructured output
Context
164K
Input /1M
$0.38
Output /1M
$1.40
DeepSeek V3.2−28%
deepseek/deepseek-v3.2

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance…

DeepSeekTool callingReasoningStructured outputPrompt caching
Context
164K
Input /1M
$0.39from $0.29
Output /1M
$0.59from $0.43
DeepSeek V3 0324
deepseek/deepseek-chat-v3-0324

DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It…

DeepSeekTool callingStructured outputPrompt caching
Context
164K
Input /1M
$0.41
Output /1M
$1.60
R1 0528
deepseek/deepseek-r1-0528

May 28th update to the original DeepSeek R1 Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B…

DeepSeekTool callingReasoningStructured outputPrompt caching
Context
164K
Input /1M
$0.70
Output /1M
$3.01
DeepSeek V4 Pro 0813−50%
deepseek/deepseek-v4-pro-0813

DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.

DeepSeekTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.77from $0.27
Output /1M
$7.00from $2.77
R1
deepseek/deepseek-r1

DeepSeek R1 is here: Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with…

DeepSeekTool callingReasoningStructured output
Context
64K
Input /1M
$0.98
Output /1M
$3.50
Qwen3.7 Flash
qwen/qwen3.7-flash

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer…

QwenImage inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$0.04
Output /1M
$0.18
Qwen3 30B A3B Instruct 2507−55%
qwen/qwen3-30b-a3b-instruct-2507

Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It…

QwenTool callingStructured output
Context
262K
Input /1M
$0.07
Output /1M
$0.27
Qwen3.5-Flash
qwen/qwen3.5-flash-02-23

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse…

QwenImage inTool callingReasoningStructured output
Context
1M
Input /1M
$0.09
Output /1M
$0.36
Qwen3 Coder 30B A3B Instruct
qwen/qwen3-coder-30b-a3b-instruct

Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for…

QwenTool callingStructured output
Context
262K
Input /1M
$0.10
Output /1M
$0.39
Qwen3 32B
qwen/qwen3-32b

Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It…

QwenTool callingReasoningStructured output
Context
131K
Input /1M
$0.11
Output /1M
$0.39
Qwen3 235B A22B Instruct 2507−75%
qwen/qwen3-235b-a22b-2507

Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B…

QwenTool callingStructured outputPrompt caching
Context
262K
Input /1M
$0.12
Output /1M
$0.49
Qwen3.5-9B
qwen/qwen3.5-9b

Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an…

QwenImage inTool callingReasoningStructured output
Context
262K
Input /1M
$0.14
Output /1M
$0.21
Qwen3 Next 80B A3B Instruct
qwen/qwen3-next-80b-a3b-instruct

Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking”…

QwenTool callingStructured outputPrompt caching
Context
262K
Input /1M
$0.14
Output /1M
$1.54
Qwen2.5 7B Instruct
qwen/qwen-2.5-7b-instruct

Qwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge…

QwenTool callingStructured output
Context
33K
Input /1M
$0.14
Output /1M
$0.28
Qwen3 VL 32B Instruct
qwen/qwen3-vl-32b-instruct

Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text…

QwenImage inTool callingStructured output
Context
131K
Input /1M
$0.15
Output /1M
$0.58
Qwen3 VL 8B Instruct
qwen/qwen3-vl-8b-instruct

Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across…

QwenImage inTool callingStructured output
Context
262K
Input /1M
$0.16
Output /1M
$0.64
Qwen3 8B
qwen/qwen3-8b

Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It…

QwenTool callingReasoningStructured output
Context
131K
Input /1M
$0.16
Output /1M
$0.64
Qwen3 Coder Next−40%
qwen/qwen3-coder-next

Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse MoE design…

QwenTool callingStructured outputPrompt caching
Context
262K
Input /1M
$0.17
Output /1M
$1.12
Qwen3 30B A3B
qwen/qwen3-30b-a3b

Qwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to excel in…

QwenTool callingReasoningStructured output
Context
131K
Input /1M
$0.17
Output /1M
$0.70
Qwen3 14B
qwen/qwen3-14b

Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It…

QwenTool callingReasoningStructured output
Context
131K
Input /1M
$0.17
Output /1M
$0.34
Qwen3.8 Omni Flash
qwen/qwen3.8-omni-flash

Qwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audio-video…

QwenImage inAudio inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$0.21
Output /1M
$0.66
Qwen3.8 Flash
qwen/qwen3.8-flash

Qwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document…

QwenImage inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$0.21
Output /1M
$0.66
Qwen3.6 35B A3B
qwen/qwen3.6-35b-a3b

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token…

QwenImage inTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.21
Output /1M
$1.40
Qwen3.5-35B-A3B
qwen/qwen3.5-35b-a3b

The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a…

QwenImage inTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.21
Output /1M
$1.40
Qwen3 VL 30B A3B Instruct
qwen/qwen3-vl-30b-a3b-instruct

Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct…

QwenImage inTool callingStructured output
Context
262K
Input /1M
$0.21
Output /1M
$0.84
Qwen3 Next 80B A3B Thinking
qwen/qwen3-next-80b-a3b-thinking

Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s…

QwenTool callingReasoningStructured output
Context
262K
Input /1M
$0.21
Output /1M
$1.68
Qwen3 VL 8B Thinking
qwen/qwen3-vl-8b-thinking

Qwen3-VL-8B-Thinking is the reasoning-optimized variant of the Qwen3-VL-8B multimodal model, designed for advanced visual and textual reasoning…

QwenImage inTool callingReasoningStructured output
Context
131K
Input /1M
$0.25
Output /1M
$2.94
Qwen3.6 Flash
qwen/qwen3.6-flash

Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context…

QwenImage inTool callingReasoningStructured output
Context
1M
Input /1M
$0.26
Output /1M
$1.57
Qwen3.5-27B
qwen/qwen3.5-27b

The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing…

QwenImage inTool callingReasoningStructured output
Context
262K
Input /1M
$0.27
Output /1M
$2.18
Qwen3 Coder Flash
qwen/qwen3-coder-flash

Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model…

QwenTool callingStructured outputPrompt caching
Context
1M
Input /1M
$0.27
Output /1M
$1.36
Qwen3 VL 30B A3B Thinking
qwen/qwen3-vl-30b-a3b-thinking

Qwen3-VL-30B-A3B-Thinking is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Thinking…

QwenImage inTool callingReasoningStructured output
Context
262K
Input /1M
$0.28
Output /1M
$3.36
Qwen3 30B A3B Thinking 2507
qwen/qwen3-30b-a3b-thinking-2507

Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step…

QwenTool callingReasoningStructured output
Context
82K
Input /1M
$0.28
Output /1M
$3.36
Qwen3 VL 235B A22B Instruct
qwen/qwen3-vl-235b-a22b-instruct

Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across images and…

QwenImage inTool callingStructured outputPrompt caching
Context
262K
Input /1M
$0.29
Output /1M
$2.66
Qwen3 235B A22B Thinking 2507
qwen/qwen3-235b-a22b-thinking-2507

Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It…

QwenTool callingReasoningStructured output
Context
131K
Input /1M
$0.32
Output /1M
$3.22
Qwen3.5-122B-A10B
qwen/qwen3.5-122b-a10b

The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse…

QwenImage inTool callingReasoningStructured output
Context
262K
Input /1M
$0.36
Output /1M
$2.91
Qwen3.5 Plus 2026-02-15
qwen/qwen3.5-plus-02-15

The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse…

QwenImage inTool callingReasoningStructured output
Context
1M
Input /1M
$0.36
Output /1M
$2.18
Qwen Plus 0728
qwen/qwen-plus-2025-07-28

Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, and cost…

QwenTool callingReasoningStructured output
Context
1M
Input /1M
$0.36
Output /1M
$1.09
Qwen-Plus
qwen/qwen-plus

Qwen-Plus, based on the Qwen2.5 foundation model, is a 131K context model with a balanced performance, speed, and cost combination.

QwenTool callingStructured outputPrompt caching
Context
1M
Input /1M
$0.36
Output /1M
$1.09
Qwen3.5 Plus 2026-04-20
qwen/qwen3.5-plus-20260420

Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text…

QwenImage inTool callingReasoningStructured output
Context
1M
Input /1M
$0.42
Output /1M
$2.52
Qwen3 Coder 480B A35B
qwen/qwen3-coder

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding…

QwenTool callingStructured outputPrompt caching
Context
262K
Input /1M
$0.42
Output /1M
$1.40
Qwen3.7 Plus
qwen/qwen3.7-plus

Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text…

QwenImage inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$0.45
Output /1M
$1.79
Qwen3.6 27B
qwen/qwen3.6-27b

Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal…

QwenImage inTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.45
Output /1M
$3.78
Qwen3.6 Plus
qwen/qwen3.6-plus

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong…

QwenImage inTool callingReasoningStructured output
Context
1M
Input /1M
$0.45
Output /1M
$2.73
Qwen2.5 72B Instruct
qwen/qwen-2.5-72b-instruct

Qwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more…

QwenTool callingStructured output
Context
33K
Input /1M
$0.50
Output /1M
$0.56
Qwen3 VL 235B A22B Thinking
qwen/qwen3-vl-235b-a22b-thinking

Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The…

QwenImage inTool callingReasoningStructured output
Context
131K
Input /1M
$0.56
Output /1M
$5.60
Qwen3.8 27B−25%
qwen/qwen3.8-27b

Qwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal…

QwenImage inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$0.59from $0.03
Output /1M
$3.57from $2.09
Qwen3 235B A22B
qwen/qwen3-235b-a22b

Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It supports…

QwenTool callingReasoningStructured output
Context
131K
Input /1M
$0.64
Output /1M
$2.55
Qwen3.5 397B A17B
qwen/qwen3.5-397b-a17b

The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a…

QwenImage inTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.77
Output /1M
$4.90
Qwen3 Coder Plus
qwen/qwen3-coder-plus

Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializing in…

QwenTool callingStructured outputPrompt caching
Context
1M
Input /1M
$0.91
Output /1M
$4.55
Qwen2.5 Coder 32B Instruct
qwen/qwen-2.5-coder-32b-instruct

Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). Qwen2.5-Coder brings the following…

QwenStructured output
Context
33K
Input /1M
$0.92
Output /1M
$1.40
Qwen3 Max Thinking
qwen/qwen3-max-thinking

Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step…

QwenTool callingReasoningStructured output
Context
262K
Input /1M
$1.09
Output /1M
$5.46
Qwen3 Max
qwen/qwen3-max

Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support…

QwenTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$1.09
Output /1M
$5.46
Qwen2.5 VL 72B Instruct
qwen/qwen2.5-vl-72b-instruct

Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing texts…

QwenImage inStructured outputPrompt caching
Context
128K
Input /1M
$1.12
Output /1M
$1.40
Qwen3.6 Max Preview
qwen/qwen3.6-max-preview

Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1…

QwenTool callingReasoningStructured output
Context
262K
Input /1M
$1.44
Output /1M
$8.63
Qwen3.7 Max
qwen/qwen3.7-max

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with…

QwenTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$2.06
Output /1M
$6.20
Qwen3.8 Max (0902)
qwen/qwen3.8-max-0902

Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that…

QwenImage inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$2.80
Output /1M
$8.40
Qwen3.8 2.4T A95B−20%
qwen/qwen3.8-2.4t-a95b

Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of Qwen3.8 Max, with 95 billion active…

QwenTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$2.80
Output /1M
$8.40
Qwen3.8 Max Prime
qwen/qwen3.8-max-prime

Qwen3.8 Max Prime is a higher-throughput variant of Qwen3.8 Max from Alibaba's Qwen team, served as a separate SKU at a higher price point. It…

QwenImage inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$5.60
Output /1M
$16.80
Kimi K2.5−5%
moonshotai/kimi-k2.5

Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm…

Moonshot AIImage inTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.63
Output /1M
$3.15
Kimi K2 0711
moonshotai/kimi-k2

Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32…

Moonshot AITool calling
Context
131K
Input /1M
$0.80
Output /1M
$3.22
Kimi K2 Thinking
moonshotai/kimi-k2-thinking

Kimi K2 Thinking is Moonshot AI’s most advanced open reasoning model to date, extending the K2 series into agentic, long-horizon reasoning. Built…

Moonshot AITool callingReasoningStructured output
Context
262K
Input /1M
$0.84
Output /1M
$3.50
Kimi K2 0905
moonshotai/kimi-k2-0905

Kimi K2 0905 is the September update of Kimi K2 0711. It is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI…

Moonshot AITool callingStructured output
Context
262K
Input /1M
$0.84
Output /1M
$3.50
Kimi K3−35%
moonshotai/kimi-k3

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and…

Moonshot AIImage inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.94from $0.82
Output /1M
$19.60from $13.65
Kimi K2.7 Code−25%
moonshotai/kimi-k2.7-code

MoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over…

Moonshot AIImage inTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.94
Output /1M
$4.69from $4.20
Kimi K2.6−37%
moonshotai/kimi-k2.6

Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent…

Moonshot AIImage inTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$1.33from $0.65
Output /1M
$5.60from $3.36
Kimi K3 (batch)
moonshotai/kimi-k3:batch

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and…

Moonshot AIImage inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$3.19
Output /1M
$15.96
GLM 5.2−54%
z-ai/glm-5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for…

Z.aiTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.03
Output /1M
$22.40from $2.52
GLM 5.3 Flash (batch)−50%
z-ai/glm-5.3-flash:batch

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear…

Z.aiImage inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.08
Output /1M
$0.28
GLM 4.7 Flash
z-ai/glm-4.7-flash

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding…

Z.aiTool callingReasoningStructured output
Context
200K
Input /1M
$0.08
Output /1M
$0.56
GLM 4.5 Air
z-ai/glm-4.5-air

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it…

Z.aiTool callingReasoningPrompt caching
Context
131K
Input /1M
$0.18
Output /1M
$1.19
GLM 5.3 Flash−50%
z-ai/glm-5.3-flash

GLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear…

Z.aiImage inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.21from $0.05
Output /1M
$0.70from $0.35
GLM 4.6V
z-ai/glm-4.6v

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed…

Z.aiImage inTool callingReasoningStructured outputPrompt caching
Context
131K
Input /1M
$0.42
Output /1M
$1.26
GLM 5.3 FlashX
z-ai/glm-5.3-flashx

GLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s…

Z.aiImage inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.52
Output /1M
$1.75
GLM 4.6
z-ai/glm-4.6

Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to…

Z.aiTool callingReasoningStructured outputPrompt caching
Context
205K
Input /1M
$0.60
Output /1M
$2.45
GLM 5.3 (batch)−38%
z-ai/glm-5.3:batch

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and…

Z.aiTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.63
Output /1M
$2.80
GLM 5−40%
z-ai/glm-5

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert…

Z.aiTool callingReasoningStructured outputPrompt caching
Context
205K
Input /1M
$0.84
Output /1M
$2.69
GLM 4.7−27%
z-ai/glm-4.7

GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step…

Z.aiTool callingReasoningStructured outputPrompt caching
Context
205K
Input /1M
$0.84from $0.56
Output /1M
$3.08from $2.45
GLM 4.5V
z-ai/glm-4.5v

GLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B…

Z.aiImage inTool callingReasoningStructured outputPrompt caching
Context
66K
Input /1M
$0.84
Output /1M
$2.52
GLM 4.5
z-ai/glm-4.5

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture…

Z.aiTool callingReasoningStructured outputPrompt caching
Context
131K
Input /1M
$0.84
Output /1M
$3.08
GLM 5V Turbo
z-ai/glm-5v-turbo

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles…

Z.aiImage inTool callingReasoningStructured outputPrompt caching
Context
203K
Input /1M
$1.68
Output /1M
$5.60
GLM 5 Turbo
z-ai/glm-5-turbo

GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It…

Z.aiTool callingReasoningStructured outputPrompt caching
Context
203K
Input /1M
$1.68
Output /1M
$5.60
GLM 5.3−70%
z-ai/glm-5.3

GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and…

Z.aiTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$1.96from $0.04
Output /1M
$6.16from $1.85
GLM 5.1−31%
z-ai/glm-5.1

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models…

Z.aiTool callingReasoningStructured outputPrompt caching
Context
205K
Input /1M
$1.96from $1.35
Output /1M
$6.16from $4.25
GLM 5.3 Prime
z-ai/glm-5.3-prime

GLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput through…

Z.aiTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$3.92
Output /1M
$12.32
Switchyard
nvidia/switchyard

Switchyard is an open-source model router that switches between multiple models to optimize the cost of requests. By default it will use OpenRouter…

NVIDIA
Context
1M
Input /1M
<$0.01
Output /1M
<$0.01
Nemotron 3 Nano 30B A3B
nvidia/nemotron-3-nano-30b-a3b

NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized…

NVIDIATool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.07
Output /1M
$0.28
Nemotron 3.5 Lightning−15%
nvidia/nemotron-3.5-lightning

NVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for…

NVIDIATool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.08from $0.05
Output /1M
$0.22
Nemotron 3 Super
nvidia/nemotron-3-super-120b-a12b

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in…

NVIDIATool callingReasoningStructured output
Context
262K
Input /1M
$0.11
Output /1M
$0.63
Nemotron 3.5 Content Safety
nvidia/nemotron-3.5-content-safety

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It…

NVIDIAImage inReasoning
Context
131K
Input /1M
$0.28
Output /1M
$0.28
Nemotron 3 Ultra
nvidia/nemotron-3-ultra-550b-a55b

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE)…

NVIDIATool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.70
Output /1M
$3.08
Command R7B (12-2024)
cohere/command-r7b-12-2024

Command R7B (12-2024) is a small, fast update of the Command R+ model, delivered in December 2024. It excels at RAG, tool use, agents, and similar…

CohereStructured output
Context
128K
Input /1M
$0.05
Output /1M
$0.21
Command R (08-2024)
cohere/command-r-08-2024

command-r-08-2024 is an update of the Command R with improved performance for multilingual retrieval-augmented generation (RAG) and tool use. More…

CohereTool callingStructured output
Context
128K
Input /1M
$0.21
Output /1M
$0.84
Command A+
cohere/command-a-plus

Command A+ is Cohere's flagship model for enterprise agentic workflows. It accepts text and image inputs with a 192K context window, supports…

CohereImage inTool callingReasoningStructured outputPrompt caching
Context
192K
Input /1M
$0.42
Output /1M
$2.10
Command A
cohere/command-a

Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual…

CohereStructured output
Context
256K
Input /1M
$3.50
Output /1M
$14.00
Command R+ (08-2024)
cohere/command-r-plus-08-2024

command-r-plus-08-2024 is an update of the Command R+ with roughly 50% higher throughput and 25% lower latencies as compared to the previous…

CohereTool callingStructured output
Context
128K
Input /1M
$3.50
Output /1M
$14.00
Sonar
perplexity/sonar

Sonar is lightweight, affordable, fast, and simple to use — now featuring citations and the ability to customize sources. It is designed for…

PerplexityImage in
Context
127K
Input /1M
$1.40
Output /1M
$1.40
Sonar Reasoning Pro
perplexity/sonar-reasoning-pro

Note: Sonar Pro pricing includes Perplexity search pricing. See details here Sonar Reasoning Pro is a premier reasoning model powered by DeepSeek…

PerplexityImage inReasoning
Context
128K
Input /1M
$2.80
Output /1M
$11.20
Sonar Deep Research
perplexity/sonar-deep-research

Sonar Deep Research is a research-focused model designed for multi-step retrieval, synthesis, and reasoning across complex topics. It autonomously…

PerplexityReasoning
Context
128K
Input /1M
$2.80
Output /1M
$11.20
Sonar Pro Search
perplexity/sonar-pro-search

Exclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system. It is designed…

PerplexityImage inReasoningStructured output
Context
200K
Input /1M
$4.20
Output /1M
$21.00
Sonar Pro
perplexity/sonar-pro

Note: Sonar Pro pricing includes Perplexity search pricing. See details here For enterprises seeking more advanced capabilities, the Sonar Pro API…

PerplexityImage in
Context
200K
Input /1M
$4.20
Output /1M
$21.00
Phi 4
microsoft/phi-4

Microsoft Research Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or…

MicrosoftStructured output
Context
16K
Input /1M
$0.10
Output /1M
$0.20
WizardLM-2 8x22B
microsoft/wizardlm-2-8x22b

WizardLM-2 8x22B is Microsoft AI's most advanced Wizard model. It demonstrates highly competitive performance compared to leading proprietary…

MicrosoftStructured output
Context
66K
Input /1M
$0.87
Output /1M
$0.87
Nova Micro 1.0
amazon/nova-micro-v1

Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the Amazon Nova family of models at a very low cost. With…

AmazonTool calling
Context
128K
Input /1M
$0.05
Output /1M
$0.20
Nova Lite 1.0
amazon/nova-lite-v1

Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to generate…

AmazonImage inTool calling
Context
300K
Input /1M
$0.08
Output /1M
$0.34
Nova 2 Lite
amazon/nova-2-lite-v1

Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2…

AmazonImage inFile inTool callingReasoning
Context
1M
Input /1M
$0.42
Output /1M
$3.50
Nova Pro 1.0
amazon/nova-pro-v1

Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide range of…

AmazonImage inTool calling
Context
300K
Input /1M
$1.12
Output /1M
$4.48
Nova Premier 1.0
amazon/nova-premier-v1

Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks and for use as the best teacher for distilling…

AmazonImage inTool callingPrompt caching
Context
1M
Input /1M
$3.50
Output /1M
$17.50
MiniMax-01
minimax/minimax-01

MiniMax-01 is a combines MiniMax-Text-01 for text generation and MiniMax-VL-01 for image understanding. It has 456 billion parameters, with 45.9…

MiniMaxImage in
Context
1.0M
Input /1M
$0.28
Output /1M
$1.54
MiniMax M2.7−30%
minimax/minimax-m2.7

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to…

MiniMaxTool callingReasoningStructured outputPrompt caching
Context
205K
Input /1M
$0.29
Output /1M
$1.18
MiniMax M2.5−10%
minimax/minimax-m2.5

MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working…

MiniMaxTool callingReasoningStructured outputPrompt caching
Context
205K
Input /1M
$0.38
Output /1M
$1.51from $1.33
MiniMax M3−50%
minimax/minimax-m3

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window…

MiniMaxImage inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.42from $0.32
Output /1M
$1.68from $1.34
MiniMax M2-her
minimax/minimax-m2-her

MiniMax M2-her is a dialogue-first large language model built for immersive roleplay, character-driven chat, and expressive multi-turn…

MiniMaxPrompt caching
Context
66K
Input /1M
$0.42
Output /1M
$1.68
MiniMax M2.1
minimax/minimax-m2.1

MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development…

MiniMaxTool callingReasoningStructured outputPrompt caching
Context
205K
Input /1M
$0.42
Output /1M
$1.68
MiniMax M2−15%
minimax/minimax-m2

MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated…

MiniMaxTool callingReasoningStructured output
Context
205K
Input /1M
$0.42from $0.36
Output /1M
$1.68from $1.43
MiniMax M1
minimax/minimax-m1

MiniMax-M1 is a large-scale, open-weight reasoning model designed for extended context and high-efficiency inference. It leverages a hybrid…

MiniMaxTool callingReasoning
Context
1M
Input /1M
$0.77
Output /1M
$3.08
Seed 1.6 Flash
bytedance-seed/seed-1.6-flash

Seed 1.6 Flash is an ultra-fast multimodal deep thinking model by ByteDance Seed, supporting both text and visual understanding. It features a 256k…

ByteDanceImage inTool callingReasoningStructured output
Context
262K
Input /1M
$0.10
Output /1M
$0.42
Seed-2.0-Mini
bytedance-seed/seed-2.0-mini

Seed-2.0-mini targets latency-sensitive, high-concurrency, and cost-sensitive scenarios, emphasizing fast response and flexible inference…

ByteDanceImage inTool callingReasoningStructured output
Context
262K
Input /1M
$0.14
Output /1M
$0.56
UI-TARS 7B
bytedance/ui-tars-1.5-7b

UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, web browsers, mobile systems…

ByteDanceImage inStructured outputPrompt caching
Context
128K
Input /1M
$0.14
Output /1M
$0.28
Seed-2.0-Lite
bytedance-seed/seed-2.0-lite

Seed-2.0-Lite is a versatile, cost‑efficient enterprise workhorse that delivers strong multimodal and agent capabilities while offering noticeably…

ByteDanceImage inTool callingReasoningStructured output
Context
262K
Input /1M
$0.35
Output /1M
$2.80
Seed 1.6
bytedance-seed/seed-1.6

Seed 1.6 is a general-purpose model released by the ByteDance Seed team. It incorporates multimodal capabilities and adaptive deep thinking with a…

ByteDanceImage inTool callingReasoningStructured output
Context
262K
Input /1M
$0.35
Output /1M
$2.80
Seed 2.1 Turbo
bytedance-seed/seed-2-1-turbo

Seed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software…

ByteDanceImage inTool callingReasoningStructured output
Context
262K
Input /1M
$0.70
Output /1M
$3.50
Seed-2.0-Code
bytedance-seed/seed-2.0-code

Seed 2.0 Code is a model from ByteDance Seed optimized for agentic coding. It is suited for frontend development, multilingual programming tasks…

ByteDanceImage inTool callingReasoningStructured output
Context
262K
Input /1M
$0.70
Output /1M
$4.20
Hy-MT2-1.8B
tencent/hy-mt2-1.8b

Hy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and…

Tencent
Context
8K
Input /1M
$0.06
Output /1M
$0.25
Hy-MT2-30B-A3B
tencent/hy-mt2-30b-a3b

Hy-MT2-30B-A3B is Tencent's flagship translation model in the Hy-MT2 family. It supports 33 language pairs and five Chinese dialect and…

TencentStructured output
Context
8K
Input /1M
$0.10
Output /1M
$0.41
Hy-MT2-7B
tencent/hy-mt2-7b

Hy-MT2-7B is a 7B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs…

TencentStructured output
Context
8K
Input /1M
$0.10
Output /1M
$0.41
Hy3
tencent/hy3

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows…

TencentTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.18
Output /1M
$0.74
Hunyuan A13B Instruct
tencent/hunyuan-a13b-instruct

Hunyuan-A13B is a 13B active parameter Mixture-of-Experts (MoE) language model developed by Tencent, with a total parameter count of 80B and…

TencentReasoningStructured output
Context
131K
Input /1M
$0.20
Output /1M
$0.80
Hy3 preview
tencent/hy3-preview

Hy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable…

TencentTool callingReasoningPrompt caching
Context
262K
Input /1M
$0.25
Output /1M
$0.84
Hy4 preview
tencent/hy4-preview

Tencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents…

TencentTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$1.17
Output /1M
$3.50
Aion 3.5 Mini
aion-labs/aion-3.5-mini

Aion 3.5 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It is the smaller, lower-cost…

AionLabsTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.98
Output /1M
$1.96
Aion-3.0-Mini
aion-labs/aion-3.0-mini

Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. It uses a collaborative…

AionLabsTool callingReasoningStructured outputPrompt caching
Context
131K
Input /1M
$0.98
Output /1M
$1.96
Aion-2.0
aion-labs/aion-2.0

Aion-2.0 is a variant of DeepSeek V3.2 optimized for immersive roleplaying and storytelling. It is particularly strong at introducing tension…

AionLabsTool callingReasoningStructured outputPrompt caching
Context
131K
Input /1M
$1.12
Output /1M
$2.24
Aion-RP 1.0 (8B)
aion-labs/aion-rp-llama-3.1-8b

Aion-RP-Llama-3.1-8B ranks the highest in the character evaluation portion of the RPBench-Auto benchmark, a roleplaying-specific variant of…

AionLabs
Context
33K
Input /1M
$1.12
Output /1M
$2.24
Aion 3.5
aion-labs/aion-3.5

Aion 3.5 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation…

AionLabsTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$4.20
Output /1M
$8.40
Aion-3.0
aion-labs/aion-3.0

Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation…

AionLabsTool callingReasoningStructured outputPrompt caching
Context
131K
Input /1M
$4.20
Output /1M
$8.40
Muse Spark 1.3 Contributor
meta/muse-spark-1.3-contributor

Muse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and…

MetaImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.14
Output /1M
$0.28
Muse Spark 1.2 Contributor
meta/muse-spark-1.2-contributor

Muse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s…

MetaImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$0.14
Output /1M
$0.28
Muse Glimmer 30B
meta/muse-glimmer-30b

Muse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous…

MetaImage inTool callingReasoningStructured outputPrompt caching
Context
131K
Input /1M
$0.49
Output /1M
$2.10
Muse Spark 1.3
meta/muse-spark-1.3

Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track…

MetaImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$1.75
Output /1M
$5.95
Muse Spark 1.2
meta/muse-spark-1.2

Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, and PDF documents, returns text…

MetaImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$1.75
Output /1M
$5.95
Muse Spark 1.1
meta/muse-spark-1.1

Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, and PDF documents and returns…

MetaImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$1.75
Output /1M
$5.95
MiMo-V2.6-Flash
xiaomi/mimo-v2.6-flash

MiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parameters and…

XiaomiImage inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$0.20
Output /1M
$0.39
MiMo-V2.5−15%
xiaomi/mimo-v2.5

MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing…

XiaomiImage inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$0.20from $0.17
Output /1M
$0.39from $0.33
MiMo-V2.6-Pro
xiaomi/mimo-v2.6-pro

MiMo-V2.6-Pro is the flagship foundation model developed by Xiaomi. Built at a scale of over 1T parameters, it is designed to push the ceiling of…

XiaomiImage inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$0.61
Output /1M
$1.22
MiMo-V2.5-Pro−30%
xiaomi/mimo-v2.5-pro

MiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and…

XiaomiTool callingReasoningStructured outputPrompt caching
Context
1.1M
Input /1M
$0.61from $0.43
Output /1M
$1.22from $0.85
MiMo-V2.6-Pro-UltraSpeed
xiaomi/mimo-v2.6-pro-ultraspeed

MiMo-V2.6-Pro-UltraSpeed is the fast speed edition of Xiaomi's flagship foundation model, MiMo-V2.6-Pro. Built from the same 1T MiMo-V2.6-Pro…

XiaomiImage inAudio inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$6.09
Output /1M
$12.18
Sakana Namazu
sakana/sakana-namazu

Sakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language and…

SakanaImage inFile inTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$1.33
Output /1M
$5.60
Fugu Max
sakana/fugu-max

Fugu Max is the cost-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent…

SakanaImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$2.80
Output /1M
$8.40
Fugu Ultra v2
sakana/fugu-ultra-v2

Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent…

SakanaImage inFile inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$7.00
Output /1M
$42.00
Fugu Ultra
sakana/fugu-ultra

Fugu Ultra is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent…

SakanaImage inTool callingReasoningStructured outputPrompt caching
Context
1M
Input /1M
$7.00
Output /1M
$42.00
Ling 3.0 Flash VL−72%
inclusionai/ling-3.0-flash-vl

Ling 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while…

InclusionaiImage inTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.03
Output /1M
$0.09
Ling 3.0 Flash−65%
inclusionai/ling-3.0-flash

Ling-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token. The model is designed…

InclusionaiTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.03
Output /1M
$0.09
Ling 3.0 Flash Fin−44%
inclusionai/ling-3.0-flash-fin

Ling 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B…

InclusionaiTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.06
Output /1M
$0.17
Hermes 3 70B Instruct
nousresearch/hermes-3-llama-3.1-70b

Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying…

Nous ResearchStructured output
Context
131K
Input /1M
$0.98
Output /1M
$0.98
Hermes 4 405B
nousresearch/hermes-4-405b

Hermes 4 is a large-scale reasoning model built on Meta-Llama-3.1-405B and released by Nous Research. It introduces a hybrid reasoning mode, where…

Nous ResearchReasoningStructured output
Context
131K
Input /1M
$1.40
Output /1M
$4.20
Hermes 3 405B Instruct
nousresearch/hermes-3-llama-3.1-405b

Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying…

Nous ResearchStructured output
Context
131K
Input /1M
$1.40
Output /1M
$1.40
Fusion
openrouter/fusion

Fusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel with web…

OpenRouter
Context
1M
Input /1M
<$0.01
Output /1M
<$0.01
Pareto Code Router
openrouter/pareto-code

The Pareto Router maintains a tiered shortlist of strong coding models, ranked by Artificial Analysis coding percentiles. Set mincodingscore…

OpenRouter
Context
2M
Input /1M
<$0.01
Output /1M
<$0.01
Body Builder (beta)
openrouter/bodybuilder

Transform your natural language requests into structured OpenRouter API request objects. Describe what you want to accomplish with AI models, and…

OpenRouter
Context
128K
Input /1M
<$0.01
Output /1M
<$0.01
Solar Mini 4−50%
upstage/solar-mini4

Solar Mini 4 is Upstage's compact, cost-efficient language model, a 35B-parameter mixture-of-experts with 3B active parameters and a 524K context…

UpstageTool callingReasoningStructured outputPrompt caching
Context
524K
Input /1M
$0.07
Output /1M
$0.28
Solar Pro 4−70%
upstage/solar-pro4

Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic…

UpstageTool callingReasoningStructured outputPrompt caching
Context
524K
Input /1M
$0.13
Output /1M
$0.50
Solar Pro 3
upstage/solar-pro-3

Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass…

UpstageTool callingReasoningStructured outputPrompt caching
Context
131K
Input /1M
$0.21
Output /1M
$0.84
Granite 4.0 Micro
ibm-granite/granite-4.0-h-micro

Granite-4.0-H-Micro is a 3B parameter from the Granite 4 family of models. These models are the latest in a series of models released by IBM. They…

Ibm GraniteStructured output
Context
131K
Input /1M
$0.02
Output /1M
$0.16
Granite 4.2 8B
ibm-granite/granite-4.2-8b

Granite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows…

Ibm GraniteTool callingReasoningStructured outputPrompt caching
Context
131K
Input /1M
$0.08
Output /1M
$0.35
Mercury 2.5−80%
inception/mercury-2.5

Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury…

InceptionTool callingReasoningStructured outputPrompt caching
Context
260K
Input /1M
$0.06
Output /1M
$0.21
Mercury 2
inception/mercury-2

Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2…

InceptionTool callingReasoningStructured outputPrompt caching
Context
128K
Input /1M
$0.35
Output /1M
$1.05
Schematron V2 Turbo
inference-net/schematron-v2-turbo

Schematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction…

Inference NetStructured outputPrompt caching
Context
128K
Input /1M
$0.04
Output /1M
$0.21
Schematron V2 Small
inference-net/schematron-v2-small

Schematron V2 Small is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes extraction quality for complex schemas and…

Inference NetStructured outputPrompt caching
Context
128K
Input /1M
$0.07
Output /1M
$0.32
Morph V3 Fast
morph/morph-v3-fast

Morph's fastest apply model for code edits. ~10,500 tokens/sec with 96% accuracy for rapid code transformations. The model requires the prompt to…

Morph
Context
82K
Input /1M
$1.12
Output /1M
$1.68
Morph V3 Large
morph/morph-v3-large

Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model requires…

MorphStructured output
Context
262K
Input /1M
$1.26
Output /1M
$2.66
Nex-N2.5-Mini
nex-agi/nex-n2.5-mini

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback…

Nex AgiImage inReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.03
Output /1M
$0.14
Nex-N2.5-Pro
nex-agi/nex-n2.5-pro

Nex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback…

Nex AgiImage inTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.10
Output /1M
$0.35
Perceptron Mk1.5
perceptron/perceptron-mk1.5

Perceptron Mk1.5 is Perceptron's embodied reasoning model for physical agents. It accepts text, image, video, and audio input, and answers with…

PerceptronImage inAudio inTool callingReasoningStructured output
Context
37K
Input /1M
$0.21
Output /1M
$2.10
Perceptron Mk1
perceptron/perceptron-mk1

Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning. It accepts image and video inputs…

PerceptronImage inReasoningStructured output
Context
33K
Input /1M
$0.21
Output /1M
$2.10
Laguna XS 2.1−40%
poolside/laguna-xs-2.1

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from Poolside and a step forward from their Laguna XS.2 model (released in…

PoolsideTool callingReasoningPrompt caching
Context
262K
Input /1M
$0.08
Output /1M
$0.17
Laguna S 2.1−10%
poolside/laguna-s-2.1

Laguna S 2.1 is the latest coding agent model from Poolside. Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2%…

PoolsideTool callingReasoningPrompt caching
Context
1.0M
Input /1M
$0.13
Output /1M
$0.25
Reka Edge
rekaai/reka-edge

Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. This model…

RekaaiImage inTool callingReasoningStructured output
Context
16K
Input /1M
$0.14
Output /1M
$0.14
Reka Flash 3
rekaai/reka-flash-3

Reka Flash 3 is a general-purpose, instruction-tuned large language model with 21 billion parameters, developed by Reka. It excels at general chat…

RekaaiReasoningStructured output
Context
66K
Input /1M
$0.14
Output /1M
$0.28
Relace Apply 3
relace/relace-apply-3

Relace Apply 3 is a specialized code-patching LLM that merges AI-suggested edits straight into your source files. It can apply updates from GPT-4o…

Relace
Context
256K
Input /1M
$1.19
Output /1M
$1.75
Relace Search
relace/relace-search

The relace-search model uses 4-12 viewfile and grep tools in parallel to explore a codebase and return relevant files to the user request. In…

RelaceTool callingStructured output
Context
256K
Input /1M
$1.40
Output /1M
$4.20
Step 3.5 Flash
stepfun/step-3.5-flash

Step 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it selectively…

StepfunTool callingReasoning
Context
262K
Input /1M
$0.14
Output /1M
$0.42
Step 3.7 Flash
stepfun/step-3.7-flash

Step 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision…

StepfunImage inTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.28
Output /1M
$1.61
Inkling Small
thinkingmachines/inkling-small

Inkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is…

ThinkingmachinesImage inAudio inTool callingReasoningPrompt caching
Context
524K
Input /1M
$0.63
Output /1M
$1.68
Inkling
thinkingmachines/inkling

Inkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is…

ThinkingmachinesImage inAudio inTool callingReasoningPrompt caching
Context
524K
Input /1M
$1.33
Output /1M
$5.67
Pareto 26.10 Preview
unbiased/pareto-26.10-preview

Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a…

UnbiasedImage inTool callingPrompt caching
Context
1.0M
Input /1M
$1.12
Output /1M
$4.48
Pareto
unbiased/pareto

Pareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a…

UnbiasedImage inTool callingStructured outputPrompt caching
Context
262K
Input /1M
$3.50
Output /1M
$10.50
Trinity Large Thinking
arcee-ai/trinity-large-thinking

Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic…

Arcee AITool callingReasoningPrompt caching
Context
262K
Input /1M
$0.35
Output /1M
$1.12
ERNIE 4.5 VL 424B A47B
baidu/ernie-4.5-vl-424b-a47b

ERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5 series, featuring 424B total parameters with 47B…

BaiduImage inReasoning
Context
123K
Input /1M
$0.59
Output /1M
$1.75
Ember-1
fireworks/ember-1

Ember-1 is a specialized reasoning model from Fireworks Research, built on Kimi K3. It is designed to make every token go further: it produces…

FireworksImage inTool callingReasoningStructured outputPrompt caching
Context
1.0M
Input /1M
$4.20
Output /1M
$21.00
KAT-Coder-Pro V2.5
kwaipilot/kat-coder-pro-v2.5

KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it…

KwaipilotTool callingStructured outputPrompt caching
Context
262K
Input /1M
$1.04
Output /1M
$4.14
LongCat 2.0−60%
meituan/longcat-2.0

LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding…

MeituanTool callingReasoningPrompt caching
Context
1.0M
Input /1M
$0.42
Output /1M
$1.68
Ternary Bonsai 2 27B
prism-ml/ternary-bonsai-2-27b

Bonsai 2 27B is a 27B-parameter reasoning model from PrismML derived from Qwen3.8-27B. It supports coding, mathematics, tool calling, and image…

Prism MlImage inTool callingReasoningStructured outputPrompt caching
Context
262K
Input /1M
$0.10
Output /1M
$0.70
Jev Router
typesafe/jev-router

Jev Router picks the best model and reasoning effort for each request, balancing quality, speed, and cost. It runs on Jev, TypeSafe's first System…

TypesafeImage inFile inAudio inTool callingReasoningStructured output
Context
1M
Input /1M
<$0.01
Output /1M
<$0.01
Palmyra X5
writer/palmyra-x5

Palmyra X5 is Writer's most advanced model, purpose-built for building and scaling AI agents across the enterprise. It delivers industry-leading…

Writer
Context
1.0M
Input /1M
$0.84
Output /1M
$8.40

Image generation 57 models billed per generated image, per megapixel or per output token — whichever the provider meters

Recraft
  • Recraft V4.1 Flashrecraft/recraft-v4.1-flash $0.0098/ image
  • Recraft V4 Stylesrecraft/recraft-v4-styles $0.049/ image
  • Recraft V4.1recraft/recraft-v4.1 $0.049/ image
  • Recraft V4.1 Utilityrecraft/recraft-v4.1-utility $0.049/ image
  • Recraft V3recraft/recraft-v3 $0.056/ image
  • Recraft V4recraft/recraft-v4 $0.056/ image
  • Recraft V4 Styles Vectorrecraft/recraft-v4-styles-vector $0.07/ image
  • Recraft V4 Vectorrecraft/recraft-v4-vector $0.112/ image
  • Recraft V4.1 Vectorrecraft/recraft-v4.1-vector $0.112/ image
  • Recraft V4 Styles Prorecraft/recraft-v4-styles-pro $0.14/ image
  • Recraft V4 Styles Pro Vectorrecraft/recraft-v4-styles-pro-vector $0.168/ image
  • Recraft V4.1 Prorecraft/recraft-v4.1-pro $0.294/ image
  • Recraft V4.1 Utility Prorecraft/recraft-v4.1-utility-pro $0.294/ image
  • Recraft V4 Prorecraft/recraft-v4-pro $0.35/ image
  • Recraft V4 Pro Vectorrecraft/recraft-v4-pro-vector $0.42/ image
  • Recraft V4.1 Pro Vectorrecraft/recraft-v4.1-pro-vector $0.42/ image
OpenAI
  • GPT Image 1 Miniopenai/gpt-image-1-mini $11.20/ 1M tokens
  • GPT-5 Image Miniopenai/gpt-5-image-mini $11.20/ 1M tokens
  • GPT Image 2openai/gpt-image-2 $42.00/ 1M tokens
  • GPT Image 2.5 Flareopenai/gpt-image-2.5-flare $42.00/ 1M tokens
  • GPT Image 2.5 Sunburstopenai/gpt-image-2.5-sunburst $42.00/ 1M tokens
  • GPT-5.4 Image 2openai/gpt-5.4-image-2 $42.00/ 1M tokens
  • GPT Image 1openai/gpt-image-1 $56.00/ 1M tokens
  • GPT-5 Imageopenai/gpt-5-image $56.00/ 1M tokens
Google
  • Nano Banana (Gemini 2.5 Flash Image)google/gemini-2.5-flash-image $42.00/ 1M tokens
  • Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image)google/gemini-3.1-flash-lite-image $42.00/ 1M tokens
  • Nano Banana 2 (Gemini 3.1 Flash Image Preview)google/gemini-3.1-flash-image-preview $84.00/ 1M tokens
  • Nano Banana 2 (Gemini 3.1 Flash Image)google/gemini-3.1-flash-image $84.00/ 1M tokens
  • Nano Banana Pro (Gemini 3 Pro Image Preview)google/gemini-3-pro-image-preview $84.00/ 1M tokens
  • Nano Banana Pro (Gemini 3 Pro Image)google/gemini-3-pro-image $168.00/ 1M tokens
Black Forest Labs
  • FLUX.2 Klein 4Bblack-forest-labs/flux.2-klein-4b $0.0196/ megapixel
  • FLUX.2 Problack-forest-labs/flux.2-pro $0.042/ megapixel
  • FLUX.3 Imageblack-forest-labs/flux-3-image $0.0574/ image
  • FLUX.2 Flexblack-forest-labs/flux.2-flex $0.084/ megapixel
  • FLUX.2 Maxblack-forest-labs/flux.2-max $0.098/ megapixel
ByteDance
  • Seedream 5.0 Flashbytedance-seed/seedream-5-0-flash $0.0252/ image
  • Seedream 5.0 Litebytedance-seed/seedream-5-0-lite $0.049/ image
  • Seedream 4.5bytedance-seed/seedream-4.5 $0.056/ image
  • Seedream 5.0 Probytedance-seed/seedream-5-0-pro $0.063/ image
Microsoft
  • MAI-Image-2.6 Flashmicrosoft/mai-image-2.6-flash $26.60/ 1M tokens
  • MAI-Image-2.6microsoft/mai-image-2.6 $53.20/ 1M tokens
  • MAI-Image-2.5microsoft/mai-image-2.5 $65.80/ 1M tokens
  • MAI-Image-2.5 Promicrosoft/mai-image-2.5-pro $151.20/ 1M tokens
Sourceful
  • Riverflow V2.5 Fastsourceful/riverflow-v2.5-fast $0.0266/ image
  • Riverflow V2 Fastsourceful/riverflow-v2-fast $0.028/ image
  • Riverflow V2.5 Prosourceful/riverflow-v2.5-pro $0.182/ image
  • Riverflow V2 Prosourceful/riverflow-v2-pro $0.21/ image
Krea
  • Krea 2 Largekrea/krea-2-large metered
  • Krea 2 Mediumkrea/krea-2-medium metered
  • Krea 2 Medium Turbokrea/krea-2-medium-turbo metered
Inclusionai
  • Ming Image 0.1 Designinclusionai/ming-image-0.1-design Free/ 1M tokens
  • Ming Image 0.1 Design Layerinclusionai/ming-image-0.1-design-layer Free/ 1M tokens
Qwen
  • Qwen Image 3qwen/qwen-image-3 $0.042/ image
  • Qwen Image 3 Proqwen/qwen-image-3-pro $0.056/ image
xAI
  • Grok Imagine Image 2.0x-ai/grok-imagine-image-2.0 $0.056/ image
  • Grok Imagine Image Qualityx-ai/grok-imagine-image-quality $0.07/ image
Meta
  • Muse Imagemeta/muse-image metered

Video generation 30 models billed per second of output — or per token where the provider meters that way; a range spans the model’s resolutions

Alibaba
  • Wan 3.0audioalibaba/wan-3.0 $0.07–$0.28/ second
  • Wan 3.0 Primeaudioalibaba/wan-3.0-prime $0.0952–$0.392/ second
  • HappyHorse 1.0alibaba/happyhorse-1.0 $0.138–$0.237/ second
  • HappyHorse 1.1alibaba/happyhorse-1.1 $0.138–$0.179/ second
  • Wan 2.7audioalibaba/wan-2.7 $0.14/ second
  • Wan 2.6audioalibaba/wan-2.6 metered
ByteDance
  • Seedance 1.5 Proaudiobytedance/seedance-1-5-pro $1.68–$3.36/ 1M tokens
  • Seedance 2.0 Miniaudiobytedance/seedance-2.0-mini $2.94–$4.90/ 1M tokens
  • Seedance 2.0audiobytedance/seedance-2.0 $3.36–$10.78/ 1M tokens
  • Seedance 2.0 Fastaudiobytedance/seedance-2.0-fast $3.46–$5.88/ 1M tokens
  • Seedance 2.5audiobytedance/seedance-2.5 $8.96–$14.98/ 1M tokens
Black Forest Labs
  • FLUX Video Upscaleblack-forest-labs/flux-video-upscale $0.105–$0.147/ MP-second
  • FLUX Video Editblack-forest-labs/flux-video-edit metered
  • FLUX.3 Videoaudioblack-forest-labs/flux-3-video metered
Google
  • Veo 3.1 Liteaudiogoogle/veo-3.1-lite $0.042–$0.112/ second
  • Veo 3.1 Fastaudiogoogle/veo-3.1-fast $0.112–$0.42/ second
  • Veo 3.1audiogoogle/veo-3.1 $0.28–$0.84/ second
Kuaishou
  • Video v3.0 Standardaudiokwaivgi/kling-v3.0-std $0.118–$0.176/ second
  • Video O1audiokwaivgi/kling-video-o1 $0.157/ second
  • Video v3.0 Proaudiokwaivgi/kling-v3.0-pro $0.157–$0.235/ second
MiniMax
  • H3 Maxminimax/hailuo-3-max $0.07–$0.112/ second
  • Hailuo 2.3minimax/hailuo-2.3 $0.114/ second
  • H3audiominimax/hailuo-3 $0.182/ second
HeyGen
  • HeyGen Videoheygen/heygen-video-1 $0.028–$0.042/ second
  • Avatar IVheygen/avatar-iv $0.07/ second
Runway
  • Aleph 2.0runway/aleph-2 metered
  • Gen-4.5runway/gen-4.5 metered
xAI
  • Grok Imagine Videox-ai/grok-imagine-video metered
  • Grok Imagine Video 1.5x-ai/grok-imagine-video-1.5 metered
OpenAI
  • Sora 2 Proaudioopenai/sora-2-pro $0.42–$0.7/ second

Music generation 17 models full tracks from a text brief — flat per track, or by the second

MiniMax
  • MiniMax Music 3minimax/music-3 $0.0028/ second
  • MiniMax Music 1.5fal-ai/minimax-music/v1.5 $0.042/ generation
  • MiniMax Music 2.0fal-ai/minimax-music/v2 $0.042/ generation
  • MiniMax Music 2.5fal-ai/minimax-music/v2.5 $0.21/ generation
  • MiniMax Music 2.6fal-ai/minimax-music/v2.6 $0.21/ generation
Stability AI
  • Stable Audio 3 Small Musicfal-ai/stable-audio-3/small/music/text-to-audio $0.0304/ generation
  • Stable Audio 3 Mediumfal-ai/stable-audio-3/medium/text-to-audio $0.0526/ generation
  • Stable Audio 3 Medium Basefal-ai/stable-audio-3/medium/base/text-to-audio $0.0671/ generation
  • Stable Audio 2.5fal-ai/stable-audio-25/text-to-audio $0.28/ generation
Google DeepMind
  • Lyria 3fal-ai/lyria3 $0.056/ generation
  • Lyria 3 Profal-ai/lyria3/pro $0.112/ generation
  • Lyria 2fal-ai/lyria2 $0.14/ generation
ACE-Step
  • ACE-Stepfal-ai/ace-step $0.00028/ second
CassetteAI
  • CassetteAI Musiccassetteai/music-generator $0.028/ minute
DiffRhythm
  • DiffRhythmfal-ai/diffrhythm $0.0014/ second
ElevenLabs
  • ElevenLabs Musicfal-ai/elevenlabs/music $0.84/ minute
Sonilo
  • Sonilo v1.1sonilo/v1.1/text-to-music $0.0035/ second

Sound effects 3 models one short sound from a description — impacts, ambiences, foley

CassetteAI
  • CassetteAI Sound Effectscassetteai/sound-effects-generator $0.014/ generation
ElevenLabs
  • ElevenLabs Sound Effects v2fal-ai/elevenlabs/sound-effects/v2 $0.0028/ second
Stability AI
  • Stable Audio 3 Small SFXfal-ai/stable-audio-3/small/sfx/text-to-audio $0.0288/ generation

Voices (text to speech) 5 models billed per 1,000 characters submitted

ElevenLabs
  • ElevenLabs Turbo v2.5fal-ai/elevenlabs/tts/turbo-v2.5 $0.07/ 1K chars
  • ElevenLabs Multilingual v2fal-ai/elevenlabs/tts/multilingual-v2 $0.14/ 1K chars
  • ElevenLabs v3fal-ai/elevenlabs/tts/eleven-v3 $0.14/ 1K chars
Kokoro
  • Kokoro American Englishfal-ai/kokoro/american-english $0.028/ 1K chars
  • Kokoro British Englishfal-ai/kokoro/british-english $0.028/ 1K chars

3D model generation 10 models image to mesh; priced by the options you pick, so each model shows its range

Tencent
  • Hunyuan 3D v3.1 Rapidfal-ai/hunyuan-3d/v3.1/rapid/image-to-3d $0.322–$0.532/ mesh
  • Hunyuan3D v3fal-ai/hunyuan3d-v3/image-to-3d $0.322–$1.26/ mesh
  • Hunyuan 3D v3.1 Profal-ai/hunyuan-3d/v3.1/pro/image-to-3d $0.532–$1.16/ mesh
Tripo3D
  • Tripo H3.1tripo3d/h3.1/image-to-3d $0.28–$0.91/ mesh
  • Tripo v2.5tripo3d/tripo/v2.5/image-to-3d $0.28–$0.7/ mesh
  • Tripo P1tripo3d/p1/image-to-3d $0.56–$0.7/ mesh
Hyper3D
  • Rodin v2.5 Fastfal-ai/hyper3d/rodin/v2.5/fast $0.14/ mesh
  • Rodin v2.5fal-ai/hyper3d/rodin/v2.5 $0.56–$1.68/ mesh
Meshy
  • Meshy v7meshy/v7/image-to-3d $1.12–$2.41/ mesh
Trellis
  • Trellis 2fal-ai/trellis-2 $0.35–$0.49/ mesh
Availability & pricing refresh with each deploy from the live provider catalogue; the app always shows current per-model pricing before you spend. Free and NSFW/roleplay models are excluded. Every price on this page is USD at the developer-API rate — see the API pricing page for how billing works.
Reading the media prices. Chat is per token; everything else is per unit of output, and the unit is the provider’s own. An image can be metered per generated image, per megapixel or per output token; video per second, with a range spanning the model’s resolutions — ask for 480p rather than 1080p and you pay the low end. Music is either flat per track or by the second, sound effects likewise, voices per 1,000 characters of text you submit, and a mesh by the options you pick, which is why 3D models show a range rather than a rate. We quote each in its own unit rather than converting: converting would mean inventing the resolution, duration or token count that your actual request decides. A row marked metered is billed at whatever the provider charges for the call — it is not free, it simply publishes no rate we can quote ahead of time.
About the discounts. A −40% chip means the model’s provider had a promotional discount running when this page was last built. The developer API bills what the provider actually charged for your call, plus our flat markup — so a live promotion reaches your bill on its own, with nothing to claim and no code to enter. The from $0.00 figure is the cheapest endpoint serving that model; which endpoint answers a given request is the provider’s routing decision, so treat it as a floor, not a quote. Input and output are quoted separately, and discounted separately — on a reasoning model the output rate is usually what decides the bill. The promotions are the providers’ own and can end at any time.