near/z-ai/glm-5.3-flashGLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear…
- Context
- —
- Input /1M
- $0.21
- Output /1M
- $0.70
One account, every major provider. Switch models mid-conversation in the
app, or call any of them by ID through the OpenAI-compatible API — point
your client at https://api.privateer.pro/v1 and set
model. Chat, image, video, music, sound effects, voices and
3D meshes — every one of them with its ID and what it costs.
No models match . Try a provider, model name or ID.
On sale now 54 models up to 90% off · in / out, USD per 1M tokens · provider promotions, billed through at the discounted rate
deepseek/deepseek-v4-flash-0731
deepseek/deepseek-v4-pro
inception/mercury-2.5
qwen/qwen3-235b-a22b-2507
inclusionai/ling-3.0-flash-vl
z-ai/glm-5.3
upstage/solar-pro4
deepseek/deepseek-v4-flash
inclusionai/ling-3.0-flash
meituan/longcat-2.0
qwen/qwen3-30b-a3b-instruct-2507
z-ai/glm-5.2
deepseek/deepseek-v4.1-flash
deepseek/deepseek-v4-flash-vision-exp
upstage/solar-mini4
google/gemini-3.8-flash
google/gemini-3.8-flash:batch
z-ai/glm-5.3-flash
z-ai/glm-5.3-flash:batch
google/gemini-3.7-flash
google/gemini-3.7-flash:batch
deepseek/deepseek-v4-pro-0813
openai/gpt-5.6-sol-pro
openai/gpt-5.6-sol-pro:batch
openai/gpt-5.6-sol
openai/gpt-5.6-sol:batch
minimax/minimax-m3
inclusionai/ling-3.0-flash-fin
poolside/laguna-xs-2.1
z-ai/glm-5
qwen/qwen3-coder-next
deepseek/deepseek-v3.1-terminus
z-ai/glm-5.3:batch
moonshotai/kimi-k2.6
moonshotai/kimi-k3
z-ai/glm-5.1
deepseek/deepseek-v4.1-flash:batch
xiaomi/mimo-v2.5-pro
minimax/minimax-m2.7
deepseek/deepseek-v3.2
z-ai/glm-4.7
mistralai/mistral-nemo
qwen/qwen3.8-27b
moonshotai/kimi-k2.7-code
google/gemma-4-26b-a4b-it
meta-llama/llama-3.3-70b-instruct
qwen/qwen3.8-2.4t-a95b
nvidia/nemotron-3.5-lightning
xiaomi/mimo-v2.5
minimax/minimax-m2
poolside/laguna-s-2.1
minimax/minimax-m2.5
deepseek/deepseek-chat
moonshotai/kimi-k2.5
Chat & reasoning 423 models 49 providers · search, filter or sort the whole list
Confidential 21 of these models run inside an attested hardware enclave — Intel TDX, NVIDIA Confidential Computing or AMD SEV-SNP, served by Tinfoil, Phala and NEAR AI. The prompt is decrypted only in there, and the app verifies the hardware attestation on every response; “verified” is claimed only where your own device does the attesting. Context and feature tags are not published per host, so those cells read “—”. What confidential compute means →
near/z-ai/glm-5.3-flashGLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear…
near/Qwen/Qwen3-VL-30B-A3B-InstructQwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct…
near/Qwen/Qwen3.6-35B-A3B-FP8
near/Qwen/Qwen3.8-27BQwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal…
near/moonshotai/kimi-k2.6Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent…
near/moonshotai/kimi-k3Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and…
phala/nvidia/nemotron-3.5-lightningNVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for…
phala/qwen/qwen-2.5-7b-instructQwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge…
phala/phala/gemma-4-26b-a4b-uncensored
phala/qwen/qwen3.8-27bQwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal…
phala/meta/muse-glimmer-30bMuse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous…
phala/phala/qwen3.8-27b-uncensored
phala/qwen/qwen3.6-27bQwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal…
phala/z-ai/glm-5.3GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and…
tinfoil/gpt-oss-120bgpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and…
tinfoil/gemma4-31b
tinfoil/glm-5-3-flash
tinfoil/deepseek-v4-1-flash
tinfoil/llama3-3-70b
tinfoil/glm-5-3
tinfoil/kimi-k3Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and…
openai/gpt-oss-20bgpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture…
openai/gpt-oss-20b:batchgpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture…
openai/gpt-5-nano:batchGPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency…
openai/gpt-oss-120b:batchgpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and…
openai/gpt-oss-120bgpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and…
openai/gpt-6-luna-pro:batchGPT-6 Luna Pro is the same underlying model as GPT-6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks…
openai/gpt-6-luna:batchGPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and…
openai/gpt-5-nanoGPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency…
openai/gpt-4.1-nano:batchFor tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a…
openai/gpt-oss-safeguard-20bgpt-oss-safeguard-20b is a safety reasoning model from OpenAI built upon gpt-oss-20b. This open-weight, 21B-parameter Mixture-of-Experts (MoE)…
openai/gpt-4o-mini:batchGPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs. As their most advanced small model…
openai/gpt-6-luna-proGPT-6 Luna Pro is the same underlying model as GPT-6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks…
openai/gpt-6-lunaGPT-6 Luna is the fast, cost-efficient model in OpenAI's GPT-6 series, positioned below GPT-6 Sol. It is suited for high-volume and…
openai/gpt-5.6-luna-pro:batchGPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks…
openai/gpt-5.6-luna:batchGPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat…
openai/gpt-5.4-nano:batchGPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It…
openai/gpt-4.1-nanoFor tasks that demand low latency, GPT‑4.1 nano is the fastest and cheapest model in the GPT-4.1 series. It delivers exceptional performance at a…
openai/gpt-5-mini:batchGPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and…
openai/gpt-4o-miniGPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs. As their most advanced small model…
openai/gpt-4o-mini-2024-07-18GPT-4o mini is OpenAI's newest model after GPT-4 Omni, supporting both text and image inputs with text outputs. As their most advanced small model…
openai/gpt-5.6-luna-proGPT-5.6 Luna Pro is the same underlying model as GPT-5.6 Luna, served with reasoning.mode set to pro for higher-quality responses on complex tasks…
openai/gpt-5.6-lunaGPT-5.6 Luna is a fast, cost-efficient model in OpenAI's GPT-5.6 series. It is suited for high-volume, latency-sensitive tasks such as chat…
openai/gpt-5.4-nanoGPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It…
openai/gpt-4.1-mini:batchGPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million…
openai/gpt-5.1-codex-miniGPT-5.1-Codex-Mini is a smaller and faster version of GPT-5.1-Codex
openai/gpt-5-miniGPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and…
openai/gpt-3.5-turbo:batchGPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional…
openai/gpt-5.4-mini:batchGPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text…
openai/gpt-4.1-miniGPT-4.1 Mini is a mid-sized model delivering performance competitive with GPT-4o at substantially lower latency and cost. It retains a 1 million…
openai/gpt-3.5-turboGPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional…
openai/o4-mini:batchOpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and…
openai/o3-mini:batchOpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding…
openai/gpt-5.1:batchGPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a…
openai/gpt-5:batchGPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex…
openai/gpt-5.4-miniGPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text…
openai/gpt-5.2:batchGPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses…
openai/gpt-6-sol-pro:batchGPT-6 Sol Pro is the same underlying model as GPT-6 Sol, served with reasoning.mode set to pro for higher-quality responses on complex tasks. Cost…
openai/gpt-6-sol:batchGPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna…
openai/gpt-5.6-terra-pro:batchGPT-5.6 Terra Pro is the same underlying model as GPT-5.6 Terra, served with reasoning.mode set to pro for higher-quality responses on complex…
openai/gpt-5.6-terra:batchGPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is…
openai/gpt-5.6-sol-pro:batchGPT-5.6 Sol Pro is the same underlying model as GPT-5.6 Sol, served with reasoning.mode set to pro for higher-quality responses on complex tasks…
openai/gpt-5.6-sol:batchGPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is…
openai/o3:batcho3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also…
openai/gpt-4.1:batchGPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context…
openai/gpt-3.5-turbo-0613GPT-3.5 Turbo is OpenAI's fastest model. It can understand and generate natural language or code, and is optimized for chat and traditional…
openai/o4-mini-highOpenAI o4-mini-high is the same model as o4-mini with reasoningeffort set to high. OpenAI o4-mini is a compact reasoning model in the o-series…
openai/o4-miniOpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining strong multimodal and…
openai/o3-mini-highOpenAI o3-mini-high is the same model as o3-mini with reasoningeffort set to high. o3-mini is a cost-efficient language model optimized for STEM…
openai/o3-miniOpenAI o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and coding…
openai/gpt-5.4:batchGPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K…
openai/gpt-5.1-codex-maxGPT-5.1-Codex-Max is OpenAI’s latest agentic coding model, designed for long-running, high-context software development tasks. It is based on an…
openai/gpt-5.1GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a…
openai/gpt-5.1-codexGPT-5.1-Codex is a specialized version of GPT-5.1 optimized for software engineering and coding workflows. It is designed for both interactive…
openai/gpt-5GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex…
openai/gpt-4o:batchGPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level…
openai/gpt-3.5-turbo-instructThis model is a variant of GPT-3.5 Turbo tuned for instructional prompts and omitting chat-related optimizations. Training data: up to Sep 2021.
openai/gpt-5.3-codexGPT-5.3-Codex is OpenAI’s most advanced agentic coding model, combining the frontier software engineering performance of GPT-5.2-Codex with the…
openai/gpt-5.2-codexGPT-5.2-Codex is an upgraded version of GPT-5.1-Codex optimized for software engineering and coding workflows. It is designed for both interactive…
openai/gpt-5.2-chatGPT-5.2 Chat (AKA Instant) is the fast, lightweight member of the 5.2 family, optimized for low-latency chat while retaining strong general…
openai/gpt-5.2GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses…
openai/gpt-6.1-sol-proGPT-6.1 Sol Pro is the same underlying model as GPT-6.1 Sol, served with reasoning.mode set to pro for higher-quality responses on complex tasks…
openai/gpt-6.1-solGPT-6.1 Sol is an upgrade to GPT-6 Sol from OpenAI, positioned below the flagship GPT-6 Astra in the GPT-6 series. It is suited for agentic coding…
openai/gpt-6-sol-proGPT-6 Sol Pro is the same underlying model as GPT-6 Sol, served with reasoning.mode set to pro for higher-quality responses on complex tasks. Cost…
openai/gpt-6-solGPT-6 Sol is the cost-efficient high-end model in OpenAI's GPT-6 series, positioned below the flagship GPT-6 Astra and above the fast GPT-6 Luna…
openai/gpt-5.6-terra-proGPT-5.6 Terra Pro is the same underlying model as GPT-5.6 Terra, served with reasoning.mode set to pro for higher-quality responses on complex…
openai/gpt-5.6-terraGPT-5.6 Terra is a balanced model in OpenAI's GPT-5.6 series, positioned between the flagship Sol tier and the cost-efficient Luna tier. It is…
openai/gpt-5.6-sol-proGPT-5.6 Sol Pro is the same underlying model as GPT-5.6 Sol, served with reasoning.mode set to pro for higher-quality responses on complex tasks…
openai/gpt-5.6-solGPT-5.6 Sol is the flagship model in OpenAI's GPT-5.6 series. It is suited for complex reasoning, coding, and agentic workflows, and is…
openai/o3o3 is a well-rounded and powerful model across domains. It sets a new standard for math, science, coding, and visual reasoning tasks. It also…
openai/gpt-4.1GPT-4.1 is a flagship large language model optimized for advanced instruction following, real-world software engineering, and long-context…
openai/gpt-5.5:batchGPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability…
openai/gpt-5.4GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K…
openai/gpt-4o-2024-11-20The 2024-11-20 version of GPT-4o offers a leveled-up creative writing ability with more natural, engaging, and tailored writing to improve…
openai/gpt-4o-2024-08-06The 2024-08-06 version of GPT-4o offers improved performance in structured outputs, with the ability to supply a JSON schema in the responeformat…
openai/gpt-4oGPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level…
openai/gpt-3.5-turbo-16kThis model offers four times the context length of gpt-3.5-turbo, allowing it to support approximately 20 pages of text in a single request at a…
openai/gpt-6-astra:batchGPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research…
openai/gpt-6-astra-pro:batchGPT-6 Astra Pro is the same underlying model as GPT-6 Astra, served with reasoning.mode set to pro for higher-quality responses on complex tasks…
openai/gpt-chat-latestGPT Chat Latest points to OpenAI's stable API alias chat-latest that always resolves to the latest Instant chat model used in ChatGPT. As OpenAI…
openai/gpt-5.5GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability…
openai/gpt-4o-2024-05-13GPT-4o ("o" for "omni") is OpenAI's latest AI model, supporting both text and image inputs with text outputs. It maintains the intelligence level…
openai/gpt-4-turbo:batchThe latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December…
openai/gpt-5-pro:batchGPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex…
openai/gpt-6-astraGPT-6 Astra is OpenAI's flagship model for demanding end-to-end work. It is suited for advanced analysis, software engineering, deep research…
openai/gpt-6-astra-proGPT-6 Astra Pro is the same underlying model as GPT-6 Astra, served with reasoning.mode set to pro for higher-quality responses on complex tasks…
openai/gpt-4-turboThe latest GPT-4 Turbo model with vision capabilities. Vision requests can now use JSON mode and function calling. Training data: up to December…
openai/gpt-5.2-pro:batchGPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is…
openai/gpt-5.5-pro:batchGPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token…
openai/gpt-5.4-pro:batchGPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex…
openai/gpt-5-proGPT-5 Pro is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex…
openai/o1The latest and strongest model family from OpenAI, o1 is designed to spend more time thinking before responding. The o1 model series is trained…
openai/o3-proThe o-series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o3-pro model uses…
openai/gpt-5.2-proGPT-5.2 Pro is OpenAI’s most advanced model, offering major improvements in agentic coding and long context performance over GPT-5 Pro. It is…
openai/gpt-5.5-proGPT-5.5 Pro is OpenAI’s high-capability model optimized for deep reasoning and accuracy on complex, high-stakes workloads. It features a 1M+ token…
openai/gpt-5.4-proGPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning capabilities for complex…
openai/gpt-4OpenAI's flagship model, GPT-4 is a large-scale multimodal language model capable of solving difficult problems with greater accuracy than previous…
openai/o1-proThe o1 series of models are trained with reinforcement learning to think before they answer and perform complex reasoning. The o1-pro model uses…
anthropic/claude-haiku-4.5:batchClaude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of…
anthropic/claude-sonnet-5.5:batchClaude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, succeeding Claude Sonnet 5 as a direct upgrade. It is especially…
anthropic/claude-sonnet-5:batchSonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports…
anthropic/claude-haiku-4.5Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of…
anthropic/claude-sonnet-4.6:batchSonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at…
anthropic/claude-sonnet-4.5:batchClaude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers…
anthropic/claude-sonnet-5.5Claude Sonnet 5.5 is Anthropic's Sonnet-class model for well-scoped everyday work, succeeding Claude Sonnet 5 as a direct upgrade. It is especially…
anthropic/claude-opus-5.5:batchClaude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is…
anthropic/claude-sonnet-5Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports…
anthropic/claude-opus-5:batchClaude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end…
anthropic/claude-opus-4.8:batchClaude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text…
anthropic/claude-opus-4.7:batchOpus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic…
anthropic/claude-opus-4.6:batchOpus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows…
anthropic/claude-opus-4.5:batchClaude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer…
anthropic/claude-sonnet-4.6Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at…
anthropic/claude-sonnet-4.5Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers…
anthropic/claude-sonnet-4Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved…
anthropic/claude-opus-5.5Claude Opus 5.5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, succeeding Claude Opus 5. It is…
anthropic/claude-fable-5.1:batchClaude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and…
anthropic/claude-opus-5Claude Opus 5 is Anthropic’s flagship model for demanding reasoning, coding, and long-horizon agentic work. It is particularly strong at end-to-end…
anthropic/claude-fable-5:batchClaude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs…
anthropic/claude-opus-4.8Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text…
anthropic/claude-opus-4.7Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic…
anthropic/claude-opus-4.6Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows…
anthropic/claude-opus-4.5Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer…
anthropic/claude-opus-4.1:batchClaude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It…
anthropic/claude-fable-5.1Claude Fable 5.1 improves on Claude Fable 5 across the board, with the biggest gains in agentic coding, long-running agentic workflows, and…
anthropic/claude-fable-5Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs…
anthropic/claude-opus-4.1Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It…
google/gemini-2.5-flash-lite:batchGemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers…
google/gemma-3-4b-itGemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over…
google/gemma-3-12b-itGemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over…
google/gemma-4-26b-a4b-itGemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate…
google/gemma-3-27b-itGemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over…
google/gemma-4-31b-itGemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token…
google/gemini-2.5-flash-liteGemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers…
google/gemini-3.1-flash-lite:batchGemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image…
google/gemini-3.5-flash-lite:batchGemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused…
google/gemini-2.5-flash:batchGemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific…
google/gemini-3.1-flash-liteGemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image…
google/gemini-3.1-flash-lite-previewGemini 3.1 Flash Lite Preview is Google's high-efficiency model optimized for high-volume use cases. It outperforms Gemini 2.5 Flash Lite on…
google/gemini-3-flash-preview:batchGemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It…
google/gemini-3.5-flash-liteGemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused…
google/gemini-2.5-flashGemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific…
google/gemini-3.8-flash:batchGemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and…
google/gemini-3.7-flash:batchGemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks…
google/gemini-3.6-flash:batchGemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce…
google/gemini-3-flash-previewGemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It…
google/gemini-2.5-pro:batchGemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs…
google/gemma-2-27b-itGemma 2 27B by Google is an open model built from the same research and technology used to create the Gemini models. Gemma models are well-suited…
google/gemini-3.8-flashGemini 3.8 Flash is Google's most intelligent Flash model with significant gains from 3.7 Flash across software engineering, agentic tasks, and…
google/gemini-3.7-flashGemini 3.7 Flash is a multimodal model from Google for fast agentic workflows, coding, and complex multi-step reasoning. It is designed for tasks…
google/gemini-3.6-flashGemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce…
google/gemini-3.5-flash:batchGemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is…
google/gemini-3.1-pro-preview:batchGemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability…
google/gemini-2.5-proGemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs…
google/gemini-2.5-pro-previewGemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs…
google/gemini-3.5-flashGemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is…
google/gemini-3.1-pro-preview-customtoolsGemini 3.1 Pro Preview Custom Tools is a variant of Gemini 3.1 Pro that improves tool selection behavior by preventing overuse of a general bash…
google/gemini-3.1-pro-previewGemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability…
x-ai/grok-build-0.1Grok Build 0.1 is SpaceXAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs…
x-ai/grok-4.3:batchGrok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows…
x-ai/grok-4.3Grok 4.3 is a reasoning model from SpaceXAI. It accepts text and image inputs with text output, and is suited for agentic workflows…
x-ai/grok-4.20-multi-agentGrok 4.20 Multi-Agent is a variant of SpaceXAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel…
x-ai/grok-4.20Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest…
x-ai/grok-4.7Grok 4.7 is SpaceXAI's flagship model for coding, agentic tasks, and knowledge work, succeeding Grok 4.6. It is particularly strong at long-running…
x-ai/grok-4.6Grok 4.6 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM. It is succeeded by Grok 4.7.
x-ai/grok-4.5Grok 4.5 is a model from SpaceXAI with frontier performance on coding, knowledge work, and STEM.
meta-llama/llama-3.2-1b-instructLlama 3.2 1B is a 1-billion-parameter language model focused on efficiently performing natural language tasks, such as summarization, dialogue, and…
meta-llama/llama-3.2-3b-instructLlama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue…
meta-llama/llama-3.1-8b-instructMeta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has…
meta-llama/llama-4-scoutLlama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of…
meta-llama/llama-guard-4-12bLlama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions…
meta-llama/llama-4-maverickLlama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with…
meta-llama/llama-3.3-70b-instructThe Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The…
meta-llama/llama-3.1-70b-instructMeta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high quality…
mistralai/mistral-nemoA 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting…
mistralai/mistral-small-24b-instruct-2501Mistral Small 3 is a 24B-parameter language model optimized for low-latency performance across common AI tasks. Released under the Apache 2.0…
mistralai/mistral-small-2603:batchMistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single…
mistralai/ministral-8b-2512:batchA balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
mistralai/mistral-small-3.2-24b-instructMistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduction, and…
mistralai/ministral-3b-2512The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.
mistralai/voxtral-small-24b-2507Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text…
mistralai/mistral-small-2603Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single…
mistralai/ministral-8b-2512A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
mistralai/codestral-2508:batchMistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as…
mistralai/ministral-14b-2512The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small…
mistralai/mistral-medium-3.1:batchMistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver…
mistralai/mistral-sabaMistral Saba is a 24B-parameter language model specifically designed for the Middle East and South Asia, delivering accurate and contextually…
mistralai/mistral-large-2512:batchMistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B…
mistralai/codestral-2508Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as…
mistralai/mistral-small-3.1-24b-instructMistral Small 3.1 24B Instruct is an upgraded variant of Mistral Small 3 (2501), featuring 24 billion parameters with advanced multimodal…
mistralai/devstral-2512Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer model…
mistralai/mistral-medium-3.1Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver…
mistralai/mistral-medium-3Mistral Medium 3 is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced…
mistralai/mistral-large-2512Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B…
mistralai/mistral-medium-3-5:batchMistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed…
mistralai/mistral-medium-3-5Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed…
mistralai/mistral-large-2407This is Mistral AI's flagship model, Mistral Large 2 (version mistral-large-2407). It's a proprietary weights-available model and excels at…
mistralai/mixtral-8x22b-instructMistral's official instruct fine-tuned version of Mixtral 8x22B. It uses 39B active parameters out of 141B, offering unparalleled cost efficiency…
mistralai/mistral-largeThis is Mistral AI's flagship model, Mistral Large 2 (version mistral-large-2407). It's a proprietary weights-available model and excels at…
deepseek/deepseek-v4.1-flashDeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED)…
deepseek/deepseek-v4-flash-0731DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained…
deepseek/deepseek-v4-flashDeepSeek V4 Flash is an efficiency-optimized Mixture-of-Experts model from DeepSeek with 284B total parameters and 13B activated parameters…
deepseek/deepseek-v4.1-flash:batchDeepSeek V4.1 Flash is a sparse mixture-of-experts model from DeepSeek, and the first built on the company's Causal Encoder-Decoder (CED)…
deepseek/deepseek-v4-proDeepSeek V4 Pro is a large-scale Mixture-of-Experts model from DeepSeek with 1.6T total parameters and 49B activated parameters, supporting a…
deepseek/deepseek-v4-flash-vision-expDeepSeek V4 Flash Vision Exp is an experimental vision-enabled version of DeepSeek V4 Flash 0731 from DeepSeek, adding image understanding while…
deepseek/deepseek-chat-v3.1DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt…
deepseek/deepseek-chatDeepSeek-V3 is the latest model from the DeepSeek team, building upon the instruction following and coding abilities of the previous versions…
deepseek/deepseek-v3.2-expDeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It…
deepseek/deepseek-v3.1-terminusDeepSeek-V3.1 Terminus is an update to DeepSeek V3.1 that maintains the model's original capabilities while addressing issues reported by users…
deepseek/deepseek-v3.2DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance…
deepseek/deepseek-chat-v3-0324DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It…
deepseek/deepseek-r1-0528May 28th update to the original DeepSeek R1 Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B…
deepseek/deepseek-v4-pro-0813DeepSeek V4 Pro 0813 is a large-scale mixture-of-experts model from DeepSeek. This is the GA release of DeepSeek V4 Pro.
deepseek/deepseek-r1DeepSeek R1 is here: Performance on par with OpenAI o1, but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with…
qwen/qwen3.7-flashQwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer…
qwen/qwen3-30b-a3b-instruct-2507Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It…
qwen/qwen3.5-flash-02-23The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse…
qwen/qwen3-coder-30b-a3b-instructQwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for…
qwen/qwen3-32bQwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It…
qwen/qwen3-235b-a22b-2507Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B…
qwen/qwen3.5-9bQwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an…
qwen/qwen3-next-80b-a3b-instructQwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking”…
qwen/qwen-2.5-7b-instructQwen2.5 7B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more knowledge…
qwen/qwen3-vl-32b-instructQwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text…
qwen/qwen3-vl-8b-instructQwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across…
qwen/qwen3-8bQwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It…
qwen/qwen3-coder-nextQwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse MoE design…
qwen/qwen3-30b-a3bQwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to excel in…
qwen/qwen3-14bQwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It…
qwen/qwen3.8-omni-flashQwen3.8 Omni Flash is an omni-modal reasoning model from Alibaba, the first Qwen model built around agentic capabilities with native audio-video…
qwen/qwen3.8-flashQwen3.8 Flash is a multimodal reasoning model from Alibaba. It is suited for coding assistance, agentic workflows, visual understanding, document…
qwen/qwen3.6-35b-a3bQwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token…
qwen/qwen3.5-35b-a3bThe Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a…
qwen/qwen3-vl-30b-a3b-instructQwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct…
qwen/qwen3-next-80b-a3b-thinkingQwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s…
qwen/qwen3-vl-8b-thinkingQwen3-VL-8B-Thinking is the reasoning-optimized variant of the Qwen3-VL-8B multimodal model, designed for advanced visual and textual reasoning…
qwen/qwen3.6-flashQwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context…
qwen/qwen3.5-27bThe Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing…
qwen/qwen3-coder-flashQwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model…
qwen/qwen3-vl-30b-a3b-thinkingQwen3-VL-30B-A3B-Thinking is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Thinking…
qwen/qwen3-30b-a3b-thinking-2507Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step…
qwen/qwen3-vl-235b-a22b-instructQwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across images and…
qwen/qwen3-235b-a22b-thinking-2507Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It…
qwen/qwen3.5-122b-a10bThe Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse…
qwen/qwen3.5-plus-02-15The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse…
qwen/qwen-plus-2025-07-28Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, and cost…
qwen/qwen-plusQwen-Plus, based on the Qwen2.5 foundation model, is a 131K context model with a balanced performance, speed, and cost combination.
qwen/qwen3.5-plus-20260420Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text…
qwen/qwen3-coderQwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding…
qwen/qwen3.7-plusQwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text…
qwen/qwen3.6-27bQwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal…
qwen/qwen3.6-plusQwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong…
qwen/qwen-2.5-72b-instructQwen2.5 72B is the latest series of Qwen large language models. Qwen2.5 brings the following improvements upon Qwen2: - Significantly more…
qwen/qwen3-vl-235b-a22b-thinkingQwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The…
qwen/qwen3.8-27bQwen3.8 27B is an open-weight dense vision-language model from Qwen. It is suited for coding, professional workflows, research, multimodal…
qwen/qwen3-235b-a22bQwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It supports…
qwen/qwen3.5-397b-a17bThe Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a…
qwen/qwen3-coder-plusQwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializing in…
qwen/qwen-2.5-coder-32b-instructQwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). Qwen2.5-Coder brings the following…
qwen/qwen3-max-thinkingQwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step…
qwen/qwen3-maxQwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support…
qwen/qwen2.5-vl-72b-instructQwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing texts…
qwen/qwen3.6-max-previewQwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1…
qwen/qwen3.7-maxQwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with…
qwen/qwen3.8-max-0902Qwen3.8 Max 0902 is an updated snapshot of Qwen3.8 Max from Alibaba's Qwen team. It is a 2.4-trillion-parameter mixture-of-experts model that…
qwen/qwen3.8-2.4t-a95bQwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of Qwen3.8 Max, with 95 billion active…
qwen/qwen3.8-max-primeQwen3.8 Max Prime is a higher-throughput variant of Qwen3.8 Max from Alibaba's Qwen team, served as a separate SKU at a higher price point. It…
moonshotai/kimi-k2.5Kimi K2.5 is Moonshot AI's native multimodal model, delivering state-of-the-art visual coding capability and a self-directed agent swarm paradigm…
moonshotai/kimi-k2Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32…
moonshotai/kimi-k2-thinkingKimi K2 Thinking is Moonshot AI’s most advanced open reasoning model to date, extending the K2 series into agentic, long-horizon reasoning. Built…
moonshotai/kimi-k2-0905Kimi K2 0905 is the September update of Kimi K2 0711. It is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI…
moonshotai/kimi-k3Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and…
moonshotai/kimi-k2.7-codeMoonshotAI: Kimi K2.7 Code is a coding-focused model in Moonshot AI's Kimi K2 family, built to complete end-to-end programming tasks reliably over…
moonshotai/kimi-k2.6Kimi K2.6 is Moonshot AI's next-generation multimodal model, designed for long-horizon coding, coding-driven UI/UX generation, and multi-agent…
moonshotai/kimi-k3:batchKimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and…
z-ai/glm-5.2GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for…
z-ai/glm-5.3-flash:batchGLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear…
z-ai/glm-4.7-flashAs a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding…
z-ai/glm-4.5-airGLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it…
z-ai/glm-5.3-flashGLM-5.3-Flash is a native multimodal model from Z.ai. It is suited for efficient coding and long-horizon agent tasks. Its hybrid sparse and linear…
z-ai/glm-4.6vGLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed…
z-ai/glm-5.3-flashxGLM-5.3-FlashX is the high-speed variant of Z.ai's GLM-5.3-Flash, a native multimodal model delivering inference speeds of up to 200 tokens/s…
z-ai/glm-4.6Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to…
z-ai/glm-5.3:batchGLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and…
z-ai/glm-5GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert…
z-ai/glm-4.7GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step…
z-ai/glm-4.5vGLM-4.5V is a vision-language foundation model for multimodal agent applications. Built on a Mixture-of-Experts (MoE) architecture with 106B…
z-ai/glm-4.5GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture…
z-ai/glm-5v-turboGLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles…
z-ai/glm-5-turboGLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It…
z-ai/glm-5.3GLM-5.3 is a large-scale reasoning model from Z.ai, built for complex software engineering and long-horizon agent tasks. It supports text input and…
z-ai/glm-5.1GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models…
z-ai/glm-5.3-primeGLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput through…
nvidia/switchyardSwitchyard is an open-source model router that switches between multiple models to optimize the cost of requests. By default it will use OpenRouter…
nvidia/nemotron-3-nano-30b-a3bNVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized…
nvidia/nemotron-3.5-lightningNVIDIA Nemotron 3.5 Lightning is an open mixture-of-experts model from NVIDIA, with 3B active parameters out of 30B total. It is suited for…
nvidia/nemotron-3-super-120b-a12bNVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in…
nvidia/nemotron-3.5-content-safetyNVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It…
nvidia/nemotron-3-ultra-550b-a55bNVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE)…
cohere/command-r7b-12-2024Command R7B (12-2024) is a small, fast update of the Command R+ model, delivered in December 2024. It excels at RAG, tool use, agents, and similar…
cohere/command-r-08-2024command-r-08-2024 is an update of the Command R with improved performance for multilingual retrieval-augmented generation (RAG) and tool use. More…
cohere/command-a-plusCommand A+ is Cohere's flagship model for enterprise agentic workflows. It accepts text and image inputs with a 192K context window, supports…
cohere/command-aCommand A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual…
cohere/command-r-plus-08-2024command-r-plus-08-2024 is an update of the Command R+ with roughly 50% higher throughput and 25% lower latencies as compared to the previous…
perplexity/sonarSonar is lightweight, affordable, fast, and simple to use — now featuring citations and the ability to customize sources. It is designed for…
perplexity/sonar-reasoning-proNote: Sonar Pro pricing includes Perplexity search pricing. See details here Sonar Reasoning Pro is a premier reasoning model powered by DeepSeek…
perplexity/sonar-deep-researchSonar Deep Research is a research-focused model designed for multi-step retrieval, synthesis, and reasoning across complex topics. It autonomously…
perplexity/sonar-pro-searchExclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system. It is designed…
perplexity/sonar-proNote: Sonar Pro pricing includes Perplexity search pricing. See details here For enterprises seeking more advanced capabilities, the Sonar Pro API…
microsoft/phi-4Microsoft Research Phi-4 is designed to perform well in complex reasoning tasks and can operate efficiently in situations with limited memory or…
microsoft/wizardlm-2-8x22bWizardLM-2 8x22B is Microsoft AI's most advanced Wizard model. It demonstrates highly competitive performance compared to leading proprietary…
amazon/nova-micro-v1Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the Amazon Nova family of models at a very low cost. With…
amazon/nova-lite-v1Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to generate…
amazon/nova-2-lite-v1Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2…
amazon/nova-pro-v1Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide range of…
amazon/nova-premier-v1Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks and for use as the best teacher for distilling…
minimax/minimax-01MiniMax-01 is a combines MiniMax-Text-01 for text generation and MiniMax-VL-01 for image understanding. It has 456 billion parameters, with 45.9…
minimax/minimax-m2.7MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to…
minimax/minimax-m2.5MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working…
minimax/minimax-m3MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window…
minimax/minimax-m2-herMiniMax M2-her is a dialogue-first large language model built for immersive roleplay, character-driven chat, and expressive multi-turn…
minimax/minimax-m2.1MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development…
minimax/minimax-m2MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated…
minimax/minimax-m1MiniMax-M1 is a large-scale, open-weight reasoning model designed for extended context and high-efficiency inference. It leverages a hybrid…
bytedance-seed/seed-1.6-flashSeed 1.6 Flash is an ultra-fast multimodal deep thinking model by ByteDance Seed, supporting both text and visual understanding. It features a 256k…
bytedance-seed/seed-2.0-miniSeed-2.0-mini targets latency-sensitive, high-concurrency, and cost-sensitive scenarios, emphasizing fast response and flexible inference…
bytedance/ui-tars-1.5-7bUI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, web browsers, mobile systems…
bytedance-seed/seed-2.0-liteSeed-2.0-Lite is a versatile, cost‑efficient enterprise workhorse that delivers strong multimodal and agent capabilities while offering noticeably…
bytedance-seed/seed-1.6Seed 1.6 is a general-purpose model released by the ByteDance Seed team. It incorporates multimodal capabilities and adaptive deep thinking with a…
bytedance-seed/seed-2-1-turboSeed 2.1 Turbo is a multimodal model from ByteDance Seed for coding and long-horizon agent workflows. It is suited for end-to-end software…
bytedance-seed/seed-2.0-codeSeed 2.0 Code is a model from ByteDance Seed optimized for agentic coding. It is suited for frontend development, multilingual programming tasks…
tencent/hy-mt2-1.8bHy-MT2-1.8B is a compact 1.8B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and…
tencent/hy-mt2-30b-a3bHy-MT2-30B-A3B is Tencent's flagship translation model in the Hy-MT2 family. It supports 33 language pairs and five Chinese dialect and…
tencent/hy-mt2-7bHy-MT2-7B is a 7B-parameter translation model from Tencent. It supports 33 language pairs and five Chinese dialect and minority-language pairs…
tencent/hy3Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows…
tencent/hunyuan-a13b-instructHunyuan-A13B is a 13B active parameter Mixture-of-Experts (MoE) language model developed by Tencent, with a total parameter count of 80B and…
tencent/hy3-previewHy3 preview is a high-efficiency Mixture-of-Experts model from Tencent designed for agentic workflows and production use. It supports configurable…
tencent/hy4-previewTencent: Hy4 preview is a mixture-of-experts model from Tencent, with 49B active parameters out of 770B total. It is designed for coding agents…
aion-labs/aion-3.5-miniAion 3.5 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It is the smaller, lower-cost…
aion-labs/aion-3.0-miniAion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. It uses a collaborative…
aion-labs/aion-2.0Aion-2.0 is a variant of DeepSeek V3.2 optimized for immersive roleplaying and storytelling. It is particularly strong at introducing tension…
aion-labs/aion-rp-llama-3.1-8bAion-RP-Llama-3.1-8B ranks the highest in the character evaluation portion of the RPBench-Auto benchmark, a roleplaying-specific variant of…
aion-labs/aion-3.5Aion 3.5 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation…
aion-labs/aion-3.0Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation…
meta/muse-spark-1.3-contributorMuse Spark 1.3 Contributor is the cost-efficient contributor tier of Meta’s multimodal reasoning model for experimentation, learning, and…
meta/muse-spark-1.2-contributorMuse Spark 1.2 contributor tier is a reasoning model from Meta designed for developers who want to start building at an even lower cost. It’s…
meta/muse-glimmer-30bMuse Glimmer 30B is a dense, open-weight multimodal model from Meta Superintelligence Labs, distilled from Muse Spark and optimized for autonomous…
meta/muse-spark-1.3Muse Spark 1.3 is a multimodal reasoning model from Meta for long-running agentic, multi-agent, and coding workflows. It is designed to keep track…
meta/muse-spark-1.2Muse Spark 1.2 is a reasoning model from Meta, designed for complex agentic tasks. It accepts text, images, video, and PDF documents, returns text…
meta/muse-spark-1.1Muse Spark 1.1 is a multimodal reasoning model from Meta, built for agentic tasks. It accepts text, images, video, and PDF documents and returns…
xiaomi/mimo-v2.6-flashMiMo-V2.6-Flash is an open-source foundation model developed by Xiaomi. Built on a Mixture-of-Experts architecture with 309B total parameters and…
xiaomi/mimo-v2.5MiMo-V2.5 is a native omnimodal model by Xiaomi. It delivers Pro-level agentic performance at roughly half the inference cost, while surpassing…
xiaomi/mimo-v2.6-proMiMo-V2.6-Pro is the flagship foundation model developed by Xiaomi. Built at a scale of over 1T parameters, it is designed to push the ceiling of…
xiaomi/mimo-v2.5-proMiMo-V2.5-Pro is Xiaomi’s flagship model, delivering strong performance in general agentic capabilities, complex software engineering, and…
xiaomi/mimo-v2.6-pro-ultraspeedMiMo-V2.6-Pro-UltraSpeed is the fast speed edition of Xiaomi's flagship foundation model, MiMo-V2.6-Pro. Built from the same 1T MiMo-V2.6-Pro…
sakana/sakana-namazuSakana Namazu is a Japanese-specialized reasoning model from Sakana AI, based on Kimi K2.6 with additional training for Japanese language and…
sakana/fugu-maxFugu Max is the cost-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent…
sakana/fugu-ultra-v2Fugu Ultra v2 is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent…
sakana/fugu-ultraFugu Ultra is the higher-performance model in Sakana AI's Fugu family. Rather than a single monolithic model, Fugu is a learned multi-agent…
inclusionai/ling-3.0-flash-vlLing 3.0 Flash VL builds on Ling 3.0 Flash (124B total / 5.5B active MoE from InclusionAI), further strengthening its language capabilities while…
inclusionai/ling-3.0-flashLing-3.0-flash is a 124B-parameter Mixture-of-Experts (MoE) model, with approximately 5.1B parameters activated per token. The model is designed…
inclusionai/ling-3.0-flash-finLing 3.0 Flash Fin is a finance-focused mixture-of-experts model from InclusionAI, built on Ling 3.0 Flash with 5.1B active parameters out of 124B…
nousresearch/hermes-3-llama-3.1-70bHermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying…
nousresearch/hermes-4-405bHermes 4 is a large-scale reasoning model built on Meta-Llama-3.1-405B and released by Nous Research. It introduces a hybrid reasoning mode, where…
nousresearch/hermes-3-llama-3.1-405bHermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying…
openrouter/fusionFusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel with web…
openrouter/pareto-codeThe Pareto Router maintains a tiered shortlist of strong coding models, ranked by Artificial Analysis coding percentiles. Set mincodingscore…
openrouter/bodybuilderTransform your natural language requests into structured OpenRouter API request objects. Describe what you want to accomplish with AI models, and…
upstage/solar-mini4Solar Mini 4 is Upstage's compact, cost-efficient language model, a 35B-parameter mixture-of-experts with 3B active parameters and a 524K context…
upstage/solar-pro4Solar Pro 4 is Upstage's cost-efficient large language model, featuring a 524K context window. It is built for long-horizon tasks and agentic…
upstage/solar-pro-3Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass…
ibm-granite/granite-4.0-h-microGranite-4.0-H-Micro is a 3B parameter from the Granite 4 family of models. These models are the latest in a series of models released by IBM. They…
ibm-granite/granite-4.2-8bGranite 4.2 8B is a dense reasoning model from IBM. It is suited for mathematics, code generation, multilingual dialogue, and agentic workflows…
inception/mercury-2.5Mercury 2.5 is the fastest reasoning LLM, and the latest diffusion LLM (dLLM) from Inception. Instead of generating tokens sequentially, Mercury…
inception/mercury-2Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2…
inference-net/schematron-v2-turboSchematron V2 Turbo is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes throughput for high-volume extraction…
inference-net/schematron-v2-smallSchematron V2 Small is a 3B-parameter HTML-to-JSON extraction model from Inference.net. It prioritizes extraction quality for complex schemas and…
morph/morph-v3-fastMorph's fastest apply model for code edits. ~10,500 tokens/sec with 96% accuracy for rapid code transformations. The model requires the prompt to…
morph/morph-v3-largeMorph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model requires…
nex-agi/nex-n2.5-miniNex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback…
nex-agi/nex-n2.5-proNex-N2.5 is an agentic model built to turn goals into working, verified outcomes. Its core strength is agentic coding within a visual feedback…
perceptron/perceptron-mk1.5Perceptron Mk1.5 is Perceptron's embodied reasoning model for physical agents. It accepts text, image, video, and audio input, and answers with…
perceptron/perceptron-mk1Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning. It accepts image and video inputs…
poolside/laguna-xs-2.1Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from Poolside and a step forward from their Laguna XS.2 model (released in…
poolside/laguna-s-2.1Laguna S 2.1 is the latest coding agent model from Poolside. Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2%…
rekaai/reka-edgeReka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. This model…
rekaai/reka-flash-3Reka Flash 3 is a general-purpose, instruction-tuned large language model with 21 billion parameters, developed by Reka. It excels at general chat…
relace/relace-apply-3Relace Apply 3 is a specialized code-patching LLM that merges AI-suggested edits straight into your source files. It can apply updates from GPT-4o…
relace/relace-searchThe relace-search model uses 4-12 viewfile and grep tools in parallel to explore a codebase and return relevant files to the user request. In…
stepfun/step-3.5-flashStep 3.5 Flash is StepFun's most capable open-source foundation model. Built on a sparse Mixture of Experts (MoE) architecture, it selectively…
stepfun/step-3.7-flashStep 3.7 Flash is StepFun's latest high-efficiency multimodal Mixture-of-Experts model. It pairs a 196B-parameter language backbone with a vision…
thinkingmachines/inkling-smallInkling Small is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 12B active parameters out of 276B total. It is…
thinkingmachines/inklingInkling is an open-weight multimodal mixture-of-experts model from Thinking Machines Lab, with 41B active parameters out of 975B total. It is…
unbiased/pareto-26.10-previewPareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a…
unbiased/paretoPareto is a multimodal composite model built for research, coding, and agentic workflows, while delivering frontier-level performance across a…
arcee-ai/trinity-large-thinkingTrinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic…
baidu/ernie-4.5-vl-424b-a47bERNIE-4.5-VL-424B-A47B is a multimodal Mixture-of-Experts (MoE) model from Baidu’s ERNIE 4.5 series, featuring 424B total parameters with 47B…
fireworks/ember-1Ember-1 is a specialized reasoning model from Fireworks Research, built on Kimi K3. It is designed to make every token go further: it produces…
kwaipilot/kat-coder-pro-v2.5KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it…
meituan/longcat-2.0LongCat 2.0 is a sparse mixture-of-experts language model from Meituan, with 48B active parameters out of 1.6T total. It is suited for coding…
prism-ml/ternary-bonsai-2-27bBonsai 2 27B is a 27B-parameter reasoning model from PrismML derived from Qwen3.8-27B. It supports coding, mathematics, tool calling, and image…
typesafe/jev-routerJev Router picks the best model and reasoning effort for each request, balancing quality, speed, and cost. It runs on Jev, TypeSafe's first System…
writer/palmyra-x5Palmyra X5 is Writer's most advanced model, purpose-built for building and scaling AI agents across the enterprise. It delivers industry-leading…
Image generation 57 models billed per generated image, per megapixel or per output token — whichever the provider meters
recraft/recraft-v4.1-flash
$0.0098/ image
recraft/recraft-v4-styles
$0.049/ image
recraft/recraft-v4.1
$0.049/ image
recraft/recraft-v4.1-utility
$0.049/ image
recraft/recraft-v3
$0.056/ image
recraft/recraft-v4
$0.056/ image
recraft/recraft-v4-styles-vector
$0.07/ image
recraft/recraft-v4-vector
$0.112/ image
recraft/recraft-v4.1-vector
$0.112/ image
recraft/recraft-v4-styles-pro
$0.14/ image
recraft/recraft-v4-styles-pro-vector
$0.168/ image
recraft/recraft-v4.1-pro
$0.294/ image
recraft/recraft-v4.1-utility-pro
$0.294/ image
recraft/recraft-v4-pro
$0.35/ image
recraft/recraft-v4-pro-vector
$0.42/ image
recraft/recraft-v4.1-pro-vector
$0.42/ image
openai/gpt-image-1-mini
$11.20/ 1M tokens
openai/gpt-5-image-mini
$11.20/ 1M tokens
openai/gpt-image-2
$42.00/ 1M tokens
openai/gpt-image-2.5-flare
$42.00/ 1M tokens
openai/gpt-image-2.5-sunburst
$42.00/ 1M tokens
openai/gpt-5.4-image-2
$42.00/ 1M tokens
openai/gpt-image-1
$56.00/ 1M tokens
openai/gpt-5-image
$56.00/ 1M tokens
google/gemini-2.5-flash-image
$42.00/ 1M tokens
google/gemini-3.1-flash-lite-image
$42.00/ 1M tokens
google/gemini-3.1-flash-image-preview
$84.00/ 1M tokens
google/gemini-3.1-flash-image
$84.00/ 1M tokens
google/gemini-3-pro-image-preview
$84.00/ 1M tokens
google/gemini-3-pro-image
$168.00/ 1M tokens
black-forest-labs/flux.2-klein-4b
$0.0196/ megapixel
black-forest-labs/flux.2-pro
$0.042/ megapixel
black-forest-labs/flux-3-image
$0.0574/ image
black-forest-labs/flux.2-flex
$0.084/ megapixel
black-forest-labs/flux.2-max
$0.098/ megapixel
bytedance-seed/seedream-5-0-flash
$0.0252/ image
bytedance-seed/seedream-5-0-lite
$0.049/ image
bytedance-seed/seedream-4.5
$0.056/ image
bytedance-seed/seedream-5-0-pro
$0.063/ image
microsoft/mai-image-2.6-flash
$26.60/ 1M tokens
microsoft/mai-image-2.6
$53.20/ 1M tokens
microsoft/mai-image-2.5
$65.80/ 1M tokens
microsoft/mai-image-2.5-pro
$151.20/ 1M tokens
sourceful/riverflow-v2.5-fast
$0.0266/ image
sourceful/riverflow-v2-fast
$0.028/ image
sourceful/riverflow-v2.5-pro
$0.182/ image
sourceful/riverflow-v2-pro
$0.21/ image
krea/krea-2-large
metered
krea/krea-2-medium
metered
krea/krea-2-medium-turbo
metered
inclusionai/ming-image-0.1-design
Free/ 1M tokens
inclusionai/ming-image-0.1-design-layer
Free/ 1M tokens
qwen/qwen-image-3
$0.042/ image
qwen/qwen-image-3-pro
$0.056/ image
x-ai/grok-imagine-image-2.0
$0.056/ image
x-ai/grok-imagine-image-quality
$0.07/ image
meta/muse-image
metered
Video generation 30 models billed per second of output — or per token where the provider meters that way; a range spans the model’s resolutions
alibaba/wan-3.0
$0.07–$0.28/ second
alibaba/wan-3.0-prime
$0.0952–$0.392/ second
alibaba/happyhorse-1.0
$0.138–$0.237/ second
alibaba/happyhorse-1.1
$0.138–$0.179/ second
alibaba/wan-2.7
$0.14/ second
alibaba/wan-2.6
metered
bytedance/seedance-1-5-pro
$1.68–$3.36/ 1M tokens
bytedance/seedance-2.0-mini
$2.94–$4.90/ 1M tokens
bytedance/seedance-2.0
$3.36–$10.78/ 1M tokens
bytedance/seedance-2.0-fast
$3.46–$5.88/ 1M tokens
bytedance/seedance-2.5
$8.96–$14.98/ 1M tokens
black-forest-labs/flux-video-upscale
$0.105–$0.147/ MP-second
black-forest-labs/flux-video-edit
metered
black-forest-labs/flux-3-video
metered
google/veo-3.1-lite
$0.042–$0.112/ second
google/veo-3.1-fast
$0.112–$0.42/ second
google/veo-3.1
$0.28–$0.84/ second
kwaivgi/kling-v3.0-std
$0.118–$0.176/ second
kwaivgi/kling-video-o1
$0.157/ second
kwaivgi/kling-v3.0-pro
$0.157–$0.235/ second
minimax/hailuo-3-max
$0.07–$0.112/ second
minimax/hailuo-2.3
$0.114/ second
minimax/hailuo-3
$0.182/ second
heygen/heygen-video-1
$0.028–$0.042/ second
heygen/avatar-iv
$0.07/ second
runway/aleph-2
metered
runway/gen-4.5
metered
x-ai/grok-imagine-video
metered
x-ai/grok-imagine-video-1.5
metered
openai/sora-2-pro
$0.42–$0.7/ second
Music generation 17 models full tracks from a text brief — flat per track, or by the second
minimax/music-3
$0.0028/ second
fal-ai/minimax-music/v1.5
$0.042/ generation
fal-ai/minimax-music/v2
$0.042/ generation
fal-ai/minimax-music/v2.5
$0.21/ generation
fal-ai/minimax-music/v2.6
$0.21/ generation
fal-ai/stable-audio-3/small/music/text-to-audio
$0.0304/ generation
fal-ai/stable-audio-3/medium/text-to-audio
$0.0526/ generation
fal-ai/stable-audio-3/medium/base/text-to-audio
$0.0671/ generation
fal-ai/stable-audio-25/text-to-audio
$0.28/ generation
fal-ai/lyria3
$0.056/ generation
fal-ai/lyria3/pro
$0.112/ generation
fal-ai/lyria2
$0.14/ generation
fal-ai/ace-step
$0.00028/ second
cassetteai/music-generator
$0.028/ minute
fal-ai/diffrhythm
$0.0014/ second
fal-ai/elevenlabs/music
$0.84/ minute
sonilo/v1.1/text-to-music
$0.0035/ second
Sound effects 3 models one short sound from a description — impacts, ambiences, foley
cassetteai/sound-effects-generator
$0.014/ generation
fal-ai/elevenlabs/sound-effects/v2
$0.0028/ second
fal-ai/stable-audio-3/small/sfx/text-to-audio
$0.0288/ generation
Voices (text to speech) 5 models billed per 1,000 characters submitted
fal-ai/elevenlabs/tts/turbo-v2.5
$0.07/ 1K chars
fal-ai/elevenlabs/tts/multilingual-v2
$0.14/ 1K chars
fal-ai/elevenlabs/tts/eleven-v3
$0.14/ 1K chars
fal-ai/kokoro/american-english
$0.028/ 1K chars
fal-ai/kokoro/british-english
$0.028/ 1K chars
3D model generation 10 models image to mesh; priced by the options you pick, so each model shows its range
fal-ai/hunyuan-3d/v3.1/rapid/image-to-3d
$0.322–$0.532/ mesh
fal-ai/hunyuan3d-v3/image-to-3d
$0.322–$1.26/ mesh
fal-ai/hunyuan-3d/v3.1/pro/image-to-3d
$0.532–$1.16/ mesh
tripo3d/h3.1/image-to-3d
$0.28–$0.91/ mesh
tripo3d/tripo/v2.5/image-to-3d
$0.28–$0.7/ mesh
tripo3d/p1/image-to-3d
$0.56–$0.7/ mesh
fal-ai/hyper3d/rodin/v2.5/fast
$0.14/ mesh
fal-ai/hyper3d/rodin/v2.5
$0.56–$1.68/ mesh
meshy/v7/image-to-3d
$1.12–$2.41/ mesh
fal-ai/trellis-2
$0.35–$0.49/ mesh