openai/gpt-5.1
Sharper GPT-5 generation for coding, product work, and tool-assisted tasks
- Context
- 400K
- Input
- $1.25/M
- Output
- $10/M
Filter 123 source-linked models by creator, price, context, modality, and published benchmark coverage.
openai/gpt-5.1
Sharper GPT-5 generation for coding, product work, and tool-assisted tasks
openai/gpt-5.1:batch
GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...
minimax/minimax-m2.5
MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1...
x-ai/grok-4.3
Grok 4.3 is a reasoning model from xAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...
anthropic/claude-opus-4.1
Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...
anthropic/claude-opus-4.1:batch
Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...
anthropic/claude-opus-4
Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workflows. It sets new benchmarks in...
anthropic/claude-sonnet-4.5
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
anthropic/claude-sonnet-4.5:batch
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
google/gemini-3-flash-preview
New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs
google/gemini-3-flash-preview:batch
Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...
openai/gpt-5.3-codex
Coding-optimized GPT model for repository edits, reviews, and agentic software work
z-ai/glm-4.6
Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...
stepfun/step-3.7-flash
Newer StepFun flash model for faster agents, coding, and multimodal prompts
deepseek/deepseek-v3.2-exp
DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...
z-ai/glm-4.5
GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...
qwen/qwen3.5-397b-a17b
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...
anthropic/claude-sonnet-4
Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),...
openai/gpt-5.1-codex
Codex GPT for repository edits, code review, and practical software agents
minimax/minimax-m2.1
MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world...
z-ai/glm-4.7-flash
As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...
tencent/hy3
Tencent Hy reasoning model for coding, instruction following, and agent tasks
deepseek/deepseek-v3.2
DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...
tencent/hy3:free
Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...
deepseek/deepseek-v3.1-terminus
DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's...
openai/gpt-5-mini
Small GPT-5 for responsive agents, coding help, and everyday automation
openai/gpt-5-mini:batch
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....
nvidia/nemotron-3-ultra-550b-a55b
Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy
nvidia/nemotron-3-ultra-550b-a55b:free
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
minimax/minimax-m2
MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning,...
google/gemini-2.5-pro
Google's proven reasoning model for coding, math, and multimodal analysis
google/gemini-2.5-pro:batch
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
qwen/qwen3.5-plus-02-15
The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of...
qwen/qwen3-coder:free
Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...
qwen/qwen3-coder
Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...
deepseek/deepseek-r1-0528
May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active...
anthropic/claude-haiku-4.5
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
anthropic/claude-haiku-4.5:batch
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
qwen/qwen3-max
Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It...
z-ai/glm-4.5-air
GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter...
openai/gpt-5.1-codex-mini
Coding-optimized GPT model for repository edits, reviews, and agentic software work
deepseek/deepseek-chat-v3.1
DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context...
mistralai/mistral-large-2512
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
openai/gpt-4.1
Long-lived GPT workhorse for coding, instruction following, and production apps
arcee-ai/trinity-large-thinking
Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloads, and reasoning tasks. Launch video: https://youtu.be/Gc82AXLa0Rg?si=4RLn6WBz33qT--B7...
mistralai/mistral-medium-3.1
Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...
openai/gpt-4.1-mini
Affordable GPT-4.1 lane for fast coding help and structured extraction
google/gemini-2.5-flash
Fast Gemini workhorse for multimodal apps where latency and price matter
google/gemini-2.5-flash:batch
Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...
deepseek/deepseek-chat
DeepSeek chat model for instruction following, coding, and analysis
| Model | Creator | Raw benchmark score | Input types | Context | Input / Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| GPT-5.1openai/gpt-5.1 | 1235.0 | 400K | $1.25 / $10 | 2025-11-13 | |||
| OpenAI: GPT-5.1 (batch)openai/gpt-5.1:batch | 1235.0 | 400K | $0.625 / $5 | Undated | |||
| MiniMax: MiniMax M2.5minimax/minimax-m2.5 | 1233.0 | 196.608K | $0.15 / $0.9 | Undated | |||
| xAI: Grok 4.3x-ai/grok-4.3 | 1233.0 | 1M | $1.25 / $2.5 | Undated | |||
| Anthropic: Claude Opus 4.1anthropic/claude-opus-4.1 | 1228.0 | 200K | $15 / $75 | Undated | |||
| Anthropic: Claude Opus 4.1 (batch)anthropic/claude-opus-4.1:batch | 1228.0 | 200K | $7.5 / $37.5 | Undated | |||
| Anthropic: Claude Opus 4anthropic/claude-opus-4 | 1227.0 | 200K | $15 / $75 | Undated | |||
| Anthropic: Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | 1223.0 | 1M | $3 / $15 | Undated | |||
| Anthropic: Claude Sonnet 4.5 (batch)anthropic/claude-sonnet-4.5:batch | 1223.0 | 1M | $1.5 / $7.5 | Undated | |||
| Gemini 3 Flash Previewgoogle/gemini-3-flash-preview | 1222.0 | 1.04858M | $0.5 / $3 | 2025-12-17 | |||
| Google: Gemini 3 Flash Preview (batch)google/gemini-3-flash-preview:batch | 1222.0 | 1.04858M | $0.25 / $1.5 | Undated | |||
| GPT-5.3 Codexopenai/gpt-5.3-codex | 1218.0 | 400K | $1.75 / $14 | 2026-02-05 | |||
| Z.ai: GLM 4.6z-ai/glm-4.6 | 1206.0 | 202.752K | $0.5 / $2 | Undated | |||
| Step 3.7 Flashstepfun/step-3.7-flash | 1204.0 | 256K | $0.185 / $1.11 | 2026-05-29 | |||
| DeepSeek: DeepSeek V3.2 Expdeepseek/deepseek-v3.2-exp | 1201.0 | 163.84K | $0.27 / $0.41 | Undated | |||
| Z.ai: GLM 4.5z-ai/glm-4.5 | 1201.0 | 131.072K | $0.6 / $2.2 | Undated | |||
| Qwen: Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b | 1196.0 | 262.144K | $0.39 / $2.34 | Undated | |||
| Anthropic: Claude Sonnet 4anthropic/claude-sonnet-4 | 1195.0 | 200K | $3 / $15 | Undated | |||
| GPT-5.1 Codexopenai/gpt-5.1-codex | 1195.0 | 400K | $1.07 / $8.5 | 2025-11-13 | |||
| MiniMax: MiniMax M2.1minimax/minimax-m2.1 | 1191.0 | 204.8K | $0.3 / $1.2 | Undated | |||
| Z.ai: GLM 4.7 Flashz-ai/glm-4.7-flash | 1191.0 | 202.752K | $0.06 / $0.4 | Undated | |||
| Hy3tencent/hy3 | 1190.0 | 256K | $0.132 / $0.528 | 2026-07-06 | |||
| DeepSeek: DeepSeek V3.2deepseek/deepseek-v3.2 | 1188.0 | 163.84K | $0.269 / $0.4 | Undated | |||
| Tencent: Hy3 (free)tencent/hy3:free | 1188.0 | 262.144K | Free / Free | Undated | |||
| DeepSeek: DeepSeek V3.1 Terminusdeepseek/deepseek-v3.1-terminus | 1187.0 | 131.072K | $0.27 / $1 | Undated | |||
| GPT-5 Miniopenai/gpt-5-mini | 1185.0 | 400K | $0.25 / $2 | 2025-08-07 | |||
| OpenAI: GPT-5 Mini (batch)openai/gpt-5-mini:batch | 1185.0 | 400K | $0.125 / $1 | Undated | |||
| Nemotron 3 Ultra 550B A55Bnvidia/nemotron-3-ultra-550b-a55b | 1183.0 | 1M | $0.5 / $2.5 | 2026-06-04 | |||
| NVIDIA: Nemotron 3 Ultra (free)nvidia/nemotron-3-ultra-550b-a55b:free | 1183.0 | 1M | Free / Free | Undated | |||
| MiniMax: MiniMax M2minimax/minimax-m2 | 1174.0 | 204.8K | $0.255 / $1.02 | Undated | |||
| Gemini 2.5 Progoogle/gemini-2.5-pro | 1168.0 | 1.04858M | $1.25 / $10 | 2025-06-17 | |||
| Google: Gemini 2.5 Pro (batch)google/gemini-2.5-pro:batch | 1168.0 | 1.04858M | $0.625 / $5 | Undated | |||
| Qwen: Qwen3.5 Plus 2026-02-15qwen/qwen3.5-plus-02-15 | 1162.0 | 1M | $0.26 / $1.56 | Undated | |||
| Qwen: Qwen3 Coder 480B A35B (free)qwen/qwen3-coder:free | 1160.0 | 262K | Free / Free | Undated | |||
| Qwen: Qwen3 Coder 480B A35Bqwen/qwen3-coder | 1157.0 | 262.144K | $0.3 / $1 | Undated | |||
| DeepSeek: R1 0528deepseek/deepseek-r1-0528 | 1154.0 | 163.84K | $0.5 / $2.15 | Undated | |||
| Anthropic: Claude Haiku 4.5anthropic/claude-haiku-4.5 | 1153.0 | 200K | $1 / $5 | Undated | |||
| Anthropic: Claude Haiku 4.5 (batch)anthropic/claude-haiku-4.5:batch | 1153.0 | 200K | $0.5 / $2.5 | Undated | |||
| Qwen: Qwen3 Maxqwen/qwen3-max | 1151.0 | 262.144K | $0.78 / $3.9 | Undated | |||
| Z.ai: GLM 4.5 Airz-ai/glm-4.5-air | 1150.0 | 131.072K | $0.13 / $0.85 | Undated | |||
| GPT-5.1 Codex miniopenai/gpt-5.1-codex-mini | 1149.0 | 400K | $0.22 / $1.8 | 2025-11-13 | |||
| DeepSeek: DeepSeek V3.1deepseek/deepseek-chat-v3.1 | 1141.0 | 163.84K | $0.25 / $0.95 | Undated | |||
| Mistral: Mistral Large 3 2512mistralai/mistral-large-2512 | 1136.0 | 262.144K | $0.5 / $1.5 | Undated | |||
| GPT-4.1openai/gpt-4.1 | 1136.0 | 1.04758M | $2 / $8 | 2025-04-14 | |||
| Arcee AI: Trinity Large Thinkingarcee-ai/trinity-large-thinking | 1135.0 | 262.144K | $0.22 / $0.85 | Undated | |||
| Mistral: Mistral Medium 3.1mistralai/mistral-medium-3.1 | 1130.0 | 131.072K | $0.4 / $2 | Undated | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 1127.0 | 1.04758M | $0.4 / $1.6 | 2025-04-14 | |||
| Gemini 2.5 Flashgoogle/gemini-2.5-flash | 1122.0 | 1.04858M | $0.3 / $2.5 | 2025-06-17 | |||
| Google: Gemini 2.5 Flash (batch)google/gemini-2.5-flash:batch | 1122.0 | 1.04858M | $0.15 / $1.25 | Undated | |||
| DeepSeek Chatdeepseek/deepseek-chat | 1110.0 | 1M | $0.14 / $0.28 | 2025-12-01 |