123 results Clear filters
Ranked by raw score for Design Arena: gamedev - . Versions are kept separate, and incompatible results are never combined.
GPT-5.1by OpenAI
2025-11-13
openai/gpt-5.1

Sharper GPT-5 generation for coding, product work, and tool-assisted tasks

T Reasoning Tools Structured output Weight access not listed
Context
400K
Input
$1.25/M
Output
$10/M
Undated
openai/gpt-5.1:batch

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...

T Weight access not listed
Context
400K
Input
$0.625/M
Output
$5/M
Undated
minimax/minimax-m2.5

MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1...

T Weight access not listed
Context
196.608K
Input
$0.15/M
Output
$0.9/M
Undated
x-ai/grok-4.3

Grok 4.3 is a reasoning model from xAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

T Weight access not listed
Context
1M
Input
$1.25/M
Output
$2.5/M
Undated
anthropic/claude-opus-4.1

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...

T Weight access not listed
Context
200K
Input
$15/M
Output
$75/M
anthropic/claude-opus-4.1:batch

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...

T Weight access not listed
Context
200K
Input
$7.5/M
Output
$37.5/M
Undated
anthropic/claude-opus-4

Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workflows. It sets new benchmarks in...

T Weight access not listed
Context
200K
Input
$15/M
Output
$75/M
Undated
anthropic/claude-sonnet-4.5

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

T Weight access not listed
Context
1M
Input
$3/M
Output
$15/M
anthropic/claude-sonnet-4.5:batch

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

T Weight access not listed
Context
1M
Input
$1.5/M
Output
$7.5/M
2025-12-17
google/gemini-3-flash-preview

New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs

T Reasoning Tools Structured output Weight access not listed
Context
1.04858M
Input
$0.5/M
Output
$3/M
google/gemini-3-flash-preview:batch

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...

T Weight access not listed
Context
1.04858M
Input
$0.25/M
Output
$1.5/M
GPT-5.3 Codexby OpenAI
2026-02-05
openai/gpt-5.3-codex

Coding-optimized GPT model for repository edits, reviews, and agentic software work

T Reasoning Tools Weight access not listed
Context
400K
Input
$1.75/M
Output
$14/M
Undated
z-ai/glm-4.6

Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...

T Weight access not listed
Context
202.752K
Input
$0.5/M
Output
$2/M
Step 3.7 Flashby StepFun
2026-05-29
stepfun/step-3.7-flash

Newer StepFun flash model for faster agents, coding, and multimodal prompts

T Reasoning Tools Open weights
Context
256K
Input
$0.185/M
Output
$1.11/M
Undated
deepseek/deepseek-v3.2-exp

DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...

T Weight access not listed
Context
163.84K
Input
$0.27/M
Output
$0.41/M
Undated
z-ai/glm-4.5

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...

T Weight access not listed
Context
131.072K
Input
$0.6/M
Output
$2.2/M
Undated
qwen/qwen3.5-397b-a17b

The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...

T Weight access not listed
Context
262.144K
Input
$0.39/M
Output
$2.34/M
Undated
anthropic/claude-sonnet-4

Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),...

T Weight access not listed
Context
200K
Input
$3/M
Output
$15/M
GPT-5.1 Codexby OpenAI
2025-11-13
openai/gpt-5.1-codex

Codex GPT for repository edits, code review, and practical software agents

T Reasoning Tools Weight access not listed
Context
400K
Input
$1.07/M
Output
$8.5/M
Undated
minimax/minimax-m2.1

MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world...

T Weight access not listed
Context
204.8K
Input
$0.3/M
Output
$1.2/M
Undated
z-ai/glm-4.7-flash

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...

T Weight access not listed
Context
202.752K
Input
$0.06/M
Output
$0.4/M
Hy3by Tencent
2026-07-06
tencent/hy3

Tencent Hy reasoning model for coding, instruction following, and agent tasks

T Reasoning Tools Open weights
Context
256K
Input
$0.132/M
Output
$0.528/M
Undated
deepseek/deepseek-v3.2

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...

T Weight access not listed
Context
163.84K
Input
$0.269/M
Output
$0.4/M
Undated
tencent/hy3:free

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...

T Weight access not listed
Context
262.144K
Input
Free
Output
Free
deepseek/deepseek-v3.1-terminus

DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's...

T Weight access not listed
Context
131.072K
Input
$0.27/M
Output
$1/M
GPT-5 Miniby OpenAI
2025-08-07
openai/gpt-5-mini

Small GPT-5 for responsive agents, coding help, and everyday automation

T Reasoning Tools Structured output Weight access not listed
Context
400K
Input
$0.25/M
Output
$2/M
Undated
openai/gpt-5-mini:batch

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....

T Weight access not listed
Context
400K
Input
$0.125/M
Output
$1/M
2026-06-04
nvidia/nemotron-3-ultra-550b-a55b

Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy

T Reasoning Tools Structured output Open weights
Context
1M
Input
$0.5/M
Output
$2.5/M
nvidia/nemotron-3-ultra-550b-a55b:free

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

T Weight access not listed
Context
1M
Input
Free
Output
Free
Undated
minimax/minimax-m2

MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning,...

T Weight access not listed
Context
204.8K
Input
$0.255/M
Output
$1.02/M
2025-06-17
google/gemini-2.5-pro

Google's proven reasoning model for coding, math, and multimodal analysis

T Reasoning Tools Structured output Weight access not listed
Context
1.04858M
Input
$1.25/M
Output
$10/M
google/gemini-2.5-pro:batch

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

T Weight access not listed
Context
1.04858M
Input
$0.625/M
Output
$5/M
Undated
qwen/qwen3.5-plus-02-15

The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of...

T Weight access not listed
Context
1M
Input
$0.26/M
Output
$1.56/M
Undated
qwen/qwen3-coder:free

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...

T Weight access not listed
Context
262K
Input
Free
Output
Free
Undated
qwen/qwen3-coder

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...

T Weight access not listed
Context
262.144K
Input
$0.3/M
Output
$1/M
Undated
deepseek/deepseek-r1-0528

May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active...

T Weight access not listed
Context
163.84K
Input
$0.5/M
Output
$2.15/M
Undated
anthropic/claude-haiku-4.5

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

T Weight access not listed
Context
200K
Input
$1/M
Output
$5/M
anthropic/claude-haiku-4.5:batch

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

T Weight access not listed
Context
200K
Input
$0.5/M
Output
$2.5/M
Qwen: Qwen3 Maxby Alibaba Qwen
Undated
qwen/qwen3-max

Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It...

T Weight access not listed
Context
262.144K
Input
$0.78/M
Output
$3.9/M
Undated
z-ai/glm-4.5-air

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter...

T Weight access not listed
Context
131.072K
Input
$0.13/M
Output
$0.85/M
2025-11-13
openai/gpt-5.1-codex-mini

Coding-optimized GPT model for repository edits, reviews, and agentic software work

T Reasoning Tools Weight access not listed
Context
400K
Input
$0.22/M
Output
$1.8/M
Undated
deepseek/deepseek-chat-v3.1

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context...

T Weight access not listed
Context
163.84K
Input
$0.25/M
Output
$0.95/M
Undated
mistralai/mistral-large-2512

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.

T Weight access not listed
Context
262.144K
Input
$0.5/M
Output
$1.5/M
GPT-4.1by OpenAI
2025-04-14
openai/gpt-4.1

Long-lived GPT workhorse for coding, instruction following, and production apps

T Tools Structured output Weight access not listed
Context
1.04758M
Input
$2/M
Output
$8/M
arcee-ai/trinity-large-thinking

Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloads, and reasoning tasks. Launch video: https://youtu.be/Gc82AXLa0Rg?si=4RLn6WBz33qT--B7...

T Weight access not listed
Context
262.144K
Input
$0.22/M
Output
$0.85/M
Undated
mistralai/mistral-medium-3.1

Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...

T Weight access not listed
Context
131.072K
Input
$0.4/M
Output
$2/M
GPT-4.1 miniby OpenAI
2025-04-14
openai/gpt-4.1-mini

Affordable GPT-4.1 lane for fast coding help and structured extraction

T Tools Structured output Weight access not listed
Context
1.04758M
Input
$0.4/M
Output
$1.6/M
2025-06-17
google/gemini-2.5-flash

Fast Gemini workhorse for multimodal apps where latency and price matter

T Reasoning Tools Structured output Weight access not listed
Context
1.04858M
Input
$0.3/M
Output
$2.5/M
google/gemini-2.5-flash:batch

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

T Weight access not listed
Context
1.04858M
Input
$0.15/M
Output
$1.25/M
DeepSeek Chatby DeepSeek
2025-12-01
deepseek/deepseek-chat

DeepSeek chat model for instruction following, coding, and analysis

T Tools Open weights
Context
1M
Input
$0.14/M
Output
$0.28/M