45 results Clear filters
Ranked by raw score for SWE-Bench Pro - . Versions are kept separate, and incompatible results are never combined.
Claude Fable 5by Anthropic
2026-06-09
anthropic/claude-fable-5

Claude model for creative writing, analysis, and controlled agent workflows

T Reasoning Tools Weight access not listed
Context
1M
Input
$10/M
Output
$50/M
Claude Opus 4.8by Anthropic
2026-05-28
anthropic/claude-opus-4-8

Top Claude Opus tier for the hardest reasoning, coding, and long-horizon agents

T Reasoning Tools Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
Grok 4.5by xAI
2026-07-08
xai/grok-4.5

xAI's latest Grok for chat, coding, agentic tools, and lower hallucination risk

T Reasoning Tools Structured output Weight access not listed
Context
500K
Input
$2/M
Output
$6/M
GPT-5.6 Solby OpenAI
2026-07-09
openai/gpt-5.6-sol

Frontier GPT-5.6 model for complex professional work, coding, and agentic workflows

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$5/M
Output
$30/M
Claude Opus 4.7by Anthropic
2026-04-16
anthropic/claude-opus-4-7

Stronger Opus tier for advanced software work and high-stakes reasoning

T Reasoning Tools Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
GPT-5.6 Terraby OpenAI
2026-07-09
openai/gpt-5.6-terra

Balanced GPT-5.6 model for capable, cost-efficient everyday work

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$2.5/M
Output
$15/M
Claude Sonnet 5by Anthropic
2026-06-30
anthropic/claude-sonnet-5

Everyday Claude agent model for coding, planning, browsing, and general work

T Reasoning Tools Weight access not listed
Context
1M
Input
$2/M
Output
$10/M
GPT-5.6 Lunaby OpenAI
2026-07-09
openai/gpt-5.6-luna

Cost-efficient GPT-5.6 model for fast, high-volume workloads

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$1/M
Output
$6/M
Ornith 1.0 397Bby Deepreinforce
2026-06-25
deepreinforce/ornith-1.0-397b

Large coding-reasoning model for agentic software tasks and RL search

T Open weights
Context
262.144K
Input
-
Output
-
GLM-5.2by Zhipu AI
2026-06-13
zhipuai/glm-5.2

Open flagship GLM for long-horizon coding agents and million-token context work

T Reasoning Tools Structured output Open weights
Context
1M
Input
$1.4/M
Output
$4.4/M
2026-04-08
meta/muse-spark-1.1

Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.

T Reasoning Tools Structured output Weight access not listed
Context
1M
Input
$1.25/M
Output
$4.25/M
Qwen3.7 Maxby Alibaba Qwen
2026-05-21
alibaba/qwen3.7-max

Qwen frontier model tuned for agent frameworks, coding assistants, and long tasks

T Reasoning Tools Weight access not listed
Context
1M
Input
$2.5/M
Output
$7.5/M
LongCat-2.0by Meituan
2026-06-30
meituan/longcat-2.0

Meituan LongCat-2.0, a reasoning model with tool calling and a 1M-token context window

T Reasoning Tools Weight access not listed
Context
1M
Input
$0.3/M
Output
$1.2/M
GPT-5.4by OpenAI
2026-03-05
openai/gpt-5.4

Agent-ready GPT for coding and computer-use workflows at a lower cost

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$2.5/M
Output
$15/M
MiniMax-M3by MiniMax
2026-06-01
minimax/MiniMax-M3

MiniMax multimodal model for long-context coding, perception, and agent planning

T Reasoning Tools Open weights
Context
512K
Input
$0.3/M
Output
$1.2/M
GPT-5.5by OpenAI
2026-04-23
openai/gpt-5.5

Default frontier GPT for coding, computer use, research, and knowledge work

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$5/M
Output
$30/M
MiMo-V2.5-Proby Xiaomi
2026-04-22
xiaomi/mimo-v2.5-pro

Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution

T Reasoning Tools Open weights
Context
1.04858M
Input
$0.435/M
Output
$0.87/M
Step 3.7 Flashby StepFun
2026-05-29
stepfun/step-3.7-flash

Newer StepFun flash model for faster agents, coding, and multimodal prompts

T Reasoning Tools Open weights
Context
256K
Input
$0.185/M
Output
$1.11/M
MiniMax-M2.7by MiniMax
2026-03-18
minimax/MiniMax-M2.7

Open MiniMax flagship for coding agents, office automation, and complex environments

T Reasoning Tools Open weights
Context
204.8K
Input
$0.3/M
Output
$1.2/M
2026-05-19
google/gemini-3.5-flash

Fast Gemini model balancing multimodal reasoning, tool use, and cost

T Reasoning Tools Structured output Weight access not listed
Context
1.04858M
Input
$1.5/M
Output
$9/M
2026-02-19
google/gemini-3.1-pro-preview

Reasoning-first Gemini preview for agentic coding and complex problem solving

T Reasoning Tools Structured output Weight access not listed
Context
1.04858M
Input
$2/M
Output
$12/M
Claude Opus 4.6by Anthropic
2026-02-05
anthropic/claude-opus-4-6

High-end Claude for difficult coding, planning, and slower expert reasoning

T Reasoning Tools Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
MAI-Code-1-Flashby Microsoft
2026-06-02
microsoft/mai-code-1-flash

Microsoft coding model built for fast, efficient assistance in everyday developer workflows

T Reasoning Tools Structured output Weight access not listed
Context
256K
Input
$0.75/M
Output
$4.5/M
Ornith 1.0 35Bby Deepreinforce
2026-06-25
deepreinforce/ornith-1.0-35b

Large coding-reasoning model for agentic software tasks and RL search

T Open weights
Context
262.144K
Input
-
Output
-
Laguna XS 2.1by Poolside
2026-07-02
poolside/laguna-xs-2.1

Agentic coding model from Poolside in the XS size class for local deployment

T Reasoning Tools Open weights
Context
262.144K
Input
$0.06/M
Output
$0.12/M
Claude Opus 4.5by Anthropic
2025-11-01
anthropic/claude-opus-4-5-20251101

Flagship Claude model for deep reasoning, coding, and long-horizon agents

T Reasoning Tools Weight access not listed
Context
200K
Input
$5/M
Output
$25/M
2025-09-29
anthropic/claude-sonnet-4-5

Balanced Claude model for coding, analysis, agent workflows, and cost control

T Reasoning Tools Weight access not listed
Context
200K
Input
$3/M
Output
$15/M
2025-11-18
google/gemini-3-pro-preview

Preview Gemini flagship for complex reasoning, coding, and rich multimodal prompts

T Reasoning Tools Structured output Weight access not listed
Context
1.04858M
Input
$2/M
Output
$12/M
Ornith 1.0 9Bby Deepreinforce
2026-06-25
deepreinforce/ornith-1.0-9b

Open coding-reasoning model for repository tasks and self-improving agents

T Open weights
Context
262.144K
Input
-
Output
-
2025-05-22
anthropic/claude-sonnet-4-0

Balanced Claude model for coding, analysis, agent workflows, and cost control

T Reasoning Tools Weight access not listed
Context
200K
Input
$3/M
Output
$15/M
GPT-5by OpenAI
2025-08-07
openai/gpt-5

Original GPT-5 workhorse for reasoning, coding, writing, and tool workflows

T Reasoning Tools Structured output Weight access not listed
Context
400K
Input
$1.25/M
Output
$10/M
GPT-5.2 Codexby OpenAI
2025-12-11
openai/gpt-5.2-codex

Code-specialist GPT for repository edits, reviews, and long-running software agents

T Reasoning Tools Weight access not listed
Context
400K
Input
$0.14/M
Output
$1.14/M
2026-06-09
cohere/north-mini-code-1-0

Cohere coding model for practical software engineering and agentic edits

T Reasoning Tools Structured output Open weights
Context
256K
Input
-
Output
-
2025-10-15
anthropic/claude-haiku-4-5

Fast Claude lane for lightweight agents, office tasks, and responsive chat

T Reasoning Tools Weight access not listed
Context
200K
Input
$1/M
Output
$5/M
2025-04
alibaba/qwen3-coder-480b-a35b-instruct

Open Qwen coding heavyweight for repository reasoning and agentic engineering

T Tools Open weights
Context
262.144K
Input
$1.5/M
Output
$7.5/M
MiniMax-M2.1by MiniMax
2025-12-23
minimax/MiniMax-M2.1

Earlier MiniMax agent model for practical coding and productivity tasks

T Reasoning Tools Open weights
Context
204.8K
Input
$0.3/M
Output
$1.2/M
2025-12-17
google/gemini-3-flash-preview

New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs

T Reasoning Tools Structured output Weight access not listed
Context
1.04858M
Input
$0.5/M
Output
$3/M
GPT-5.2by OpenAI
2025-12-11
openai/gpt-5.2

Reliable GPT generation for broad coding, writing, and tool-assisted product work

T Reasoning Tools Structured output Weight access not listed
Context
400K
Input
$1.75/M
Output
$14/M
Kimi K2.6by Moonshot AI
2026-04-21
moonshotai/kimi-k2.6

Multimodal Kimi workhorse for agent loops, coding tasks, and visual context

T Reasoning Tools Structured output Open weights
Context
262.144K
Input
$0.95/M
Output
$4/M
Qwen3 235B-A22Bby Alibaba Qwen
2025-04
alibaba/qwen3-235b-a22b

Large open Qwen MoE for multilingual reasoning, coding, and tool use

T Reasoning Tools Open weights
Context
131.072K
Input
$0.7/M
Output
$2.8/M
GLM-5.1by Zhipu AI
2026-04-07
zhipuai/glm-5.1

Strong GLM coding model for agentic engineering, terminals, and repository generation

T Reasoning Tools Structured output Open weights
Context
200K
Input
$1.4/M
Output
$4.4/M
DeepSeek V4 Proby DeepSeek
2026-04-24
deepseek/deepseek-v4-pro

Open MoE flagship with million-token context for coding and long agent runs

T Reasoning Tools Structured output Open weights
Context
1M
Input
$0.435/M
Output
$0.87/M
Claude Sonnet 4.6by Anthropic
2026-02-17
anthropic/claude-sonnet-4-6

Claude workhorse for coding agents, careful analysis, and production cost control

T Reasoning Tools Weight access not listed
Context
1M
Input
$3/M
Output
$15/M
GLM-4.6by Zhipu AI
2025-09-30
zhipuai/glm-4.6

Late GLM-4 workhorse for coding agents, reasoning, and structured tasks

T Reasoning Tools Open weights
Context
204.8K
Input
$0.6/M
Output
$2.2/M
2025-04-05
meta/llama-4-maverick-17b-instruct

Open multimodal Llama for strong reasoning with efficient everyday serving

T Open weights
Context
1M
Input
$0.124/M
Output
$0.603/M