38 results Clear filters
Ranked by raw score for SWE-Bench Verified - . Versions are kept separate, and incompatible results are never combined.
Claude Fable 5by Anthropic
2026-06-09
anthropic/claude-fable-5

Claude model for creative writing, analysis, and controlled agent workflows

T Reasoning Tools Weight access not listed
Context
1M
Input
$10/M
Output
$50/M
Claude Opus 4.8by Anthropic
2026-05-28
anthropic/claude-opus-4-8

Top Claude Opus tier for the hardest reasoning, coding, and long-horizon agents

T Reasoning Tools Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
Claude Sonnet 5by Anthropic
2026-06-30
anthropic/claude-sonnet-5

Everyday Claude agent model for coding, planning, browsing, and general work

T Reasoning Tools Weight access not listed
Context
1M
Input
$2/M
Output
$10/M
Ornith 1.0 397Bby Deepreinforce
2026-06-25
deepreinforce/ornith-1.0-397b

Large coding-reasoning model for agentic software tasks and RL search

T Open weights
Context
262.144K
Input
-
Output
-
DeepSeek V4 Proby DeepSeek
2026-04-24
deepseek/deepseek-v4-pro

Open MoE flagship with million-token context for coding and long agent runs

T Reasoning Tools Structured output Open weights
Context
1M
Input
$0.435/M
Output
$0.87/M
MiniMax-M3by MiniMax
2026-06-01
minimax/MiniMax-M3

MiniMax multimodal model for long-context coding, perception, and agent planning

T Reasoning Tools Open weights
Context
512K
Input
$0.3/M
Output
$1.2/M
Qwen3.7 Maxby Alibaba Qwen
2026-05-21
alibaba/qwen3.7-max

Qwen frontier model tuned for agent frameworks, coding assistants, and long tasks

T Reasoning Tools Weight access not listed
Context
1M
Input
$2.5/M
Output
$7.5/M
Kimi K2.6by Moonshot AI
2026-04-21
moonshotai/kimi-k2.6

Multimodal Kimi workhorse for agent loops, coding tasks, and visual context

T Reasoning Tools Structured output Open weights
Context
262.144K
Input
$0.95/M
Output
$4/M
MiniMax-M2.7by MiniMax
2026-03-18
minimax/MiniMax-M2.7

Open MiniMax flagship for coding agents, office automation, and complex environments

T Reasoning Tools Open weights
Context
204.8K
Input
$0.3/M
Output
$1.2/M
2026-04-24
deepseek/deepseek-v4-flash

Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work

T Reasoning Tools Structured output Open weights
Context
1M
Input
$0.14/M
Output
$0.28/M
MiMo-V2.5-Proby Xiaomi
2026-04-22
xiaomi/mimo-v2.5-pro

Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution

T Reasoning Tools Open weights
Context
1.04858M
Input
$0.435/M
Output
$0.87/M
Hy3by Tencent
2026-07-06
tencent/hy3

Tencent Hy reasoning model for coding, instruction following, and agent tasks

T Reasoning Tools Open weights
Context
256K
Input
$0.132/M
Output
$0.528/M
Mistral Medium 3.5by Mistral AI
2026-04-29
mistral/mistral-medium-2604

Balanced Mistral model for enterprise assistants, multilingual work, and tools

T Reasoning Tools Structured output Open weights
Context
262.144K
Input
$1.5/M
Output
$7.5/M
2026-04-29
mistral/mistral-medium-latest

Balanced Mistral model for enterprise assistants, multilingual work, and tools

T Reasoning Tools Structured output Open weights
Context
262.144K
Input
$1.5/M
Output
$7.5/M
Qwen3.6 27Bby Alibaba Qwen
2026-04-22
alibaba/qwen3.6-27b

Qwen vision-language model for visual reasoning, documents, and agent tasks

T Reasoning Tools Structured output Open weights
Context
262.144K
Input
$0.6/M
Output
$3.6/M
Step 3.7 Flashby StepFun
2026-05-29
stepfun/step-3.7-flash

Newer StepFun flash model for faster agents, coding, and multimodal prompts

T Reasoning Tools Open weights
Context
256K
Input
$0.185/M
Output
$1.11/M
Qwen3.5 397B-A17Bby Alibaba Qwen
2026-02-15
alibaba/qwen3.5-397b-a17b

Large open Qwen multimodal MoE for visual agents and long technical tasks

T Reasoning Tools Structured output Open weights
Context
262.144K
Input
$0.6/M
Output
$3.6/M
MiniMax-M2.5by MiniMax
2026-02-12
minimax/MiniMax-M2.5

Prior MiniMax coding model for agent workflows, office edits, and automation

T Reasoning Tools Open weights
Context
204.8K
Input
$0.3/M
Output
$1.2/M
Ornith 1.0 35Bby Deepreinforce
2026-06-25
deepreinforce/ornith-1.0-35b

Large coding-reasoning model for agentic software tasks and RL search

T Open weights
Context
262.144K
Input
-
Output
-
Step 3.5 Flashby StepFun
2026-01-29
stepfun/step-3.5-flash

StepFun flash lane for quick multimodal reasoning and coding assistance

T Reasoning Tools Open weights
Context
256K
Input
$0.1/M
Output
$0.3/M
Hy3 previewby Tencent
2026-04-20
tencent/hy3-preview

Tencent Hy reasoning model for coding, instruction following, and agent tasks

T Reasoning Tools Open weights
Context
256K
Input
$0.063/M
Output
$0.21/M
MiniMax-M2.1by MiniMax
2025-12-23
minimax/MiniMax-M2.1

Earlier MiniMax agent model for practical coding and productivity tasks

T Reasoning Tools Open weights
Context
204.8K
Input
$0.3/M
Output
$1.2/M
GLM-4.7by Zhipu AI
2025-12-22
zhipuai/glm-4.7

Mature GLM model for dependable coding, reasoning, and structured agent tasks

T Reasoning Tools Open weights
Context
204.8K
Input
$0.6/M
Output
$2.2/M
Qwen3.6 35B-A3Bby Alibaba Qwen
2026-04-17
alibaba/qwen3.6-35b-a3b

Open multimodal Qwen MoE for local agents that need vision, audio, and code

T Reasoning Tools Structured output Open weights
Context
262.144K
Input
$0.248/M
Output
$1.485/M
GLM-5by Zhipu AI
2026-02-12
zhipuai/glm-5

General GLM flagship for coding, analysis, and tool-heavy engineering workflows

T Reasoning Tools Open weights
Context
204.8K
Input
$1/M
Output
$3.2/M
Qwen3.5 27Bby Alibaba Qwen
2026-02-23
alibaba/qwen3.5-27b

Qwen vision-language model for visual reasoning, documents, and agent tasks

T Tools Structured output Open weights
Context
262.144K
Input
$0.3/M
Output
$2.4/M
Qwen3.5 122B-A10Bby Alibaba Qwen
2026-02-23
alibaba/qwen3.5-122b-a10b

Qwen vision-language model for visual reasoning, documents, and agent tasks

T Reasoning Tools Structured output Open weights
Context
262.144K
Input
$0.4/M
Output
$3.2/M
MAI-Code-1-Flashby Microsoft
2026-06-02
microsoft/mai-code-1-flash

Microsoft coding model built for fast, efficient assistance in everyday developer workflows

T Reasoning Tools Structured output Weight access not listed
Context
256K
Input
$0.75/M
Output
$4.5/M
Kimi K2 Thinkingby Moonshot AI
2025-11-06
moonshotai/kimi-k2-thinking

Thinking Kimi model for slower research passes, planning, and hard technical questions

T Tools Structured output Open weights
Context
262.144K
Input
$0.6/M
Output
$2.5/M
Laguna XS 2.1by Poolside
2026-07-02
poolside/laguna-xs-2.1

Agentic coding model from Poolside in the XS size class for local deployment

T Reasoning Tools Open weights
Context
262.144K
Input
$0.06/M
Output
$0.12/M
Kimi K2.5by Moonshot AI
2026-01
moonshotai/kimi-k2.5

Earlier Kimi frontier model for long-context agents, coding, and multimodal work

T Reasoning Tools Structured output Open weights
Context
262.144K
Input
$0.6/M
Output
$3/M
2026-06-04
nvidia/nemotron-3-ultra-550b-a55b

Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy

T Reasoning Tools Structured output Open weights
Context
1M
Input
$0.5/M
Output
$2.5/M
Ornith 1.0 9Bby Deepreinforce
2026-06-25
deepreinforce/ornith-1.0-9b

Open coding-reasoning model for repository tasks and self-improving agents

T Open weights
Context
262.144K
Input
-
Output
-
MiniMax-M2by MiniMax
2025-10-27
minimax/MiniMax-M2

Efficient open MiniMax model built for coding agents and tool-heavy workflows

T Reasoning Tools Open weights
Context
196.608K
Input
$0.3/M
Output
$1.2/M
2026-06-09
cohere/north-mini-code-1-0

Cohere coding model for practical software engineering and agentic edits

T Reasoning Tools Structured output Open weights
Context
256K
Input
-
Output
-
Devstral Mediumby Mistral AI
2025-07-10
mistral/devstral-medium-2507

Mistral coding agent model for repository tasks and software engineering workflows

T Tools Weight access not listed
Context
128K
Input
$0.4/M
Output
$2/M
GLM-4.7-Flashby Zhipu AI
2026-01-19
zhipuai/glm-4.7-flash

Budget GLM lane for fast coding help, routing, and everyday automation

T Reasoning Tools Open weights
Context
200K
Input
$0.04/M
Output
$0.3/M
Devstral Smallby Mistral AI
2025-07-10
mistral/devstral-small-2507

Mistral coding agent model for repository tasks and software engineering workflows

T Tools Open weights
Context
128K
Input
$0.1/M
Output
$0.3/M