9 results Clear filters
Ranked by raw score for SWE-Bench Multilingual - . Versions are kept separate, and incompatible results are never combined.
Ornith 1.0 397Bby Deepreinforce
2026-06-25
deepreinforce/ornith-1.0-397b

Large coding-reasoning model for agentic software tasks and RL search

T Open weights
Context
262.144K
Input
-
Output
-
Qwen3.7 Maxby Alibaba Qwen
2026-05-21
alibaba/qwen3.7-max

Qwen frontier model tuned for agent frameworks, coding assistants, and long tasks

T Reasoning Tools Weight access not listed
Context
1M
Input
$2.5/M
Output
$7.5/M
Claude Sonnet 5by Anthropic
2026-06-30
anthropic/claude-sonnet-5

Everyday Claude agent model for coding, planning, browsing, and general work

T Reasoning Tools Weight access not listed
Context
1M
Input
$2/M
Output
$10/M
Grok 4.5by xAI
2026-07-08
xai/grok-4.5

xAI's latest Grok for chat, coding, agentic tools, and lower hallucination risk

T Reasoning Tools Structured output Weight access not listed
Context
500K
Input
$2/M
Output
$6/M
LongCat-2.0by Meituan
2026-06-30
meituan/longcat-2.0

Meituan LongCat-2.0, a reasoning model with tool calling and a 1M-token context window

T Reasoning Tools Weight access not listed
Context
1M
Input
$0.3/M
Output
$1.2/M
Ornith 1.0 35Bby Deepreinforce
2026-06-25
deepreinforce/ornith-1.0-35b

Large coding-reasoning model for agentic software tasks and RL search

T Open weights
Context
262.144K
Input
-
Output
-
2026-06-04
nvidia/nemotron-3-ultra-550b-a55b

Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy

T Reasoning Tools Structured output Open weights
Context
1M
Input
$0.5/M
Output
$2.5/M
Laguna XS 2.1by Poolside
2026-07-02
poolside/laguna-xs-2.1

Agentic coding model from Poolside in the XS size class for local deployment

T Reasoning Tools Open weights
Context
262.144K
Input
$0.06/M
Output
$0.12/M
Ornith 1.0 9Bby Deepreinforce
2026-06-25
deepreinforce/ornith-1.0-9b

Open coding-reasoning model for repository tasks and self-improving agents

T Open weights
Context
262.144K
Input
-
Output
-