24 results Clear filters
Ranked by raw score for Terminal-Bench - 2.1. Versions are kept separate, and incompatible results are never combined.
GPT-5.6 Solby OpenAI
2026-07-09
openai/gpt-5.6-sol

Frontier GPT-5.6 model for complex professional work, coding, and agentic workflows

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$5/M
Output
$30/M
Claude Fable 5by Anthropic
2026-06-09
anthropic/claude-fable-5

Claude model for creative writing, analysis, and controlled agent workflows

T Reasoning Tools Weight access not listed
Context
1M
Input
$10/M
Output
$50/M
GPT-5.6 Terraby OpenAI
2026-07-09
openai/gpt-5.6-terra

Balanced GPT-5.6 model for capable, cost-efficient everyday work

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$2.5/M
Output
$15/M
GPT-5.6 Lunaby OpenAI
2026-07-09
openai/gpt-5.6-luna

Cost-efficient GPT-5.6 model for fast, high-volume workloads

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$1/M
Output
$6/M
Grok 4.5by xAI
2026-07-08
xai/grok-4.5

xAI's latest Grok for chat, coding, agentic tools, and lower hallucination risk

T Reasoning Tools Structured output Weight access not listed
Context
500K
Input
$2/M
Output
$6/M
GLM-5.2by Zhipu AI
2026-06-13
zhipuai/glm-5.2

Open flagship GLM for long-horizon coding agents and million-token context work

T Reasoning Tools Structured output Open weights
Context
1M
Input
$1.4/M
Output
$4.4/M
Claude Sonnet 5by Anthropic
2026-06-30
anthropic/claude-sonnet-5

Everyday Claude agent model for coding, planning, browsing, and general work

T Reasoning Tools Weight access not listed
Context
1M
Input
$2/M
Output
$10/M
2026-04-08
meta/muse-spark-1.1

Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.

T Reasoning Tools Structured output Weight access not listed
Context
1M
Input
$1.25/M
Output
$4.25/M
GPT-5.5by OpenAI
2026-04-23
openai/gpt-5.5

Default frontier GPT for coding, computer use, research, and knowledge work

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$5/M
Output
$30/M
2026-05-19
google/gemini-3.5-flash

Fast Gemini model balancing multimodal reasoning, tool use, and cost

T Reasoning Tools Structured output Weight access not listed
Context
1.04858M
Input
$1.5/M
Output
$9/M
Claude Opus 4.8by Anthropic
2026-05-28
anthropic/claude-opus-4-8

Top Claude Opus tier for the hardest reasoning, coding, and long-horizon agents

T Reasoning Tools Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
Claude Opus 4.7by Anthropic
2026-04-16
anthropic/claude-opus-4-7

Stronger Opus tier for advanced software work and high-stakes reasoning

T Reasoning Tools Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
LongCat-2.0by Meituan
2026-06-30
meituan/longcat-2.0

Meituan LongCat-2.0, a reasoning model with tool calling and a 1M-token context window

T Reasoning Tools Weight access not listed
Context
1M
Input
$0.3/M
Output
$1.2/M
2026-02-19
google/gemini-3.1-pro-preview

Reasoning-first Gemini preview for agentic coding and complex problem solving

T Reasoning Tools Structured output Weight access not listed
Context
1.04858M
Input
$2/M
Output
$12/M
Claude Opus 4.6by Anthropic
2026-02-05
anthropic/claude-opus-4-6

High-end Claude for difficult coding, planning, and slower expert reasoning

T Reasoning Tools Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
GPT-5.4by OpenAI
2026-03-05
openai/gpt-5.4

Agent-ready GPT for coding and computer-use workflows at a lower cost

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$2.5/M
Output
$15/M
Claude Sonnet 4.6by Anthropic
2026-02-17
anthropic/claude-sonnet-4-6

Claude workhorse for coding agents, careful analysis, and production cost control

T Reasoning Tools Weight access not listed
Context
1M
Input
$3/M
Output
$15/M
MiniMax-M3by MiniMax
2026-06-01
minimax/MiniMax-M3

MiniMax multimodal model for long-context coding, perception, and agent planning

T Reasoning Tools Open weights
Context
512K
Input
$0.3/M
Output
$1.2/M
GLM-5.1by Zhipu AI
2026-04-07
zhipuai/glm-5.1

Strong GLM coding model for agentic engineering, terminals, and repository generation

T Reasoning Tools Structured output Open weights
Context
200K
Input
$1.4/M
Output
$4.4/M
DeepSeek V4 Proby DeepSeek
2026-04-24
deepseek/deepseek-v4-pro

Open MoE flagship with million-token context for coding and long agent runs

T Reasoning Tools Structured output Open weights
Context
1M
Input
$0.435/M
Output
$0.87/M
Kimi K2.6by Moonshot AI
2026-04-21
moonshotai/kimi-k2.6

Multimodal Kimi workhorse for agent loops, coding tasks, and visual context

T Reasoning Tools Structured output Open weights
Context
262.144K
Input
$0.95/M
Output
$4/M
Step 3.7 Flashby StepFun
2026-05-29
stepfun/step-3.7-flash

Newer StepFun flash model for faster agents, coding, and multimodal prompts

T Reasoning Tools Open weights
Context
256K
Input
$0.185/M
Output
$1.11/M
2026-06-04
nvidia/nemotron-3-ultra-550b-a55b

Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy

T Reasoning Tools Structured output Open weights
Context
1M
Input
$0.5/M
Output
$2.5/M
MiniMax-M2.7by MiniMax
2026-03-18
minimax/MiniMax-M2.7

Open MiniMax flagship for coding agents, office automation, and complex environments

T Reasoning Tools Open weights
Context
204.8K
Input
$0.3/M
Output
$1.2/M