5 results Clear filters
Ranked by raw score for Terminal-Bench - 2.0. Versions are kept separate, and incompatible results are never combined.
GPT-5.5by OpenAI
2026-04-23
openai/gpt-5.5

Default frontier GPT for coding, computer use, research, and knowledge work

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$5/M
Output
$30/M
GPT-5.4by OpenAI
2026-03-05
openai/gpt-5.4

Agent-ready GPT for coding and computer-use workflows at a lower cost

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$2.5/M
Output
$15/M
Qwen3.7 Maxby Alibaba Qwen
2026-05-21
alibaba/qwen3.7-max

Qwen frontier model tuned for agent frameworks, coding assistants, and long tasks

T Reasoning Tools Weight access not listed
Context
1M
Input
$2.5/M
Output
$7.5/M
MAI-Code-1-Flashby Microsoft
2026-06-02
microsoft/mai-code-1-flash

Microsoft coding model built for fast, efficient assistance in everyday developer workflows

T Reasoning Tools Structured output Weight access not listed
Context
256K
Input
$0.75/M
Output
$4.5/M
Laguna XS 2.1by Poolside
2026-07-02
poolside/laguna-xs-2.1

Agentic coding model from Poolside in the XS size class for local deployment

T Reasoning Tools Open weights
Context
262.144K
Input
$0.06/M
Output
$0.12/M