14 results Clear filters
Ranked by raw score for GPQA Diamond - . Versions are kept separate, and incompatible results are never combined.
Fuguby Sakana
2026-06-15
sakana/fugu

Multi-agent model for routing expert agents across complex analytical tasks

T Reasoning Tools Structured output Weight access not listed
Context
1M
Input
-
Output
-
Fugu Ultraby Sakana
2026-06-15
sakana/fugu-ultra

Quality-first multi-agent model for hard research, analysis, and competitions

T Reasoning Tools Structured output Weight access not listed
Context
1M
Input
$5/M
Output
$30/M
GPT-5.6 Solby OpenAI
2026-07-09
openai/gpt-5.6-sol

Frontier GPT-5.6 model for complex professional work, coding, and agentic workflows

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$5/M
Output
$30/M
GPT-5.4 Proby OpenAI
2026-03-05
openai/gpt-5.4-pro

More exact GPT-5.4 tier for demanding professional reasoning and agent tasks

T Reasoning Tools Weight access not listed
Context
1.05M
Input
$30/M
Output
$180/M
2026-02-19
google/gemini-3.1-pro-preview

Reasoning-first Gemini preview for agentic coding and complex problem solving

T Reasoning Tools Structured output Weight access not listed
Context
1.04858M
Input
$2/M
Output
$12/M
Claude Opus 4.7by Anthropic
2026-04-16
anthropic/claude-opus-4-7

Stronger Opus tier for advanced software work and high-stakes reasoning

T Reasoning Tools Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
GPT-5.5by OpenAI
2026-04-23
openai/gpt-5.5

Default frontier GPT for coding, computer use, research, and knowledge work

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$5/M
Output
$30/M
GPT-5.6 Terraby OpenAI
2026-07-09
openai/gpt-5.6-terra

Balanced GPT-5.6 model for capable, cost-efficient everyday work

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$2.5/M
Output
$15/M
GPT-5.4by OpenAI
2026-03-05
openai/gpt-5.4

Agent-ready GPT for coding and computer-use workflows at a lower cost

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$2.5/M
Output
$15/M
Qwen3.7 Maxby Alibaba Qwen
2026-05-21
alibaba/qwen3.7-max

Qwen frontier model tuned for agent frameworks, coding assistants, and long tasks

T Reasoning Tools Weight access not listed
Context
1M
Input
$2.5/M
Output
$7.5/M
GPT-5.6 Lunaby OpenAI
2026-07-09
openai/gpt-5.6-luna

Cost-efficient GPT-5.6 model for fast, high-volume workloads

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$1/M
Output
$6/M
LongCat-2.0by Meituan
2026-06-30
meituan/longcat-2.0

Meituan LongCat-2.0, a reasoning model with tool calling and a 1M-token context window

T Reasoning Tools Weight access not listed
Context
1M
Input
$0.3/M
Output
$1.2/M
MiMo-V2.5-Proby Xiaomi
2026-04-22
xiaomi/mimo-v2.5-pro

Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution

T Reasoning Tools Open weights
Context
1.04858M
Input
$0.435/M
Output
$0.87/M
MAI-Code-1-Flashby Microsoft
2026-06-02
microsoft/mai-code-1-flash

Microsoft coding model built for fast, efficient assistance in everyday developer workflows

T Reasoning Tools Structured output Weight access not listed
Context
256K
Input
$0.75/M
Output
$4.5/M