46 results Clear filters
Ranked by raw score for Design Arena: fullstack - . Versions are kept separate, and incompatible results are never combined.
Undated
anthropic/claude-opus-4.7

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

T Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
anthropic/claude-opus-4.7:batch

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

T Weight access not listed
Context
1M
Input
$2.5/M
Output
$12.5/M
Claude Fable 5by Anthropic
2026-06-09
anthropic/claude-fable-5

Claude model for creative writing, analysis, and controlled agent workflows

T Reasoning Tools Weight access not listed
Context
1M
Input
$10/M
Output
$50/M
Undated
anthropic/claude-fable-5:batch

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

T Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
Undated
anthropic/claude-opus-4.8

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

T Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
anthropic/claude-opus-4.8:batch

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

T Weight access not listed
Context
1M
Input
$2.5/M
Output
$12.5/M
Undated
x-ai/grok-4.5

Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

T Weight access not listed
Context
500K
Input
$2/M
Output
$6/M
Undated
z-ai/glm-5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

T Weight access not listed
Context
1.024M
Input
$0.966/M
Output
$3.036/M
Claude Sonnet 5by Anthropic
2026-06-30
anthropic/claude-sonnet-5

Everyday Claude agent model for coding, planning, browsing, and general work

T Reasoning Tools Weight access not listed
Context
1M
Input
$2/M
Output
$10/M
anthropic/claude-sonnet-5:batch

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

T Weight access not listed
Context
1M
Input
$1/M
Output
$5/M
Undated
anthropic/claude-opus-4.6

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...

T Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
anthropic/claude-opus-4.6:batch

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...

T Weight access not listed
Context
1M
Input
$2.5/M
Output
$12.5/M
Undated
anthropic/claude-sonnet-4.6

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...

T Weight access not listed
Context
1M
Input
$3/M
Output
$15/M
Undated
minimax/minimax-m3

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

T Weight access not listed
Context
524.288K
Input
$0.3/M
Output
$1.2/M
Undated
minimax/minimax-m3:batch

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

T Weight access not listed
Context
524.288K
Input
$0.15/M
Output
$0.6/M
2026-05-19
google/gemini-3.5-flash

Fast Gemini model balancing multimodal reasoning, tool use, and cost

T Reasoning Tools Structured output Weight access not listed
Context
1.04858M
Input
$1.5/M
Output
$9/M
google/gemini-3.5-flash:batch

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

T Weight access not listed
Context
1.04858M
Input
$0.75/M
Output
$4.5/M
Kimi K2.7 Codeby Moonshot AI
2026-06-12
moonshotai/kimi-k2.7-code

Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking

T Reasoning Tools Open weights
Context
262.144K
Input
$0.95/M
Output
$4/M
Qwen: Qwen3.7 Maxby Alibaba Qwen
Undated
qwen/qwen3.7-max

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...

T Weight access not listed
Context
1M
Input
$1.475/M
Output
$4.425/M
Undated
z-ai/glm-5.1

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...

T Weight access not listed
Context
200K
Input
$0.966/M
Output
$3.036/M
Kimi K2.6by Moonshot AI
2026-04-21
moonshotai/kimi-k2.6

Multimodal Kimi workhorse for agent loops, coding tasks, and visual context

T Reasoning Tools Structured output Open weights
Context
262.144K
Input
$0.95/M
Output
$4/M
Undated
anthropic/claude-opus-4.5

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...

T Weight access not listed
Context
200K
Input
$5/M
Output
$25/M
anthropic/claude-opus-4.5:batch

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...

T Weight access not listed
Context
200K
Input
$2.5/M
Output
$12.5/M
Undated
z-ai/glm-5v-turbo

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...

T Weight access not listed
Context
202.752K
Input
$1.2/M
Output
$4/M
Undated
z-ai/glm-5

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading...

T Weight access not listed
Context
204.8K
Input
$0.95/M
Output
$2.55/M
Kimi K2.5by Moonshot AI
2026-01
moonshotai/kimi-k2.5

Earlier Kimi frontier model for long-context agents, coding, and multimodal work

T Reasoning Tools Structured output Open weights
Context
262.144K
Input
$0.6/M
Output
$3/M
GPT-5.5by OpenAI
2026-04-23
openai/gpt-5.5

Default frontier GPT for coding, computer use, research, and knowledge work

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$5/M
Output
$30/M
Undated
openai/gpt-5.5:batch

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...

T Weight access not listed
Context
1.05M
Input
$2.5/M
Output
$15/M
2026-02-19
google/gemini-3.1-pro-preview

Reasoning-first Gemini preview for agentic coding and complex problem solving

T Reasoning Tools Structured output Weight access not listed
Context
1.04858M
Input
$2/M
Output
$12/M
google/gemini-3.1-pro-preview:batch

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...

T Weight access not listed
Context
1.04858M
Input
$1/M
Output
$6/M
2025-12-17
google/gemini-3-flash-preview

New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs

T Reasoning Tools Structured output Weight access not listed
Context
1.04858M
Input
$0.5/M
Output
$3/M
google/gemini-3-flash-preview:batch

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...

T Weight access not listed
Context
1.04858M
Input
$0.25/M
Output
$1.5/M
Undated
x-ai/grok-4.20

Grok 4.20 is a reasoning model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...

T Weight access not listed
Context
2M
Input
$1.25/M
Output
$2.5/M
Undated
anthropic/claude-sonnet-4.5

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

T Weight access not listed
Context
1M
Input
$3/M
Output
$15/M
anthropic/claude-sonnet-4.5:batch

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

T Weight access not listed
Context
1M
Input
$1.5/M
Output
$7.5/M
Undated
z-ai/glm-4.7

GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...

T Weight access not listed
Context
202.752K
Input
$0.4/M
Output
$1.75/M
GPT-5.1 Codexby OpenAI
2025-11-13
openai/gpt-5.1-codex

Codex GPT for repository edits, code review, and practical software agents

T Reasoning Tools Weight access not listed
Context
400K
Input
$1.07/M
Output
$8.5/M
GPT-5.2by OpenAI
2025-12-11
openai/gpt-5.2

Reliable GPT generation for broad coding, writing, and tool-assisted product work

T Reasoning Tools Structured output Weight access not listed
Context
400K
Input
$1.75/M
Output
$14/M
Undated
openai/gpt-5.2:batch

GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly...

T Weight access not listed
Context
400K
Input
$0.875/M
Output
$7/M
Undated
z-ai/glm-4.6

Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...

T Weight access not listed
Context
202.752K
Input
$0.5/M
Output
$2/M
GPT-5.4by OpenAI
2026-03-05
openai/gpt-5.4

Agent-ready GPT for coding and computer-use workflows at a lower cost

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$2.5/M
Output
$15/M
Undated
openai/gpt-5.4:batch

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for...

T Weight access not listed
Context
1.05M
Input
$1.25/M
Output
$7.5/M
Undated
x-ai/grok-4.3

Grok 4.3 is a reasoning model from xAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

T Weight access not listed
Context
1M
Input
$1.25/M
Output
$2.5/M
GPT-5.2 Codexby OpenAI
2025-12-11
openai/gpt-5.2-codex

Code-specialist GPT for repository edits, reviews, and long-running software agents

T Reasoning Tools Weight access not listed
Context
400K
Input
$0.14/M
Output
$1.14/M
GPT-5.3 Codexby OpenAI
2026-02-05
openai/gpt-5.3-codex

Coding-optimized GPT model for repository edits, reviews, and agentic software work

T Reasoning Tools Weight access not listed
Context
400K
Input
$1.75/M
Output
$14/M
DeepSeek V4 Proby DeepSeek
2026-04-24
deepseek/deepseek-v4-pro

Open MoE flagship with million-token context for coding and long agent runs

T Reasoning Tools Structured output Open weights
Context
1M
Input
$0.435/M
Output
$0.87/M