72 results Clear filters
Ranked by raw score for Design Arena: asciiart - . Versions are kept separate, and incompatible results are never combined.
Undated
anthropic/claude-haiku-4.5

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

T Weight access not listed
Context
200K
Input
$1/M
Output
$5/M
anthropic/claude-haiku-4.5:batch

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

T Weight access not listed
Context
200K
Input
$0.5/M
Output
$2.5/M
Undated
minimax/minimax-m2.7

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...

T Weight access not listed
Context
196.608K
Input
$0.25/M
Output
$1/M
MiMo-V2.5by Xiaomi
2026-04-22
xiaomi/mimo-v2.5

Open MiMo model for multimodal coding agents and long-context automation

T Reasoning Tools Open weights
Context
1.04858M
Input
$0.14/M
Output
$0.28/M
Qwen: Qwen3 Maxby Alibaba Qwen
Undated
qwen/qwen3-max

Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It...

T Weight access not listed
Context
262.144K
Input
$0.78/M
Output
$3.9/M
Qwen: Qwen3.6 Plusby Alibaba Qwen
Undated
qwen/qwen3.6-plus

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...

T Weight access not listed
Context
1M
Input
$0.325/M
Output
$1.95/M
GPT-5 Miniby OpenAI
2025-08-07
openai/gpt-5-mini

Small GPT-5 for responsive agents, coding help, and everyday automation

T Reasoning Tools Structured output Weight access not listed
Context
400K
Input
$0.25/M
Output
$2/M
Undated
openai/gpt-5-mini:batch

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....

T Weight access not listed
Context
400K
Input
$0.125/M
Output
$1/M
GPT-5.1by OpenAI
2025-11-13
openai/gpt-5.1

Sharper GPT-5 generation for coding, product work, and tool-assisted tasks

T Reasoning Tools Structured output Weight access not listed
Context
400K
Input
$1.25/M
Output
$10/M
Undated
openai/gpt-5.1:batch

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...

T Weight access not listed
Context
400K
Input
$0.625/M
Output
$5/M
2026-04-24
deepseek/deepseek-v4-flash

Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work

T Reasoning Tools Structured output Open weights
Context
1M
Input
$0.14/M
Output
$0.28/M
2025-11-13
openai/gpt-5.1-codex-mini

Coding-optimized GPT model for repository edits, reviews, and agentic software work

T Reasoning Tools Weight access not listed
Context
400K
Input
$0.22/M
Output
$1.8/M
Undated
z-ai/glm-5v-turbo

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...

T Weight access not listed
Context
202.752K
Input
$1.2/M
Output
$4/M
Undated
nex-agi/nex-n2-pro

Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total. Built on the Qwen3.5 architecture, it accepts text and image input and produces...

T Weight access not listed
Context
262.144K
Input
$0.25/M
Output
$1/M
Undated
qwen/qwen3.5-plus-02-15

The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of...

T Weight access not listed
Context
1M
Input
$0.26/M
Output
$1.56/M
Undated
deepseek/deepseek-v3.2

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...

T Weight access not listed
Context
163.84K
Input
$0.269/M
Output
$0.4/M
2026-06-04
nvidia/nemotron-3-ultra-550b-a55b

Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy

T Reasoning Tools Structured output Open weights
Context
1M
Input
$0.5/M
Output
$2.5/M
nvidia/nemotron-3-ultra-550b-a55b:free

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

T Weight access not listed
Context
1M
Input
Free
Output
Free
Undated
mistralai/mistral-large-2512

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.

T Weight access not listed
Context
262.144K
Input
$0.5/M
Output
$1.5/M
arcee-ai/trinity-large-thinking

Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloads, and reasoning tasks. Launch video: https://youtu.be/Gc82AXLa0Rg?si=4RLn6WBz33qT--B7...

T Weight access not listed
Context
262.144K
Input
$0.22/M
Output
$0.85/M
Undated
inception/mercury-2

Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving...

T Weight access not listed
Context
128K
Input
$0.25/M
Output
$0.75/M
Undated
mistralai/mistral-medium-3.1

Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...

T Weight access not listed
Context
131.072K
Input
$0.4/M
Output
$2/M