133 results Clear filters
Ranked by raw score for Design Arena: website - . Versions are kept separate, and incompatible results are never combined.
2026-06-04
nvidia/nemotron-3-ultra-550b-a55b

Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy

T Reasoning Tools Structured output Open weights
Context
1M
Input
$0.5/M
Output
$2.5/M
nvidia/nemotron-3-ultra-550b-a55b:free

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

T Weight access not listed
Context
1M
Input
Free
Output
Free
GPT-5 Nanoby OpenAI
2025-08-07
openai/gpt-5-nano

Tiny GPT-5 lane for routing, extraction, classification, and bulk jobs

T Reasoning Tools Structured output Weight access not listed
Context
400K
Input
$0.05/M
Output
$0.4/M
Undated
openai/gpt-5-nano:batch

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger...

T Weight access not listed
Context
400K
Input
$0.025/M
Output
$0.2/M
Undated
qwen/qwen3-coder-30b-a3b-instruct

Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use. Built on the...

T Weight access not listed
Context
160K
Input
$0.07/M
Output
$0.27/M
2026-03-03
google/gemini-3.1-flash-lite-preview

Low-latency Gemini model for high-volume multimodal and agent workloads

T Reasoning Tools Structured output Weight access not listed
Context
1.04858M
Input
$0.25/M
Output
$1.5/M
Undated
mistralai/ministral-14b-2512

The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language...

T Weight access not listed
Context
262.144K
Input
$0.2/M
Output
$0.2/M
Undated
mistralai/mistral-medium-3

Mistral Medium 3 is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances state-of-the-art reasoning and multimodal performance with 8× lower cost...

T Weight access not listed
Context
131.072K
Input
$0.4/M
Output
$2/M
Undated
mistralai/ministral-8b-2512

A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.

T Weight access not listed
Context
262.144K
Input
$0.15/M
Output
$0.15/M
Undated
qwen/qwen3-235b-a22b-2507

Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following,...

T Weight access not listed
Context
262.144K
Input
$0.09/M
Output
$0.55/M
Undated
qwen/qwen3-235b-a22b-thinking-2507

Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144...

T Weight access not listed
Context
131.072K
Input
$0.23/M
Output
$2.3/M
Undated
moonshotai/kimi-k2

Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It is optimized for...

T Weight access not listed
Context
131.072K
Input
$0.57/M
Output
$2.3/M
GPT-4.1by OpenAI
2025-04-14
openai/gpt-4.1

Long-lived GPT workhorse for coding, instruction following, and production apps

T Tools Structured output Weight access not listed
Context
1.04758M
Input
$2/M
Output
$8/M
o3by OpenAI
2025-04-16
openai/o3

Deliberate o-series reasoner for hard math, coding, and multi-step analysis

T Reasoning Tools Structured output Weight access not listed
Context
200K
Input
$2/M
Output
$8/M
Qwen: Qwen3 235B A22Bby Alibaba Qwen
Undated
qwen/qwen3-235b-a22b

Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It supports seamless switching between a "thinking" mode for complex reasoning, math, and...

T Weight access not listed
Context
131.072K
Input
$0.455/M
Output
$1.82/M
Undated
mistralai/ministral-3b-2512

The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.

T Weight access not listed
Context
131.072K
Input
$0.1/M
Output
$0.1/M
Undated
mistralai/codestral-2508

Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation. [Blog Post](https://mistral.ai/news/codestral-25-08)

T Weight access not listed
Context
256K
Input
$0.3/M
Output
$0.9/M
GPT-4.1 miniby OpenAI
2025-04-14
openai/gpt-4.1-mini

Affordable GPT-4.1 lane for fast coding help and structured extraction

T Tools Structured output Weight access not listed
Context
1.04758M
Input
$0.4/M
Output
$1.6/M
Undated
inception/mercury-2

Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving...

T Weight access not listed
Context
128K
Input
$0.25/M
Output
$0.75/M
o4-miniby OpenAI
2025-04-16
openai/o4-mini

Fast o-series model for compact reasoning, coding, and tool use

T Reasoning Tools Structured output Weight access not listed
Context
200K
Input
$1.1/M
Output
$4.4/M
Undated
openai/gpt-oss-120b:free

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...

T Weight access not listed
Context
131.072K
Input
Free
Output
Free
GPT-4.1 nanoby OpenAI
2025-04-14
openai/gpt-4.1-nano

Tiny GPT-4.1 option for classification, routing, and very high-volume tasks

T Tools Weight access not listed
Context
1.04758M
Input
$0.1/M
Output
$0.4/M
GPT OSS 120Bby OpenAI
2025-08-05
openai/gpt-oss-120b

Open GPT reasoning model for self-hosted agents and controllable deployments

T Reasoning Tools Open weights
Context
131.072K
Input
$0.03/M
Output
$0.17/M
Qwen: Qwen3 30B A3Bby Alibaba Qwen
Undated
qwen/qwen3-30b-a3b

Qwen3, the latest generation in the Qwen large language model series, features both dense and mixture-of-experts (MoE) architectures to excel in reasoning, multilingual support, and advanced agent tasks. Its unique...

T Weight access not listed
Context
40.96K
Input
$0.12/M
Output
$0.5/M
Undated
qwen/qwen3-30b-a3b-thinking-2507

Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step thinking. The model is designed specifically for “thinking mode,” where internal reasoning traces are separated...

T Weight access not listed
Context
81.92K
Input
$0.2/M
Output
$2.4/M
Undated
mistralai/mistral-small-3.2-24b-instruct

Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduction, and improved function calling. Compared to the 3.1 release, version 3.2 significantly improves accuracy on...

T Weight access not listed
Context
131.072K
Input
$0.1/M
Output
$0.3/M
Undated
meta-llama/llama-4-maverick

Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...

T Weight access not listed
Context
1.04858M
Input
$0.2/M
Output
$0.8/M
GPT OSS 20Bby OpenAI
2025-08-05
openai/gpt-oss-20b

Open GPT reasoning model for self-hosted agents and controllable deployments

T Reasoning Tools Open weights
Context
131.072K
Input
$0.03/M
Output
$0.13/M
Undated
openai/gpt-oss-20b:free

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...

T Weight access not listed
Context
131.072K
Input
Free
Output
Free
Undated
amazon/nova-premier-v1

Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks and for use as the best teacher for distilling custom models.

T Weight access not listed
Context
1M
Input
$2.5/M
Output
$12.5/M
GPT-4oby OpenAI
2024-05-13
openai/gpt-4o

Omni-era GPT for multimodal chat, practical coding, and general assistants

T Tools Weight access not listed
Context
128K
Input
$2.5/M
Output
$10/M
Undated
amazon/nova-pro-v1

Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide range of tasks. As of December...

T Weight access not listed
Context
300K
Input
$0.8/M
Output
$3.2/M
Undated
meta-llama/llama-4-scout

Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...

T Weight access not listed
Context
327.68K
Input
$0.1/M
Output
$0.3/M