moonshotai/kimi-k2-thinking
Thinking Kimi model for slower research passes, planning, and hard technical questions
- Context
- 262.144K
- Input
- $0.6/M
- Output
- $2.5/M
Filter 128 source-linked models by creator, price, context, modality, and published benchmark coverage.
moonshotai/kimi-k2-thinking
Thinking Kimi model for slower research passes, planning, and hard technical questions
qwen/qwen3-next-80b-a3b-thinking
Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic...
mistralai/mistral-large-2512
Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.
openai/o3-mini-high
OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...
deepseek/deepseek-chat-v3-0324
DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the [DeepSeek V3](/deepseek/deepseek-chat-v3) model and performs really well...
openai/gpt-oss-20b
Open GPT reasoning model for self-hosted agents and controllable deployments
openai/gpt-oss-20b:free
gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...
openai/gpt-4.1-mini
Affordable GPT-4.1 lane for fast coding help and structured extraction
mistralai/mistral-medium-3.1
Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...
qwen/qwen3-30b-a3b-thinking-2507
Qwen3-30B-A3B-Thinking-2507 is a 30B parameter Mixture-of-Experts reasoning model optimized for complex tasks requiring extended multi-step thinking. The model is designed specifically for “thinking mode,” where internal reasoning traces are separated...
meta-llama/llama-4-maverick
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...
nvidia/nemotron-3-nano-30b-a3b
Small Nemotron 3 MoE for efficient coding, math, and long-context agents
nvidia/nemotron-3-nano-30b-a3b:free
NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...
inclusionai/ling-2.6-flash
Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency....
upstage/solar-pro-3
Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass, it delivers exceptional performance while maintaining computational efficiency. Optimized...
qwen/qwen3-32b
Qwen3-32B is a dense 32.8B parameter causal language model from the Qwen3 series, optimized for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
mistralai/ministral-14b-2512
The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language...
qwen/qwen3-14b
Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...
meta-llama/llama-4-scout
Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...
openai/gpt-4.1-nano
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
meta-llama/llama-3.3-70b-instruct
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
meta-llama/llama-3.3-70b-instruct:free
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
mistralai/ministral-8b-2512
A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.
qwen/qwen3-8b
Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,...
meta-llama/llama-3.1-8b-instruct
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...
google/gemma-3-27b-it
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...
mistralai/ministral-3b-2512
The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.
google/gemma-3-12b-it
Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...
| Model | Creator | Raw benchmark score | Input types | Context | Input / Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| Kimi K2 Thinkingmoonshotai/kimi-k2-thinking | 17.3 | 262.144K | $0.6 / $2.5 | 2025-11-06 | |||
| Qwen: Qwen3 Next 80B A3B Thinkingqwen/qwen3-next-80b-a3b-thinking | 16.7 | 131.072K | $0.15 / $1.2 | Undated | |||
| Mistral: Mistral Large 3 2512mistralai/mistral-large-2512 | 15.9 | 262.144K | $0.5 / $1.5 | Undated | |||
| OpenAI: o3 Mini Highopenai/o3-mini-high | 15.6 | 200K | $1.1 / $4.4 | Undated | |||
| DeepSeek: DeepSeek V3 0324deepseek/deepseek-chat-v3-0324 | 15.4 | 163.84K | $0.27 / $1.12 | Undated | |||
| GPT OSS 20Bopenai/gpt-oss-20b | 14.9 | 131.072K | $0.03 / $0.13 | 2025-08-05 | |||
| OpenAI: gpt-oss-20b (free)openai/gpt-oss-20b:free | 14.9 | 131.072K | Free / Free | Undated | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 14.8 | 1.04758M | $0.4 / $1.6 | 2025-04-14 | |||
| Mistral: Mistral Medium 3.1mistralai/mistral-medium-3.1 | 14.7 | 131.072K | $0.4 / $2 | Undated | |||
| Qwen: Qwen3 30B A3B Thinking 2507qwen/qwen3-30b-a3b-thinking-2507 | 14.4 | 81.92K | $0.2 / $2.4 | Undated | |||
| Meta: Llama 4 Maverickmeta-llama/llama-4-maverick | 14.3 | 1.04858M | $0.2 / $0.8 | Undated | |||
| Nemotron 3 Nano 30B A3Bnvidia/nemotron-3-nano-30b-a3b | 14.2 | 262.144K | $0.05 / $0.2 | 2025-12-15 | |||
| NVIDIA: Nemotron 3 Nano 30B A3B (free)nvidia/nemotron-3-nano-30b-a3b:free | 14.2 | 256K | Free / Free | Undated | |||
| inclusionAI: Ling-2.6-flashinclusionai/ling-2.6-flash | 14.1 | 262.144K | $0.01 / $0.03 | Undated | |||
| Upstage: Solar Pro 3upstage/solar-pro-3 | 14.1 | 128K | $0.15 / $0.6 | Undated | |||
| Qwen: Qwen3 32Bqwen/qwen3-32b | 11.5 | 40.96K | $0.08 / $0.28 | Undated | |||
| Mistral: Ministral 3 14B 2512mistralai/ministral-14b-2512 | 11.1 | 262.144K | $0.2 / $0.2 | Undated | |||
| Qwen: Qwen3 14Bqwen/qwen3-14b | 10.4 | 131.072K | $0.227 / $0.91 | Undated | |||
| Meta: Llama 4 Scoutmeta-llama/llama-4-scout | 10.0 | 327.68K | $0.1 / $0.3 | Undated | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 9.6 | 1.04758M | $0.1 / $0.4 | 2025-04-14 | |||
| Meta: Llama 3.3 70B Instructmeta-llama/llama-3.3-70b-instruct | 9.4 | 131.072K | $0.13 / $0.4 | Undated | |||
| Meta: Llama 3.3 70B Instruct (free)meta-llama/llama-3.3-70b-instruct:free | 9.4 | 65.536K | Free / Free | Undated | |||
| Mistral: Ministral 3 8B 2512mistralai/ministral-8b-2512 | 9.0 | 262.144K | $0.15 / $0.15 | Undated | |||
| Qwen: Qwen3 8Bqwen/qwen3-8b | 8.3 | 131.072K | $0.117 / $0.455 | Undated | |||
| Meta: Llama 3.1 8B Instructmeta-llama/llama-3.1-8b-instruct | 7.6 | 131.072K | $0.05 / $0.08 | Undated | |||
| Google: Gemma 3 27Bgoogle/gemma-3-27b-it | 7.4 | 131.072K | $0.08 / $0.45 | Undated | |||
| Mistral: Ministral 3 3B 2512mistralai/ministral-3b-2512 | 6.8 | 131.072K | $0.1 / $0.1 | Undated | |||
| Google: Gemma 3 12Bgoogle/gemma-3-12b-it | 5.5 | 131.072K | $0.05 / $0.15 | Undated |