xiaomi/mimo-v2.5
Open MiMo model for multimodal coding agents and long-context automation
- Context
- 1.04858M
- Input
- $0.14/M
- Output
- $0.28/M
Filter 128 source-linked models by creator, price, context, modality, and published benchmark coverage.
xiaomi/mimo-v2.5
Open MiMo model for multimodal coding agents and long-context automation
qwen/qwen3.6-27b
Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...
openai/gpt-5.1
Sharper GPT-5 generation for coding, product work, and tool-assisted tasks
openai/gpt-5.1:batch
GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...
google/gemini-3.5-flash-lite
Fast Gemini model balancing multimodal reasoning, tool use, and cost
google/gemini-3.5-flash-lite:batch
Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.
anthropic/claude-sonnet-4.5
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
anthropic/claude-sonnet-4.5:batch
Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...
moonshotai/kimi-k2.5
Earlier Kimi frontier model for long-context agents, coding, and multimodal work
openai/gpt-5
Original GPT-5 workhorse for reasoning, coding, writing, and tool workflows
openai/gpt-5:batch
GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...
kwaipilot/kat-coder-pro-v2
KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engineering and SaaS integration. It builds on the agentic coding strengths of earlier versions,...
qwen/qwen3.5-397b-a17b
The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...
z-ai/glm-4.7
GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...
qwen/qwen3.5-122b-a10b
The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of...
deepseek/deepseek-v3.2
DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...
qwen/qwen3.6-35b-a3b
Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...
inclusionai/ring-2.6-1t
Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both strong capability and operational efficiency. It is optimized for coding agents, tool...
deepseek/deepseek-v3.1-terminus
DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's...
stepfun/step-3.7-flash
Newer StepFun flash model for faster agents, coding, and multimodal prompts
mistralai/mistral-medium-3-5
Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...
anthropic/claude-haiku-4.5
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
anthropic/claude-haiku-4.5:batch
Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...
google/gemma-4-31b-it
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
google/gemma-4-31b-it:free
Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...
anthropic/claude-sonnet-4
Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),...
z-ai/glm-4.6
Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...
google/gemini-2.5-pro
Google's proven reasoning model for coding, math, and multimodal analysis
google/gemini-2.5-pro:batch
Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...
google/gemma-4-26b-a4b-it
Open Gemma instruction model for efficient chat and self-hosted deployments
google/gemma-4-26b-a4b-it:free
Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...
nvidia/nemotron-3-super-120b-a12b
Nemotron middle tier for collaborative agents and high-volume reasoning workloads
nvidia/nemotron-3-super-120b-a12b:free
NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...
openai/gpt-5-mini
Small GPT-5 for responsive agents, coding help, and everyday automation
openai/gpt-5-mini:batch
GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....
google/gemini-3.1-flash-lite-preview
Low-latency Gemini model for high-volume multimodal and agent workloads
qwen/qwen3.5-35b-a3b
The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...
openai/gpt-oss-120b
Open GPT reasoning model for self-hosted agents and controllable deployments
openai/gpt-oss-120b:free
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...
cohere/command-a
Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding use cases. Compared to other leading proprietary...
inception/mercury-2
Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving...
qwen/qwen3.5-9b
Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...
qwen/qwen3-coder-next
Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse MoE design with 80B total parameters and only 3B activated per...
cohere/north-mini-code:free
North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized...
mistralai/mistral-small-2603
Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...
qwen/qwen3-235b-a22b-thinking-2507
Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144...
mistralai/devstral-2512
Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer model supporting a 256K context window. Devstral 2 supports exploring...
deepseek/deepseek-r1
Classic open reasoning model for transparent math, coding, and deliberate problem solving
amazon/nova-2-lite-v1
Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing...
arcee-ai/trinity-large-thinking
Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloads, and reasoning tasks. Launch video: https://youtu.be/Gc82AXLa0Rg?si=4RLn6WBz33qT--B7...
| Model | Creator | Raw benchmark score | Input types | Context | Input / Output | Released | Compare |
|---|---|---|---|---|---|---|---|
| MiMo-V2.5xiaomi/mimo-v2.5 | 37.2 | 1.04858M | $0.14 / $0.28 | 2026-04-22 | |||
| Qwen: Qwen3.6 27Bqwen/qwen3.6-27b | 37.1 | 262.144K | $0.3 / $2 | Undated | |||
| GPT-5.1openai/gpt-5.1 | 36.9 | 400K | $1.25 / $10 | 2025-11-13 | |||
| OpenAI: GPT-5.1 (batch)openai/gpt-5.1:batch | 36.9 | 400K | $0.625 / $5 | Undated | |||
| Gemini 3.5 Flash Litegoogle/gemini-3.5-flash-lite | 36.5 | 1.04858M | $0.3 / $2.5 | 2026-07-21 | |||
| Google: Gemini 3.5 Flash Lite (batch)google/gemini-3.5-flash-lite:batch | 36.5 | 1.04858M | $0.15 / $1.25 | Undated | |||
| Anthropic: Claude Sonnet 4.5anthropic/claude-sonnet-4.5 | 36.4 | 1M | $3 / $15 | Undated | |||
| Anthropic: Claude Sonnet 4.5 (batch)anthropic/claude-sonnet-4.5:batch | 36.4 | 1M | $1.5 / $7.5 | Undated | |||
| Kimi K2.5moonshotai/kimi-k2.5 | 35.4 | 262.144K | $0.6 / $3 | 2026-01 | |||
| GPT-5openai/gpt-5 | 34.7 | 400K | $1.25 / $10 | 2025-08-07 | |||
| OpenAI: GPT-5 (batch)openai/gpt-5:batch | 34.7 | 400K | $0.625 / $5 | Undated | |||
| Kwaipilot: KAT-Coder-Pro V2kwaipilot/kat-coder-pro-v2 | 33.7 | 256K | $0.3 / $1.2 | Undated | |||
| Qwen: Qwen3.5 397B A17Bqwen/qwen3.5-397b-a17b | 33.7 | 262.144K | $0.39 / $2.34 | Undated | |||
| Z.ai: GLM 4.7z-ai/glm-4.7 | 33.7 | 202.752K | $0.4 / $1.75 | Undated | |||
| Qwen: Qwen3.5-122B-A10Bqwen/qwen3.5-122b-a10b | 32.3 | 262.144K | $0.26 / $2.08 | Undated | |||
| DeepSeek: DeepSeek V3.2deepseek/deepseek-v3.2 | 32.0 | 163.84K | $0.269 / $0.4 | Undated | |||
| Qwen: Qwen3.6 35B A3Bqwen/qwen3.6-35b-a3b | 31.6 | 262.144K | $0.14 / $1 | Undated | |||
| inclusionAI: Ring-2.6-1Tinclusionai/ring-2.6-1t | 30.6 | 262.144K | $0.075 / $0.625 | Undated | |||
| DeepSeek: DeepSeek V3.1 Terminusdeepseek/deepseek-v3.1-terminus | 30.4 | 131.072K | $0.27 / $1 | Undated | |||
| Step 3.7 Flashstepfun/step-3.7-flash | 30.3 | 256K | $0.185 / $1.11 | 2026-05-29 | |||
| Mistral: Mistral Medium 3.5mistralai/mistral-medium-3-5 | 29.9 | 262.144K | $1.5 / $7.5 | Undated | |||
| Anthropic: Claude Haiku 4.5anthropic/claude-haiku-4.5 | 29.6 | 200K | $1 / $5 | Undated | |||
| Anthropic: Claude Haiku 4.5 (batch)anthropic/claude-haiku-4.5:batch | 29.6 | 200K | $0.5 / $2.5 | Undated | |||
| Gemma 4 31B ITgoogle/gemma-4-31b-it | 29.4 | 262.144K | $0.1 / $0.35 | 2026-04-02 | |||
| Google: Gemma 4 31B (free)google/gemma-4-31b-it:free | 29.4 | 262.144K | Free / Free | Undated | |||
| Anthropic: Claude Sonnet 4anthropic/claude-sonnet-4 | 28.9 | 200K | $3 / $15 | Undated | |||
| Z.ai: GLM 4.6z-ai/glm-4.6 | 28.7 | 202.752K | $0.5 / $2 | Undated | |||
| Gemini 2.5 Progoogle/gemini-2.5-pro | 25.8 | 1.04858M | $1.25 / $10 | 2025-06-17 | |||
| Google: Gemini 2.5 Pro (batch)google/gemini-2.5-pro:batch | 25.8 | 1.04858M | $0.625 / $5 | Undated | |||
| Gemma 4 26B A4B ITgoogle/gemma-4-26b-a4b-it | 25.7 | 262.144K | $0.06 / $0.33 | 2026-04-02 | |||
| Google: Gemma 4 26B A4B (free)google/gemma-4-26b-a4b-it:free | 25.7 | 131.072K | Free / Free | Undated | |||
| Nemotron 3 Super 120B A12Bnvidia/nemotron-3-super-120b-a12b | 25.4 | 262.144K | $0.2 / $0.8 | 2026-03-11 | |||
| NVIDIA: Nemotron 3 Super (free)nvidia/nemotron-3-super-120b-a12b:free | 25.4 | 262.144K | Free / Free | Undated | |||
| GPT-5 Miniopenai/gpt-5-mini | 25.3 | 400K | $0.25 / $2 | 2025-08-07 | |||
| OpenAI: GPT-5 Mini (batch)openai/gpt-5-mini:batch | 25.3 | 400K | $0.125 / $1 | Undated | |||
| Gemini 3.1 Flash Lite Previewgoogle/gemini-3.1-flash-lite-preview | 25.0 | 1.04858M | $0.25 / $1.5 | 2026-03-03 | |||
| Qwen: Qwen3.5-35B-A3Bqwen/qwen3.5-35b-a3b | 24.0 | 262.144K | $0.14 / $1 | Undated | |||
| GPT OSS 120Bopenai/gpt-oss-120b | 23.8 | 131.072K | $0.03 / $0.17 | 2025-08-05 | |||
| OpenAI: gpt-oss-120b (free)openai/gpt-oss-120b:free | 23.8 | 131.072K | Free / Free | Undated | |||
| Cohere: Command Acohere/command-a | 22.5 | 256K | $2.5 / $10 | Undated | |||
| Inception: Mercury 2inception/mercury-2 | 21.4 | 128K | $0.25 / $0.75 | Undated | |||
| Qwen: Qwen3.5-9Bqwen/qwen3.5-9b | 21.4 | 262.144K | $0.1 / $0.15 | Undated | |||
| Qwen: Qwen3 Coder Nextqwen/qwen3-coder-next | 21.1 | 262.144K | $0.12 / $0.8 | Undated | |||
| Cohere: North Mini Code (free)cohere/north-mini-code:free | 19.8 | 256K | Free / Free | Undated | |||
| Mistral: Mistral Small 4mistralai/mistral-small-2603 | 19.6 | 262.144K | $0.15 / $0.6 | Undated | |||
| Qwen: Qwen3 235B A22B Thinking 2507qwen/qwen3-235b-a22b-thinking-2507 | 19.6 | 131.072K | $0.3 / $3 | Undated | |||
| Mistral: Devstral 2 2512mistralai/devstral-2512 | 19.2 | 262.144K | $0.4 / $2 | Undated | |||
| DeepSeek-R1deepseek/deepseek-r1 | 18.5 | 128K | $0.7 / $2.5 | 2025-01-20 | |||
| Amazon: Nova 2 Liteamazon/nova-2-lite-v1 | 18.2 | 1M | $0.3 / $2.5 | Undated | |||
| Arcee AI: Trinity Large Thinkingarcee-ai/trinity-large-thinking | 18.2 | 262.144K | $0.22 / $0.85 | Undated |