494 results Clear filters
Undated
deepseek/deepseek-v3.2

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...

T Weight access not listed
Context
163.84K
Input
$0.269/M
Output
$0.4/M
Undated
anthropic/claude-opus-4.5

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...

T Weight access not listed
Context
200K
Input
$5/M
Output
$25/M
Undated
deepcogito/cogito-v2.1-671b

Cogito v2.1 671B MoE represents one of the strongest open models globally, matching performance of frontier closed and open models. This model is trained using self play with reinforcement learning...

T Weight access not listed
Context
128K
Input
$1.25/M
Output
$1.25/M
Undated
openai/gpt-5.1-chat

GPT-5.1 Chat (AKA Instant is the fast, lightweight member of the 5.1 family, optimized for low-latency chat while retaining strong general intelligence. It uses adaptive reasoning to selectively “think” on...

T Weight access not listed
Context
128K
Input
$1.25/M
Output
$10/M
Undated
amazon/nova-premier-v1

Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks and for use as the best teacher for distilling custom models.

T Weight access not listed
Context
1M
Input
$2.5/M
Output
$12.5/M
Undated
perplexity/sonar-pro-search

Exclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system. It is designed for deeper reasoning and analysis. Pricing is based...

T Weight access not listed
Context
200K
Input
$3/M
Output
$15/M
openai/gpt-oss-safeguard-20b

gpt-oss-safeguard-20b is a safety reasoning model from OpenAI built upon gpt-oss-20b. This open-weight, 21B-parameter Mixture-of-Experts (MoE) model offers lower latency for safety tasks like content classification, LLM filtering, and trust...

T Weight access not listed
Context
131.072K
Input
$0.075/M
Output
$0.3/M
nvidia/nemotron-nano-12b-v2-vl:free

NVIDIA Nemotron Nano 2 VL is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligence. It introduces a hybrid Transformer-Mamba architecture, combining transformer-level accuracy with Mamba’s...

T Weight access not listed
Context
128K
Input
Free
Output
Free
Undated
minimax/minimax-m2

MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning,...

T Weight access not listed
Context
204.8K
Input
$0.255/M
Output
$1.02/M
Undated
qwen/qwen3-vl-32b-instruct

Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. With 32 billion parameters, it combines deep visual perception with advanced text...

T Weight access not listed
Context
131.072K
Input
$0.104/M
Output
$0.416/M
Undated
ibm-granite/granite-4.0-h-micro

Granite-4.0-H-Micro is a 3B parameter from the Granite 4 family of models. These models are the latest in a series of models released by IBM. They are fine-tuned for long...

T Weight access not listed
Context
131K
Input
$0.017/M
Output
$0.112/M
Undated
openai/gpt-5-image-mini

GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by [GPT-5 Mini](https://openrouter.ai/openai/gpt-5-mini), with GPT Image 1 Mini for efficient image generation. This natively multimodal model features superior instruction following, text...

T Weight access not listed
Context
400K
Input
$2.5/M
Output
$2/M
Undated
anthropic/claude-haiku-4.5

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

T Weight access not listed
Context
200K
Input
$1/M
Output
$5/M
Undated
qwen/qwen3-vl-8b-thinking

Qwen3-VL-8B-Thinking is the reasoning-optimized variant of the Qwen3-VL-8B multimodal model, designed for advanced visual and textual reasoning across complex scenes, documents, and temporal sequences. It integrates enhanced multimodal alignment and...

T Weight access not listed
Context
131.072K
Input
$0.18/M
Output
$2.1/M
Undated
qwen/qwen3-vl-8b-instruct

Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon...

T Weight access not listed
Context
131.072K
Input
$0.117/M
Output
$0.455/M
Undated
openai/gpt-5-image

[GPT-5](https://openrouter.ai/openai/gpt-5) Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It offers major improvements in reasoning, code quality, and user experience while incorporating GPT Image 1's superior instruction following,...

T Weight access not listed
Context
400K
Input
$10/M
Output
$10/M
Undated
qwen/qwen3-vl-30b-a3b-thinking

Qwen3-VL-30B-A3B-Thinking is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Thinking variant enhances reasoning in STEM, math, and complex tasks. It excels...

T Weight access not listed
Context
131.072K
Input
$0.2/M
Output
$2.4/M
Undated
qwen/qwen3-vl-30b-a3b-instruct

Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception...

T Weight access not listed
Context
262.144K
Input
$0.15/M
Output
$0.6/M
Undated
z-ai/glm-4.6

Compared with GLM-4.5, this generation brings several key improvements: Longer context window: The context window has been expanded from 128K to 200K tokens, enabling the model to handle more complex...

T Weight access not listed
Context
202.752K
Input
$0.5/M
Output
$2/M
Undated
anthropic/claude-sonnet-4.5

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

T Weight access not listed
Context
1M
Input
$3/M
Output
$15/M
Undated
deepseek/deepseek-v3.2-exp

DeepSeek-V3.2-Exp is an experimental large language model released by DeepSeek as an intermediate step between V3.1 and future architectures. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...

T Weight access not listed
Context
163.84K
Input
$0.27/M
Output
$0.41/M
Undated
thedrummer/cydonia-24b-v4.1

Uncensored and creative writing model based on Mistral Small 3.2 24B with good recall, prompt adherence, and intelligence.

T Weight access not listed
Context
131.072K
Input
$0.3/M
Output
$0.5/M
Undated
relace/relace-apply-3

Relace Apply 3 is a specialized code-patching LLM that merges AI-suggested edits straight into your source files. It can apply updates from GPT-4o, Claude, and others into your files at...

T Weight access not listed
Context
256K
Input
$0.85/M
Output
$1.25/M
Undated
qwen/qwen3-vl-235b-a22b-thinking

Qwen3-VL-235B-A22B Thinking is a multimodal model that unifies strong text generation with visual understanding across images and video. The Thinking model is optimized for multimodal reasoning in STEM and math....

T Weight access not listed
Context
131.072K
Input
$0.4/M
Output
$4/M
Undated
qwen/qwen3-vl-235b-a22b-instruct

Qwen3-VL-235B-A22B Instruct is an open-weight multimodal model that unifies strong text generation with visual understanding across images and video. The Instruct model targets general vision-language use (VQA, document parsing, chart/table...

T Weight access not listed
Context
131.072K
Input
$0.21/M
Output
$1.9/M
Qwen: Qwen3 Maxby Alibaba Qwen
Undated
qwen/qwen3-max

Qwen3-Max is an updated release built on the Qwen3 series, offering major improvements in reasoning, instruction following, multilingual support, and long-tail knowledge coverage compared to the January 2025 version. It...

T Weight access not listed
Context
262.144K
Input
$0.78/M
Output
$3.9/M
Qwen: Qwen3 Coder Plusby Alibaba Qwen
Undated
qwen/qwen3-coder-plus

Qwen3 Coder Plus is Alibaba's proprietary version of the Open Source Qwen3 Coder 480B A35B. It is a powerful coding agent model specializing in autonomous programming via tool calling and...

T Weight access not listed
Context
1M
Input
$0.65/M
Output
$3.25/M
deepseek/deepseek-v3.1-terminus

DeepSeek-V3.1 Terminus is an update to [DeepSeek V3.1](/deepseek/deepseek-chat-v3.1) that maintains the model's original capabilities while addressing issues reported by users, including language consistency and agent capabilities, further optimizing the model's...

T Weight access not listed
Context
131.072K
Input
$0.27/M
Output
$1/M
Undated
qwen/qwen3-coder-flash

Qwen3 Coder Flash is Alibaba's fast and cost efficient version of their proprietary Qwen3 Coder Plus. It is a powerful coding agent model specializing in autonomous programming via tool calling...

T Weight access not listed
Context
1M
Input
$0.195/M
Output
$0.975/M
Undated
qwen/qwen3-next-80b-a3b-thinking

Qwen3-Next-80B-A3B-Thinking is a reasoning-first chat model in the Qwen3-Next line that outputs structured “thinking” traces by default. It’s designed for hard multi-step problems; math proofs, code synthesis/debugging, logic, and agentic...

T Weight access not listed
Context
131.072K
Input
$0.15/M
Output
$1.2/M
qwen/qwen3-next-80b-a3b-instruct:free

Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces. It targets complex tasks across reasoning, code generation, knowledge QA, and multilingual...

T Weight access not listed
Context
262.144K
Input
Free
Output
Free
Undated
qwen/qwen3-next-80b-a3b-instruct

Qwen3-Next-80B-A3B-Instruct is an instruction-tuned chat model in the Qwen3-Next series optimized for fast, stable responses without “thinking” traces. It targets complex tasks across reasoning, code generation, knowledge QA, and multilingual...

T Weight access not listed
Context
262.144K
Input
$0.1/M
Output
$1.1/M
Undated
qwen/qwen-plus-2025-07-28:thinking

Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, and cost combination.

T Weight access not listed
Context
1M
Input
$0.4/M
Output
$1.2/M
Qwen: Qwen Plus 0728by Alibaba Qwen
Undated
qwen/qwen-plus-2025-07-28

Qwen Plus 0728, based on the Qwen3 foundation model, is a 1 million context hybrid reasoning model with a balanced performance, speed, and cost combination.

T Weight access not listed
Context
1M
Input
$0.26/M
Output
$0.78/M
nvidia/nemotron-nano-9b-v2:free

NVIDIA-Nemotron-Nano-9B-v2 is a large language model (LLM) trained from scratch by NVIDIA, and designed as a unified model for both reasoning and non-reasoning tasks. It responds to user queries and...

T Weight access not listed
Context
128K
Input
Free
Output
Free
Undated
moonshotai/kimi-k2-0905

Kimi K2 0905 is the September update of [Kimi K2 0711](moonshotai/kimi-k2). It is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32...

T Weight access not listed
Context
262.144K
Input
$0.6/M
Output
$2.5/M
Nous: Hermes 4 70Bby Nousresearch
Undated
nousresearch/hermes-4-70b

Hermes 4 70B is a hybrid reasoning model from Nous Research, built on Meta-Llama-3.1-70B. It introduces the same hybrid mode as the larger 405B release, allowing the model to either...

T Weight access not listed
Context
131.072K
Input
$0.13/M
Output
$0.4/M
Nous: Hermes 4 405Bby Nousresearch
Undated
nousresearch/hermes-4-405b

Hermes 4 is a large-scale reasoning model built on Meta-Llama-3.1-405B and released by Nous Research. It introduces a hybrid reasoning mode, where the model can choose to deliberate internally with...

T Weight access not listed
Context
131.072K
Input
$1/M
Output
$3/M
Undated
deepseek/deepseek-chat-v3.1

DeepSeek-V3.1 is a large hybrid reasoning model (671B parameters, 37B active) that supports both thinking and non-thinking modes via prompt templates. It extends the DeepSeek-V3 base with a two-phase long-context...

T Weight access not listed
Context
163.84K
Input
$0.25/M
Output
$0.95/M
Undated
mistralai/mistral-medium-3.1

Mistral Medium 3.1 is an updated version of Mistral Medium 3, which is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances...

T Weight access not listed
Context
131.072K
Input
$0.4/M
Output
$2/M
Undated
ai21/jamba-large-1.7

Jamba Large 1.7 is the latest model in the Jamba open family, offering improvements in grounding, instruction-following, and overall efficiency. Built on a hybrid SSM-Transformer architecture with a 256K context...

T Weight access not listed
Context
256K
Input
$2/M
Output
$8/M
Undated
openai/gpt-5-chat

GPT-5 Chat is designed for advanced, natural, multimodal, and context-aware conversations for enterprise applications.

T Weight access not listed
Context
128K
Input
$1.25/M
Output
$10/M
Undated
openai/gpt-oss-120b:free

gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...

T Weight access not listed
Context
131.072K
Input
Free
Output
Free
Undated
openai/gpt-oss-20b:free

gpt-oss-20b is an open-weight 21B parameter model released by OpenAI under the Apache 2.0 license. It uses a Mixture-of-Experts (MoE) architecture with 3.6B active parameters per forward pass, optimized for...

T Weight access not listed
Context
131.072K
Input
Free
Output
Free
Undated
anthropic/claude-opus-4.1

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...

T Weight access not listed
Context
200K
Input
$15/M
Output
$75/M
Undated
mistralai/codestral-2508

Mistral's cutting-edge language model for coding released end of July 2025. Codestral specializes in low-latency, high-frequency tasks such as fill-in-the-middle (FIM), code correction and test generation. [Blog Post](https://mistral.ai/news/codestral-25-08)

T Weight access not listed
Context
256K
Input
$0.3/M
Output
$0.9/M
Undated
qwen/qwen3-coder-30b-a3b-instruct

Qwen3-Coder-30B-A3B-Instruct is a 30.5B parameter Mixture-of-Experts (MoE) model with 128 experts (8 active per forward pass), designed for advanced code generation, repository-scale understanding, and agentic tool use. Built on the...

T Weight access not listed
Context
160K
Input
$0.07/M
Output
$0.27/M
Undated
qwen/qwen3-30b-a3b-instruct-2507

Qwen3-30B-A3B-Instruct-2507 is a 30.5B-parameter mixture-of-experts language model from Qwen, with 3.3B active parameters per inference. It operates in non-thinking mode and is designed for high-quality instruction following, multilingual understanding, and...

T Weight access not listed
Context
128K
Input
$0.048/M
Output
$0.193/M
Undated
z-ai/glm-4.5

GLM-4.5 is our latest flagship foundation model, purpose-built for agent-based applications. It leverages a Mixture-of-Experts (MoE) architecture and supports a context length of up to 128k tokens. GLM-4.5 delivers significantly...

T Weight access not listed
Context
131.072K
Input
$0.6/M
Output
$2.2/M
Undated
z-ai/glm-4.5-air

GLM-4.5-Air is the lightweight variant of our latest flagship model family, also purpose-built for agent-centric applications. Like GLM-4.5, it adopts the Mixture-of-Experts (MoE) architecture but with a more compact parameter...

T Weight access not listed
Context
131.072K
Input
$0.13/M
Output
$0.85/M