384 results Clear filters
Undated
qwen/qwen3.5-plus-20260420

Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This...

T Weight access not listed
Context
1M
Input
$0.3/M
Output
$1.8/M
Qwen: Qwen3.6 Flashby Alibaba Qwen
Undated
qwen/qwen3.6-flash

Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in...

T Weight access not listed
Context
1M
Input
$0.188/M
Output
$1.125/M
Qwen: Qwen3.6 35B A3Bby Alibaba Qwen
Undated
qwen/qwen3.6-35b-a3b

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...

T Weight access not listed
Context
262.144K
Input
$0.14/M
Output
$1/M
Qwen: Qwen3.6 27Bby Alibaba Qwen
Undated
qwen/qwen3.6-27b

Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...

T Weight access not listed
Context
262.144K
Input
$0.3/M
Output
$2/M
Undated
openai/gpt-5.4-image-2

[GPT-5.4](https://openrouter.ai/openai/gpt-5.4) Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2. It enables rich multimodal workflows, allowing users to seamlessly move between reasoning, coding, and...

T Weight access not listed
Context
272K
Input
$8/M
Output
$15/M
Undated
~anthropic/claude-opus-latest

This model always redirects to the latest model in the Claude Opus family.

T Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
Undated
anthropic/claude-opus-4.7

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

T Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
google/gemma-4-26b-a4b-it:free

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

T Weight access not listed
Context
131.072K
Input
Free
Output
Free
Undated
google/gemma-4-31b-it:free

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

T Weight access not listed
Context
262.144K
Input
Free
Output
Free
Qwen: Qwen3.6 Plusby Alibaba Qwen
Undated
qwen/qwen3.6-plus

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...

T Weight access not listed
Context
1M
Input
$0.325/M
Output
$1.95/M
Undated
z-ai/glm-5v-turbo

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...

T Weight access not listed
Context
202.752K
Input
$1.2/M
Output
$4/M
x-ai/grok-4.20-multi-agent

Grok 4.20 Multi-Agent is a variant of xAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and synthesize information...

T Weight access not listed
Context
2M
Input
$1.25/M
Output
$2.5/M
Undated
x-ai/grok-4.20

Grok 4.20 is a reasoning model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...

T Weight access not listed
Context
2M
Input
$1.25/M
Output
$2.5/M
Reka Edgeby Rekaai
Undated
rekaai/reka-edge

Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. This model is optimized specifically to deliver industry-leading performance in image understanding,...

T Weight access not listed
Context
16.384K
Input
$0.1/M
Output
$0.1/M
Undated
mistralai/mistral-small-2603

Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...

T Weight access not listed
Context
262.144K
Input
$0.15/M
Output
$0.6/M
Undated
bytedance-seed/seed-2.0-lite

Seed-2.0-Lite is a versatile, cost‑efficient enterprise workhorse that delivers strong multimodal and agent capabilities while offering noticeably lower latency, making it a practical default choice for most production workloads across...

T Weight access not listed
Context
262.144K
Input
$0.25/M
Output
$2/M
Qwen: Qwen3.5-9Bby Alibaba Qwen
Undated
qwen/qwen3.5-9b

Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...

T Weight access not listed
Context
262.144K
Input
$0.1/M
Output
$0.15/M
Undated
openai/gpt-5.3-chat

GPT-5.3 Chat is an update to ChatGPT's most-used model that makes everyday conversations smoother, more useful, and more directly helpful. It delivers more accurate answers with better contextualization and significantly...

T Weight access not listed
Context
128K
Input
$1.75/M
Output
$14/M
Undated
bytedance-seed/seed-2.0-mini

Seed-2.0-mini targets latency-sensitive, high-concurrency, and cost-sensitive scenarios, emphasizing fast response and flexible inference deployment. It delivers performance comparable to ByteDance-Seed-1.6, supports 256k context, four reasoning effort modes (minimal/low/medium/high), multimodal understanding,...

T Weight access not listed
Context
262.144K
Input
$0.1/M
Output
$0.4/M
Qwen: Qwen3.5-35B-A3Bby Alibaba Qwen
Undated
qwen/qwen3.5-35b-a3b

The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...

T Weight access not listed
Context
262.144K
Input
$0.14/M
Output
$1/M
Qwen: Qwen3.5-27Bby Alibaba Qwen
Undated
qwen/qwen3.5-27b

The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of...

T Weight access not listed
Context
262.144K
Input
$0.195/M
Output
$1.56/M
Undated
qwen/qwen3.5-122b-a10b

The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of...

T Weight access not listed
Context
262.144K
Input
$0.26/M
Output
$2.08/M
Qwen: Qwen3.5-Flashby Alibaba Qwen
Undated
qwen/qwen3.5-flash-02-23

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the...

T Weight access not listed
Context
1M
Input
$0.065/M
Output
$0.26/M
Undated
anthropic/claude-sonnet-4.6

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...

T Weight access not listed
Context
1M
Input
$3/M
Output
$15/M
Undated
qwen/qwen3.5-plus-02-15

The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of...

T Weight access not listed
Context
1M
Input
$0.26/M
Output
$1.56/M
Undated
qwen/qwen3.5-397b-a17b

The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...

T Weight access not listed
Context
262.144K
Input
$0.39/M
Output
$2.34/M
Undated
anthropic/claude-opus-4.6

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...

T Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
Free Models Routerby Openrouter
Undated
openrouter/free

The simplest way to get free inference. openrouter/free is a router that selects free models at random from the models available on OpenRouter. The router smartly filters for models that...

T Weight access not listed
Context
200K
Input
Free
Output
Free
Undated
bytedance-seed/seed-1.6-flash

Seed 1.6 Flash is an ultra-fast multimodal deep thinking model by ByteDance Seed, supporting both text and visual understanding. It features a 256k context window and can generate outputs of...

T Weight access not listed
Context
262.144K
Input
$0.075/M
Output
$0.3/M
ByteDance Seed: Seed 1.6by Bytedance Seed
Undated
bytedance-seed/seed-1.6

Seed 1.6 is a general-purpose model released by the ByteDance Seed team. It incorporates multimodal capabilities and adaptive deep thinking with a 256K context window.

T Weight access not listed
Context
262.144K
Input
$0.25/M
Output
$2/M
Undated
openai/gpt-5.2-chat

GPT-5.2 Chat (AKA Instant) is the fast, lightweight member of the 5.2 family, optimized for low-latency chat while retaining strong general intelligence. It uses adaptive reasoning to selectively “think” on...

T Weight access not listed
Context
128K
Input
$1.75/M
Output
$14/M
Undated
z-ai/glm-4.6v

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts...

T Weight access not listed
Context
131.072K
Input
$0.3/M
Output
$0.9/M
Undated
amazon/nova-2-lite-v1

Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing...

T Weight access not listed
Context
1M
Input
$0.3/M
Output
$2.5/M
Undated
mistralai/ministral-14b-2512

The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language...

T Weight access not listed
Context
262.144K
Input
$0.2/M
Output
$0.2/M
Undated
mistralai/ministral-8b-2512

A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.

T Weight access not listed
Context
262.144K
Input
$0.15/M
Output
$0.15/M
Undated
mistralai/ministral-3b-2512

The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.

T Weight access not listed
Context
131.072K
Input
$0.1/M
Output
$0.1/M
Undated
mistralai/mistral-large-2512

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.

T Weight access not listed
Context
262.144K
Input
$0.5/M
Output
$1.5/M
Undated
anthropic/claude-opus-4.5

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...

T Weight access not listed
Context
200K
Input
$5/M
Output
$25/M
Undated
openai/gpt-5.1-chat

GPT-5.1 Chat (AKA Instant is the fast, lightweight member of the 5.1 family, optimized for low-latency chat while retaining strong general intelligence. It uses adaptive reasoning to selectively “think” on...

T Weight access not listed
Context
128K
Input
$1.25/M
Output
$10/M
Undated
amazon/nova-premier-v1

Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks and for use as the best teacher for distilling custom models.

T Weight access not listed
Context
1M
Input
$2.5/M
Output
$12.5/M
Undated
perplexity/sonar-pro-search

Exclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system. It is designed for deeper reasoning and analysis. Pricing is based...

T Weight access not listed
Context
200K
Input
$3/M
Output
$15/M
nvidia/nemotron-nano-12b-v2-vl:free

NVIDIA Nemotron Nano 2 VL is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligence. It introduces a hybrid Transformer-Mamba architecture, combining transformer-level accuracy with Mamba’s...

T Weight access not listed
Context
128K
Input
Free
Output
Free
Undated
qwen/qwen3-vl-32b-instruct

Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. With 32 billion parameters, it combines deep visual perception with advanced text...

T Weight access not listed
Context
131.072K
Input
$0.104/M
Output
$0.416/M
Undated
openai/gpt-5-image-mini

GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by [GPT-5 Mini](https://openrouter.ai/openai/gpt-5-mini), with GPT Image 1 Mini for efficient image generation. This natively multimodal model features superior instruction following, text...

T Weight access not listed
Context
400K
Input
$2.5/M
Output
$2/M
Undated
anthropic/claude-haiku-4.5

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

T Weight access not listed
Context
200K
Input
$1/M
Output
$5/M
Undated
qwen/qwen3-vl-8b-thinking

Qwen3-VL-8B-Thinking is the reasoning-optimized variant of the Qwen3-VL-8B multimodal model, designed for advanced visual and textual reasoning across complex scenes, documents, and temporal sequences. It integrates enhanced multimodal alignment and...

T Weight access not listed
Context
131.072K
Input
$0.18/M
Output
$2.1/M
Undated
qwen/qwen3-vl-8b-instruct

Qwen3-VL-8B-Instruct is a multimodal vision-language model from the Qwen3-VL series, built for high-fidelity understanding and reasoning across text, images, and video. It features improved multimodal fusion with Interleaved-MRoPE for long-horizon...

T Weight access not listed
Context
131.072K
Input
$0.117/M
Output
$0.455/M
Undated
openai/gpt-5-image

[GPT-5](https://openrouter.ai/openai/gpt-5) Image combines OpenAI's GPT-5 model with state-of-the-art image generation capabilities. It offers major improvements in reasoning, code quality, and user experience while incorporating GPT Image 1's superior instruction following,...

T Weight access not listed
Context
400K
Input
$10/M
Output
$10/M
Undated
qwen/qwen3-vl-30b-a3b-thinking

Qwen3-VL-30B-A3B-Thinking is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Thinking variant enhances reasoning in STEM, math, and complex tasks. It excels...

T Weight access not listed
Context
131.072K
Input
$0.2/M
Output
$2.4/M
Undated
qwen/qwen3-vl-30b-a3b-instruct

Qwen3-VL-30B-A3B-Instruct is a multimodal model that unifies strong text generation with visual understanding for images and videos. Its Instruct variant optimizes instruction-following for general multimodal tasks. It excels in perception...

T Weight access not listed
Context
262.144K
Input
$0.15/M
Output
$0.6/M