978 results
Undated
qwen/qwen3.5-122b-a10b

The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of...

T Weight access not listed
Context
262.144K
Input
$0.26/M
Output
$2.08/M
Qwen: Qwen3.5-Flashby Alibaba Qwen
Undated
qwen/qwen3.5-flash-02-23

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the...

T Weight access not listed
Context
1M
Input
$0.065/M
Output
$0.26/M
AionLabs: Aion-2.0by Aion Labs
Undated
aion-labs/aion-2.0

Aion-2.0 is a variant of DeepSeek V3.2 optimized for immersive roleplaying and storytelling. It is particularly strong at introducing tension, crises, and conflict into stories, making narratives feel more engaging....

T Weight access not listed
Context
131.072K
Input
$0.8/M
Output
$1.6/M
Undated
anthropic/claude-sonnet-4.6

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...

T Weight access not listed
Context
1M
Input
$3/M
Output
$15/M
Undated
qwen/qwen3.5-plus-02-15

The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of...

T Weight access not listed
Context
1M
Input
$0.26/M
Output
$1.56/M
Undated
qwen/qwen3.5-397b-a17b

The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...

T Weight access not listed
Context
262.144K
Input
$0.39/M
Output
$2.34/M
Undated
minimax/minimax-m2.5

MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1...

T Weight access not listed
Context
196.608K
Input
$0.15/M
Output
$0.9/M
Undated
z-ai/glm-5

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading...

T Weight access not listed
Context
204.8K
Input
$0.95/M
Output
$2.55/M
Undated
qwen/qwen3-max-thinking

Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step reasoning. By significantly scaling model capacity and reinforcement learning compute, it...

T Weight access not listed
Context
262.144K
Input
$0.78/M
Output
$3.9/M
Undated
anthropic/claude-opus-4.6

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...

T Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
Qwen: Qwen3 Coder Nextby Alibaba Qwen
Undated
qwen/qwen3-coder-next

Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse MoE design with 80B total parameters and only 3B activated per...

T Weight access not listed
Context
262.144K
Input
$0.12/M
Output
$0.8/M
Free Models Routerby Openrouter
Undated
openrouter/free

The simplest way to get free inference. openrouter/free is a router that selects free models at random from the models available on OpenRouter. The router smartly filters for models that...

T Weight access not listed
Context
200K
Input
Free
Output
Free
Undated
upstage/solar-pro-3

Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass, it delivers exceptional performance while maintaining computational efficiency. Optimized...

T Weight access not listed
Context
128K
Input
$0.15/M
Output
$0.6/M
Undated
minimax/minimax-m2-her

MiniMax M2-her is a dialogue-first large language model built for immersive roleplay, character-driven chat, and expressive multi-turn conversations. Designed to stay consistent in tone and personality, it supports rich message...

T Weight access not listed
Context
65.536K
Input
$0.3/M
Output
$1.2/M
Undated
writer/palmyra-x5

Palmyra X5 is Writer's most advanced model, purpose-built for building and scaling AI agents across the enterprise. It delivers industry-leading speed and efficiency on context windows up to 1 million...

T Weight access not listed
Context
1.04M
Input
$0.6/M
Output
$6/M
liquid/lfm-2.5-1.2b-thinking:free

LFM2.5-1.2B-Thinking is a lightweight reasoning-focused model optimized for agentic tasks, data extraction, and RAG—while still running comfortably on edge devices. It supports long context (up to 32K tokens) and is...

T Weight access not listed
Context
32.768K
Input
Free
Output
Free
liquid/lfm-2.5-1.2b-instruct:free

LFM2.5-1.2B-Instruct is a compact, high-performance instruction-tuned model built for fast on-device AI. It delivers strong chat quality in a 1.2B parameter footprint, with efficient edge inference and broad runtime support.

T Weight access not listed
Context
32.768K
Input
Free
Output
Free
Undated
openai/gpt-audio

The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Audio is priced...

T Weight access not listed
Context
128K
Input
$2.5/M
Output
$10/M
Undated
openai/gpt-audio-mini

A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Input is priced at $0.60 per million...

T Weight access not listed
Context
128K
Input
$0.6/M
Output
$2.4/M
Undated
z-ai/glm-4.7-flash

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...

T Weight access not listed
Context
202.752K
Input
$0.06/M
Output
$0.4/M
Undated
bytedance-seed/seed-1.6-flash

Seed 1.6 Flash is an ultra-fast multimodal deep thinking model by ByteDance Seed, supporting both text and visual understanding. It features a 256k context window and can generate outputs of...

T Weight access not listed
Context
262.144K
Input
$0.075/M
Output
$0.3/M
ByteDance Seed: Seed 1.6by Bytedance Seed
Undated
bytedance-seed/seed-1.6

Seed 1.6 is a general-purpose model released by the ByteDance Seed team. It incorporates multimodal capabilities and adaptive deep thinking with a 256K context window.

T Weight access not listed
Context
262.144K
Input
$0.25/M
Output
$2/M
Undated
minimax/minimax-m2.1

MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world...

T Weight access not listed
Context
204.8K
Input
$0.3/M
Output
$1.2/M
Undated
z-ai/glm-4.7

GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...

T Weight access not listed
Context
202.752K
Input
$0.4/M
Output
$1.75/M
nvidia/nemotron-3-nano-30b-a3b:free

NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...

T Weight access not listed
Context
256K
Input
Free
Output
Free
Undated
openai/gpt-5.2-chat

GPT-5.2 Chat (AKA Instant) is the fast, lightweight member of the 5.2 family, optimized for low-latency chat while retaining strong general intelligence. It uses adaptive reasoning to selectively “think” on...

T Weight access not listed
Context
128K
Input
$1.75/M
Output
$14/M
Undated
mistralai/devstral-2512

Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer model supporting a 256K context window. Devstral 2 supports exploring...

T Weight access not listed
Context
262.144K
Input
$0.4/M
Output
$2/M
Undated
relace/relace-search

The relace-search model uses 4-12 `view_file` and `grep` tools in parallel to explore a codebase and return relevant files to the user request. In contrast to RAG, relace-search performs agentic...

T Weight access not listed
Context
256K
Input
$1/M
Output
$3/M
Undated
z-ai/glm-4.6v

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts...

T Weight access not listed
Context
131.072K
Input
$0.3/M
Output
$0.9/M
Body Builder (beta)by Openrouter
Undated
openrouter/bodybuilder

Transform your natural language requests into structured OpenRouter API request objects. Describe what you want to accomplish with AI models, and Body Builder will construct the appropriate API calls. Example:...

T Weight access not listed
Context
128K
Input
-
Output
-
Undated
amazon/nova-2-lite-v1

Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing...

T Weight access not listed
Context
1M
Input
$0.3/M
Output
$2.5/M
Undated
mistralai/ministral-14b-2512

The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language...

T Weight access not listed
Context
262.144K
Input
$0.2/M
Output
$0.2/M
Undated
mistralai/ministral-8b-2512

A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.

T Weight access not listed
Context
262.144K
Input
$0.15/M
Output
$0.15/M
Undated
mistralai/ministral-3b-2512

The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.

T Weight access not listed
Context
131.072K
Input
$0.1/M
Output
$0.1/M
Undated
mistralai/mistral-large-2512

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.

T Weight access not listed
Context
262.144K
Input
$0.5/M
Output
$1.5/M
Undated
deepseek/deepseek-v3.2

DeepSeek-V3.2 is a large language model designed to harmonize high computational efficiency with strong reasoning and agentic tool-use performance. It introduces DeepSeek Sparse Attention (DSA), a fine-grained sparse attention mechanism...

T Weight access not listed
Context
163.84K
Input
$0.269/M
Output
$0.4/M
Undated
anthropic/claude-opus-4.5

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...

T Weight access not listed
Context
200K
Input
$5/M
Output
$25/M
Undated
allenai/olmo-3-32b-think

Olmo 3 32B Think is a large-scale, 32-billion-parameter model purpose-built for deep reasoning, complex logic chains and advanced instruction-following scenarios. Its capacity enables strong performance on demanding evaluation tasks and...

T Weight access not listed
Context
65.536K
Input
$0.15/M
Output
$0.5/M
Undated
deepcogito/cogito-v2.1-671b

Cogito v2.1 671B MoE represents one of the strongest open models globally, matching performance of frontier closed and open models. This model is trained using self play with reinforcement learning...

T Weight access not listed
Context
128K
Input
$1.25/M
Output
$1.25/M
Undated
openai/gpt-5.1-chat

GPT-5.1 Chat (AKA Instant is the fast, lightweight member of the 5.1 family, optimized for low-latency chat while retaining strong general intelligence. It uses adaptive reasoning to selectively “think” on...

T Weight access not listed
Context
128K
Input
$1.25/M
Output
$10/M
Undated
amazon/nova-premier-v1

Amazon Nova Premier is the most capable of Amazon’s multimodal models for complex reasoning tasks and for use as the best teacher for distilling custom models.

T Weight access not listed
Context
1M
Input
$2.5/M
Output
$12.5/M
Undated
perplexity/sonar-pro-search

Exclusively available on the OpenRouter API, Sonar Pro's new Pro Search mode is Perplexity's most advanced agentic search system. It is designed for deeper reasoning and analysis. Pricing is based...

T Weight access not listed
Context
200K
Input
$3/M
Output
$15/M
Undated
mistralai/voxtral-small-24b-2507

Voxtral Small is an enhancement of Mistral Small 3, incorporating state-of-the-art audio input capabilities while retaining best-in-class text performance. It excels at speech transcription, translation and audio understanding. Input audio...

T Weight access not listed
Context
32K
Input
$0.1/M
Output
$0.3/M
openai/gpt-oss-safeguard-20b

gpt-oss-safeguard-20b is a safety reasoning model from OpenAI built upon gpt-oss-20b. This open-weight, 21B-parameter Mixture-of-Experts (MoE) model offers lower latency for safety tasks like content classification, LLM filtering, and trust...

T Weight access not listed
Context
131.072K
Input
$0.075/M
Output
$0.3/M
nvidia/nemotron-nano-12b-v2-vl:free

NVIDIA Nemotron Nano 2 VL is a 12-billion-parameter open multimodal reasoning model designed for video understanding and document intelligence. It introduces a hybrid Transformer-Mamba architecture, combining transformer-level accuracy with Mamba’s...

T Weight access not listed
Context
128K
Input
Free
Output
Free
Undated
minimax/minimax-m2

MiniMax-M2 is a compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters (230 billion total), it delivers near-frontier intelligence across general reasoning,...

T Weight access not listed
Context
204.8K
Input
$0.255/M
Output
$1.02/M
Undated
qwen/qwen3-vl-32b-instruct

Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. With 32 billion parameters, it combines deep visual perception with advanced text...

T Weight access not listed
Context
131.072K
Input
$0.104/M
Output
$0.416/M
Undated
ibm-granite/granite-4.0-h-micro

Granite-4.0-H-Micro is a 3B parameter from the Granite 4 family of models. These models are the latest in a series of models released by IBM. They are fine-tuned for long...

T Weight access not listed
Context
131K
Input
$0.017/M
Output
$0.112/M
Undated
openai/gpt-5-image-mini

GPT-5 Image Mini combines OpenAI's advanced language capabilities, powered by [GPT-5 Mini](https://openrouter.ai/openai/gpt-5-mini), with GPT Image 1 Mini for efficient image generation. This natively multimodal model features superior instruction following, text...

T Weight access not listed
Context
400K
Input
$2.5/M
Output
$2/M
Undated
anthropic/claude-haiku-4.5

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

T Weight access not listed
Context
200K
Input
$1/M
Output
$5/M