978 results
2024-10-24
cohere/c4ai-aya-expanse-8b

Compact open multilingual model optimized for generation across 23 languages

T Open weights
Context
8K
Input
-
Output
-
Claude Haiku 3.5by Anthropic
2024-10-22
anthropic/claude-3-5-haiku-20241022

Fast Claude model for responsive assistance, classification, and lightweight agents

T Tools Weight access not listed
Context
200K
Input
$0.8/M
Output
$4/M
2024-10-22
anthropic/claude-3-5-sonnet-20241022

Balanced Claude model for coding, analysis, agent workflows, and cost control

T Tools Weight access not listed
Context
200K
Input
$3/M
Output
$15/M
2024-10-01
openai/whisper-large-v3-turbo

Speech transcription model for accurate audio-to-text and captioning workflows

Open weights
Context
448
Input
$0.002/M
Output
$0.002/M
2024-10-01
openai/whisper-large-v3

Open Whisper checkpoint for robust multilingual transcription and captioning

Open weights
Context
448
Input
$0.002/M
Output
$0.002/M
Pixtral 12Bby Mistral AI
2024-09-01
mistral/pixtral-12b

Mistral vision-language model for image understanding and multimodal chat

T Tools Open weights
Context
128K
Input
$0.15/M
Output
$0.15/M
2024-09
alibaba/qwen2-5-vl-72b-instruct

Qwen vision-language model for visual reasoning, documents, and agent tasks

T Tools Open weights
Context
131.072K
Input
$2.8/M
Output
$8.4/M
Command Rby Cohere
2024-08-30
cohere/command-r-08-2024

Cohere retrieval model for long-context chat and enterprise RAG workflows

T Tools Open weights
Context
128K
Input
$0.15/M
Output
$0.6/M
Command R+by Cohere
2024-08-30
cohere/command-r-plus-08-2024

Cohere's RAG workhorse for long-context enterprise search and tool use

T Tools Open weights
Context
128K
Input
$2.5/M
Output
$10/M
2024-08-21
nvidia/nemotron-mini-4b-instruct

Compact Nemotron model for efficient reasoning and deployable AI agents

T Tools Open weights
Context
128K
Input
-
Output
-
2024-08-06
openai/gpt-4o-2024-08-06

GPT model for general reasoning, writing, coding, and tool-assisted tasks

T Weight access not listed
Context
128K
Input
$2.5/M
Output
$10/M
GPT-4o miniby OpenAI
2024-07-18
openai/gpt-4o-mini

Small omni GPT for cheap multimodal assistance and production-scale traffic

T Tools Structured output Weight access not listed
Context
128K
Input
$0.15/M
Output
$0.6/M
Mistral Nemoby Mistral AI
2024-07-01
mistral/mistral-nemo

Efficient Mistral-NVIDIA open model for multilingual chat and local deployment

T Open weights
Context
128K
Input
$0.15/M
Output
$0.15/M
2024-06-26
openai/o3-deep-research

Research model for long-horizon investigation, synthesis, and analytical reports

T Reasoning Tools Weight access not listed
Context
200K
Input
$9/M
Output
$36/M
2024-06-26
openai/o4-mini-deep-research

Research model for long-horizon investigation, synthesis, and analytical reports

T Reasoning Tools Weight access not listed
Context
200K
Input
$1.8/M
Output
$7.2/M
Codestral (latest)by Mistral AI
2024-05-29
mistral/codestral-latest

Mistral code model for completions, refactors, and developer IDE workflows

T Tools Open weights
Context
256K
Input
$0.3/M
Output
$0.9/M
2024-05-13
openai/gpt-4o-2024-05-13

GPT model for general reasoning, writing, coding, and tool-assisted tasks

T Tools Structured output Weight access not listed
Context
128K
Input
$5/M
Output
$15/M
GPT-4oby OpenAI
2024-05-13
openai/gpt-4o

Omni-era GPT for multimodal chat, practical coding, and general assistants

T Tools Weight access not listed
Context
128K
Input
$2.5/M
Output
$10/M
Qwen-VL Maxby Alibaba Qwen
2024-04-08
alibaba/qwen-vl-max

Qwen vision-language model for visual reasoning, documents, and agent tasks

T Tools Weight access not listed
Context
131.072K
Input
$0.8/M
Output
$3.2/M
Qwen Maxby Alibaba Qwen
2024-04-03
alibaba/qwen-max

Flagship Qwen model for complex reasoning, coding, and agentic workflows

T Tools Weight access not listed
Context
32.768K
Input
$1.6/M
Output
$6.4/M
Claude Haiku 3by Anthropic
2024-03-13
anthropic/claude-3-haiku-20240307

Legacy model retained for compatibility with older integrations

T Tools Weight access not listed
Context
200K
Input
$0.25/M
Output
$1.25/M
Qwen Plusby Alibaba Qwen
2024-01-25
alibaba/qwen-plus

Qwen instruction model for multilingual chat, reasoning, and tool use

T Reasoning Tools Weight access not listed
Context
1M
Input
$0.4/M
Output
$1.2/M
Qwen-VL Plusby Alibaba Qwen
2024-01-25
alibaba/qwen-vl-plus

Qwen vision-language model for visual reasoning, documents, and agent tasks

T Tools Weight access not listed
Context
131.072K
Input
$0.21/M
Output
$0.63/M
Sonar Reasoning Proby Perplexity
2024-01-01
perplexity/sonar-reasoning-pro

Web-grounded Sonar for multi-step research questions that need cited reasoning

T Reasoning Weight access not listed
Context
128K
Input
$2/M
Output
$8/M
Sonar Proby Perplexity
2024-01-01
perplexity/sonar-pro

Deeper Sonar search model with broader retrieval and stronger synthesis

T Weight access not listed
Context
200K
Input
$3/M
Output
$15/M
Sonarby Perplexity
2024-01-01
perplexity/sonar

Fast web-grounded Sonar for current answers, citations, and lightweight retrieval

T Weight access not listed
Context
128K
Input
$1/M
Output
$1/M
GPT-4 Turboby OpenAI
2023-11-06
openai/gpt-4-turbo

Compact GPT model for low-latency assistance and high-volume workloads

T Tools Weight access not listed
Context
128K
Input
$10/M
Output
$30/M
GPT-4by OpenAI
2023-11-06
openai/gpt-4

GPT model for general reasoning, writing, coding, and tool-assisted tasks

T Tools Weight access not listed
Context
8.192K
Input
$30/M
Output
$60/M
GPT-3.5-turboby OpenAI
2023-03-01
openai/gpt-3.5-turbo

Compact GPT model for low-latency assistance and high-volume workloads

T Tools Weight access not listed
Context
16.385K
Input
$0.5/M
Output
$1.5/M
Undated
openai/gpt-5.6-luna-pro

GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

T Weight access not listed
Context
1.05M
Input
$0.1/M
Output
$0.6/M
Undated
openai/gpt-5.6-terra-pro

GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

T Weight access not listed
Context
1.05M
Input
$1/M
Output
$6/M
Undated
openai/gpt-5.6-sol-pro

GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode

T Weight access not listed
Context
1.05M
Input
$5/M
Output
$30/M
Undated
x-ai/grok-4.5

Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

T Weight access not listed
Context
500K
Input
$2/M
Output
$6/M
Undated
~x-ai/grok-latest

This model always redirects to the latest Grok model from xAI.

T Weight access not listed
Context
500K
Input
$2/M
Output
$6/M
Undated
aion-labs/aion-3.0-mini

Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. It uses a collaborative generation process in which multiple specialized models each...

T Weight access not listed
Context
131.072K
Input
$0.7/M
Output
$1.4/M
AionLabs: Aion-3.0by Aion Labs
Undated
aion-labs/aion-3.0

Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation process in which multiple specialized models each contribute...

T Weight access not listed
Context
131.072K
Input
$3/M
Output
$6/M
Undated
tencent/hy3:free

Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...

T Weight access not listed
Context
262.144K
Input
Free
Output
Free
Undated
poolside/laguna-xs-2.1:free

Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...

T Weight access not listed
Context
262.144K
Input
Free
Output
Free
Undated
nex-agi/nex-n2-mini

Nex-N2-Mini is an open-source agentic mixture-of-experts model from Nex AGI, the smaller sibling in the Nex-N2 series. It accepts text and image input and is built for coding, tool use,...

T Weight access not listed
Context
262.144K
Input
$0.025/M
Output
$0.1/M
cohere/north-mini-code:free

North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized...

T Weight access not listed
Context
256K
Input
Free
Output
Free
Undated
z-ai/glm-5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

T Weight access not listed
Context
1.024M
Input
$1.19/M
Output
$3.74/M
OpenRouter: Fusionby Openrouter
Undated
openrouter/fusion

Fusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel with web search and web fetch enabled, then a...

T Weight access not listed
Context
1M
Input
-
Output
-
Undated
~anthropic/claude-fable-latest

This model always redirects to the latest model in the Claude Fable family.

T Weight access not listed
Context
1M
Input
$10/M
Output
$50/M
Undated
nex-agi/nex-n2-pro

Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total. Built on the Qwen3.5 architecture, it accepts text and image input and produces...

T Weight access not listed
Context
262.144K
Input
$0.25/M
Output
$1/M
nvidia/nemotron-3.5-content-safety:free

NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...

T Weight access not listed
Context
128K
Input
Free
Output
Free
nvidia/nemotron-3-ultra-550b-a55b:free

NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...

T Weight access not listed
Context
1M
Input
Free
Output
Free
Qwen: Qwen3.7 Plusby Alibaba Qwen
Undated
qwen/qwen3.7-plus

Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...

T Weight access not listed
Context
1M
Input
$0.32/M
Output
$1.28/M
Undated
minimax/minimax-m3

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

T Weight access not listed
Context
524.288K
Input
$0.3/M
Output
$1.2/M
Undated
anthropic/claude-opus-4.8-fast

Fast-mode variant of [Opus 4.8](/anthropic/claude-opus-4.8) - identical capabilities with higher output speed at 2x pricing relative to regular Opus 4.8. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode

T Weight access not listed
Context
1M
Input
$10/M
Output
$50/M
Undated
anthropic/claude-opus-4.8

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

T Weight access not listed
Context
1M
Input
$5/M
Output
$25/M