cohere/c4ai-aya-expanse-8b
Compact open multilingual model optimized for generation across 23 languages
- Context
- 8K
- Input
- -
- Output
- -
Filter 978 source-linked models by creator, price, context, modality, and published benchmark coverage.
cohere/c4ai-aya-expanse-8b
Compact open multilingual model optimized for generation across 23 languages
anthropic/claude-3-5-haiku-20241022
Fast Claude model for responsive assistance, classification, and lightweight agents
anthropic/claude-3-5-sonnet-20241022
Balanced Claude model for coding, analysis, agent workflows, and cost control
openai/whisper-large-v3-turbo
Speech transcription model for accurate audio-to-text and captioning workflows
openai/whisper-large-v3
Open Whisper checkpoint for robust multilingual transcription and captioning
mistral/pixtral-12b
Mistral vision-language model for image understanding and multimodal chat
alibaba/qwen2-5-vl-72b-instruct
Qwen vision-language model for visual reasoning, documents, and agent tasks
cohere/command-r-08-2024
Cohere retrieval model for long-context chat and enterprise RAG workflows
cohere/command-r-plus-08-2024
Cohere's RAG workhorse for long-context enterprise search and tool use
nvidia/nemotron-mini-4b-instruct
Compact Nemotron model for efficient reasoning and deployable AI agents
openai/gpt-4o-2024-08-06
GPT model for general reasoning, writing, coding, and tool-assisted tasks
openai/gpt-4o-mini
Small omni GPT for cheap multimodal assistance and production-scale traffic
mistral/mistral-nemo
Efficient Mistral-NVIDIA open model for multilingual chat and local deployment
openai/o3-deep-research
Research model for long-horizon investigation, synthesis, and analytical reports
openai/o4-mini-deep-research
Research model for long-horizon investigation, synthesis, and analytical reports
mistral/codestral-latest
Mistral code model for completions, refactors, and developer IDE workflows
openai/gpt-4o-2024-05-13
GPT model for general reasoning, writing, coding, and tool-assisted tasks
openai/gpt-4o
Omni-era GPT for multimodal chat, practical coding, and general assistants
alibaba/qwen-vl-max
Qwen vision-language model for visual reasoning, documents, and agent tasks
alibaba/qwen-max
Flagship Qwen model for complex reasoning, coding, and agentic workflows
anthropic/claude-3-haiku-20240307
Legacy model retained for compatibility with older integrations
alibaba/qwen-plus
Qwen instruction model for multilingual chat, reasoning, and tool use
alibaba/qwen-vl-plus
Qwen vision-language model for visual reasoning, documents, and agent tasks
perplexity/sonar-reasoning-pro
Web-grounded Sonar for multi-step research questions that need cited reasoning
perplexity/sonar-pro
Deeper Sonar search model with broader retrieval and stronger synthesis
perplexity/sonar
Fast web-grounded Sonar for current answers, citations, and lightweight retrieval
openai/gpt-4-turbo
Compact GPT model for low-latency assistance and high-volume workloads
openai/gpt-4
GPT model for general reasoning, writing, coding, and tool-assisted tasks
openai/gpt-3.5-turbo
Compact GPT model for low-latency assistance and high-volume workloads
openai/gpt-5.6-luna-pro
GPT-5.6 Luna Pro is the same underlying model as [GPT-5.6 Luna](https://openrouter.ai/openai/gpt-5.6-luna), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode
openai/gpt-5.6-terra-pro
GPT-5.6 Terra Pro is the same underlying model as [GPT-5.6 Terra](https://openrouter.ai/openai/gpt-5.6-terra), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode
openai/gpt-5.6-sol-pro
GPT-5.6 Sol Pro is the same underlying model as [GPT-5.6 Sol](https://openrouter.ai/openai/gpt-5.6-sol), served with `reasoning.mode` set to `pro` for higher-quality responses on complex tasks. Learn more in OpenAI's docs: https://developers.openai.com/api/docs/guides/reasoning#reasoning-mode
x-ai/grok-4.5
Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.
~x-ai/grok-latest
This model always redirects to the latest Grok model from xAI.
aion-labs/aion-3.0-mini
Aion-3.0 Mini is a multi-model roleplaying and storytelling system from AionLabs, built on the DeepSeek family of models. It uses a collaborative generation process in which multiple specialized models each...
aion-labs/aion-3.0
Aion-3.0 is a multi-model roleplaying and storytelling system from AionLabs, built on the GLM family of models. It uses a collaborative generation process in which multiple specialized models each contribute...
tencent/hy3:free
Hy3 is a 295B-parameter Mixture-of-Experts model from Tencent (21B active, 192 experts with top-8 routing) built for reasoning, agentic workflows, and real-world production use. It supports a configurable reasoning effort:...
poolside/laguna-xs-2.1:free
Laguna XS 2.1 is the latest coding agent model in the 33B-A3B category from [Poolside](https://poolside.ai/) and a step forward from their Laguna XS.2 model (released in April 2026). It combines...
nex-agi/nex-n2-mini
Nex-N2-Mini is an open-source agentic mixture-of-experts model from Nex AGI, the smaller sibling in the Nex-N2 series. It accepts text and image input and is built for coding, tool use,...
cohere/north-mini-code:free
North Mini Code is Cohere's first agentic coding model and the debut of its North family. A sparse mixture-of-experts model with 30B total parameters and 3B active, it is optimized...
z-ai/glm-5.2
GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...
openrouter/fusion
Fusion turns your prompt into a small multi-model deliberation. A panel of expert models (see below) analyzes your prompt in parallel with web search and web fetch enabled, then a...
~anthropic/claude-fable-latest
This model always redirects to the latest model in the Claude Fable family.
nex-agi/nex-n2-pro
Nex-N2-Pro is an agentic mixture-of-experts model from Nex AGI, with 17B active parameters out of 397B total. Built on the Qwen3.5 architecture, it accepts text and image input and produces...
nvidia/nemotron-3.5-content-safety:free
NVIDIA Nemotron 3.5 Content Safety is a compact 4B-parameter multimodal guardrail model from NVIDIA, fine-tuned from Google Gemma-3-4B. It moderates both inputs to and responses from LLMs and VLMs, accepting...
nvidia/nemotron-3-ultra-550b-a55b:free
NVIDIA Nemotron 3 Ultra is an open frontier-reasoning and orchestration model from NVIDIA, with 55B active parameters out of 550B total (MoE). Built on a hybrid Transformer-Mamba mixture-of-experts architecture, it...
qwen/qwen3.7-plus
Qwen3.7-Plus is a cost-effective model in Alibaba's Qwen3.7 series. It supports text and image input with text output, building on the series' text capabilities with a comprehensive upgrade to its...
minimax/minimax-m3
MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...
anthropic/claude-opus-4.8-fast
Fast-mode variant of [Opus 4.8](/anthropic/claude-opus-4.8) - identical capabilities with higher output speed at 2x pricing relative to regular Opus 4.8. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode
anthropic/claude-opus-4.8
Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...
| Model | Creator | Input types | Context | Input / Output | Released | Compare |
|---|---|---|---|---|---|---|
| Aya Expanse 8Bcohere/c4ai-aya-expanse-8b | 8K | - / - | 2024-10-24 | |||
| Claude Haiku 3.5anthropic/claude-3-5-haiku-20241022 | 200K | $0.8 / $4 | 2024-10-22 | |||
| Claude Sonnet 3.5 v2anthropic/claude-3-5-sonnet-20241022 | 200K | $3 / $15 | 2024-10-22 | |||
| Whisper Large v3 Turboopenai/whisper-large-v3-turbo | 448 | $0.002 / $0.002 | 2024-10-01 | |||
| Whisper 3 Largeopenai/whisper-large-v3 | 448 | $0.002 / $0.002 | 2024-10-01 | |||
| Pixtral 12Bmistral/pixtral-12b | 128K | $0.15 / $0.15 | 2024-09-01 | |||
| Qwen2.5-VL 72B Instructalibaba/qwen2-5-vl-72b-instruct | 131.072K | $2.8 / $8.4 | 2024-09 | |||
| Command Rcohere/command-r-08-2024 | 128K | $0.15 / $0.6 | 2024-08-30 | |||
| Command R+cohere/command-r-plus-08-2024 | 128K | $2.5 / $10 | 2024-08-30 | |||
| Nemotron Mini 4B Instructnvidia/nemotron-mini-4b-instruct | 128K | - / - | 2024-08-21 | |||
| GPT-4o (2024-08-06)openai/gpt-4o-2024-08-06 | 128K | $2.5 / $10 | 2024-08-06 | |||
| GPT-4o miniopenai/gpt-4o-mini | 128K | $0.15 / $0.6 | 2024-07-18 | |||
| Mistral Nemomistral/mistral-nemo | 128K | $0.15 / $0.15 | 2024-07-01 | |||
| o3-deep-researchopenai/o3-deep-research | 200K | $9 / $36 | 2024-06-26 | |||
| o4-mini-deep-researchopenai/o4-mini-deep-research | 200K | $1.8 / $7.2 | 2024-06-26 | |||
| Codestral (latest)mistral/codestral-latest | 256K | $0.3 / $0.9 | 2024-05-29 | |||
| GPT-4o (2024-05-13)openai/gpt-4o-2024-05-13 | 128K | $5 / $15 | 2024-05-13 | |||
| GPT-4oopenai/gpt-4o | 128K | $2.5 / $10 | 2024-05-13 | |||
| Qwen-VL Maxalibaba/qwen-vl-max | 131.072K | $0.8 / $3.2 | 2024-04-08 | |||
| Qwen Maxalibaba/qwen-max | 32.768K | $1.6 / $6.4 | 2024-04-03 | |||
| Claude Haiku 3anthropic/claude-3-haiku-20240307 | 200K | $0.25 / $1.25 | 2024-03-13 | |||
| Qwen Plusalibaba/qwen-plus | 1M | $0.4 / $1.2 | 2024-01-25 | |||
| Qwen-VL Plusalibaba/qwen-vl-plus | 131.072K | $0.21 / $0.63 | 2024-01-25 | |||
| Sonar Reasoning Properplexity/sonar-reasoning-pro | 128K | $2 / $8 | 2024-01-01 | |||
| Sonar Properplexity/sonar-pro | 200K | $3 / $15 | 2024-01-01 | |||
| Sonarperplexity/sonar | 128K | $1 / $1 | 2024-01-01 | |||
| GPT-4 Turboopenai/gpt-4-turbo | 128K | $10 / $30 | 2023-11-06 | |||
| GPT-4openai/gpt-4 | 8.192K | $30 / $60 | 2023-11-06 | |||
| GPT-3.5-turboopenai/gpt-3.5-turbo | 16.385K | $0.5 / $1.5 | 2023-03-01 | |||
| OpenAI: GPT-5.6 Luna Proopenai/gpt-5.6-luna-pro | 1.05M | $0.1 / $0.6 | Undated | |||
| OpenAI: GPT-5.6 Terra Proopenai/gpt-5.6-terra-pro | 1.05M | $1 / $6 | Undated | |||
| OpenAI: GPT-5.6 Sol Proopenai/gpt-5.6-sol-pro | 1.05M | $5 / $30 | Undated | |||
| xAI: Grok 4.5x-ai/grok-4.5 | 500K | $2 / $6 | Undated | |||
| xAI: Grok Latest~x-ai/grok-latest | 500K | $2 / $6 | Undated | |||
| AionLabs: Aion-3.0-Miniaion-labs/aion-3.0-mini | 131.072K | $0.7 / $1.4 | Undated | |||
| AionLabs: Aion-3.0aion-labs/aion-3.0 | 131.072K | $3 / $6 | Undated | |||
| Tencent: Hy3 (free)tencent/hy3:free | 262.144K | Free / Free | Undated | |||
| Poolside: Laguna XS 2.1 (free)poolside/laguna-xs-2.1:free | 262.144K | Free / Free | Undated | |||
| Nex AGI: Nex-N2-Mininex-agi/nex-n2-mini | 262.144K | $0.025 / $0.1 | Undated | |||
| Cohere: North Mini Code (free)cohere/north-mini-code:free | 256K | Free / Free | Undated | |||
| Z.ai: GLM 5.2z-ai/glm-5.2 | 1.024M | $1.19 / $3.74 | Undated | |||
| OpenRouter: Fusionopenrouter/fusion | 1M | - / - | Undated | |||
| Anthropic: Claude Fable Latest~anthropic/claude-fable-latest | 1M | $10 / $50 | Undated | |||
| Nex AGI: Nex-N2-Pronex-agi/nex-n2-pro | 262.144K | $0.25 / $1 | Undated | |||
| NVIDIA: Nemotron 3.5 Content Safety (free)nvidia/nemotron-3.5-content-safety:free | 128K | Free / Free | Undated | |||
| NVIDIA: Nemotron 3 Ultra (free)nvidia/nemotron-3-ultra-550b-a55b:free | 1M | Free / Free | Undated | |||
| Qwen: Qwen3.7 Plusqwen/qwen3.7-plus | 1M | $0.32 / $1.28 | Undated | |||
| MiniMax: MiniMax M3minimax/minimax-m3 | 524.288K | $0.3 / $1.2 | Undated | |||
| Anthropic: Claude Opus 4.8 (Fast)anthropic/claude-opus-4.8-fast | 1M | $10 / $50 | Undated | |||
| Anthropic: Claude Opus 4.8anthropic/claude-opus-4.8 | 1M | $5 / $25 | Undated |