mistral/mistral-medium-2604
Balanced Mistral model for enterprise assistants, multilingual work, and tools
- Context
- 262.144K
- Input
- $1.5/M
- Output
- $7.5/M
Filter 978 source-linked models by creator, price, context, modality, and published benchmark coverage.
mistral/mistral-medium-2604
Balanced Mistral model for enterprise assistants, multilingual work, and tools
poolside/laguna-m.1
Poolside's open-weight model for agentic coding and long-horizon work
poolside/laguna-xs.2
Agentic coding model from Poolside in the XS size class for local deployment
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning
Open Nemotron omni model combining reasoning with text, vision, and audio
alibaba/qwen3.6-flash
Qwen vision-language model for visual reasoning, documents, and agent tasks
deepseek/deepseek-v4-pro
Open MoE flagship with million-token context for coding and long agent runs
deepseek/deepseek-v4-flash
Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work
openai/gpt-5.5-pro
Highest-accuracy GPT-5.5 tier for slower, precision-heavy reasoning and coding
openai/gpt-5.5
Default frontier GPT for coding, computer use, research, and knowledge work
xiaomi/mimo-v2.5-pro
Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution
xiaomi/mimo-v2.5
Open MiMo model for multimodal coding agents and long-context automation
alibaba/qwen3.6-27b
Qwen vision-language model for visual reasoning, documents, and agent tasks
google/gemini-embedding-2
Multimodal embedding model mapping text, images, video, audio, and PDFs into a unified embedding space
moonshotai/kimi-k2.6
Multimodal Kimi workhorse for agent loops, coding tasks, and visual context
openai/gpt-image-2
Image model for prompt-driven generation, editing, and visual design workflows
google/deep-research-preview-04-2026
Agentic model for autonomous multi-step research, synthesis, and cited reports
google/deep-research-max-preview-04-2026
Maximum-comprehensiveness agentic researcher for multi-step investigation, synthesis, and cited reports
tencent/hy3-preview
Tencent Hy reasoning model for coding, instruction following, and agent tasks
alibaba/qwen3.6-max-preview
Flagship Qwen model for complex reasoning, coding, and agentic workflows
xai/grok-4.3
xAI's default Grok for chat, coding, agentic tools, and lower hallucination risk
alibaba/qwen3.6-35b-a3b
Open multimodal Qwen MoE for local agents that need vision, audio, and code
xai/grok-build-0.1
Fast Grok coding model tuned for agentic engineering and iterative edits
nvidia/nemotron-3-content-safety
Safety model for policy screening, moderation, and risk-aware routing workflows
anthropic/claude-opus-4-7
Stronger Opus tier for advanced software work and high-stakes reasoning
google/gemini-3.1-flash-tts-preview
Low-latency speech generation with steerable prompts and expressive audio tags
google/gemini-robotics-er-1.6-preview
Vision-language model for embodied reasoning: spatial understanding, task planning, and physical-world agentic robotics
meta/muse-spark-1.1
Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.
zhipuai/glm-5.1
Strong GLM coding model for agentic engineering, terminals, and repository generation
google/gemma-4-26b-a4b-it
Open Gemma instruction model for efficient chat and self-hosted deployments
google/gemma-4-31b-it
Largest Gemma 4 instruction model for open, self-hosted chat and reasoning
google/gemma-4-E4B-it
Open Gemma instruction model for efficient chat and self-hosted deployments
google/gemma-4-E2B-it
Open Gemma instruction model for efficient chat and self-hosted deployments
stepfun/step-3.5-flash-2603
StepFun flash model for efficient multimodal reasoning, coding, and tool use
alibaba/qwen3.6-plus
Earlier Qwen multimodal workhorse for million-token agent and document tasks
zhipuai/glm-5v-turbo
Fast GLM vision model for screenshots, documents, and multimodal agent tasks
nvidia/llama-nemotron-rerank-vl-1b-v2
Reranking model for improving retrieval quality in search and recommendation systems
google/veo-3.1-lite-generate-preview
Video model for prompt-guided generation, editing, and motion workflows
google/gemini-3.1-flash-live-preview
High-quality, low-latency Live API model for real-time dialogue and voice-first AI applications
google/lyria-3-pro-preview
Music generation model for full-length songs from text or images with vocals and structure
google/lyria-3-clip-preview
Music generation model for short 30-second clips, loops, and previews from text or image prompts
nvidia/nemotron-cascade-2-30b-a3b
Nemotron model for efficient reasoning, coding, and specialized AI agents
xiaomi/mimo-v2-omni
MiMo omni model for text, image, video, audio, and agents
xiaomi/mimo-v2-pro
Earlier MiMo Pro model for multimodal agents, reasoning, and code tasks
minimax/MiniMax-M2.7-highspeed
Low-latency M2.7 variant for interactive coding plans and agent loops
minimax/MiniMax-M2.7
Open MiniMax flagship for coding agents, office automation, and complex environments
openai/gpt-5.4-nano
Cheapest GPT-5.4 lane for simple routing, extraction, and bulk automation
openai/gpt-5.4-mini
Strong small GPT for coding subagents, quick tool use, and high-volume work
zhipuai/glm-5-turbo
Faster GLM-5 lane for coding agents that need lower latency
mistral/mistral-small-latest
Efficient Mistral model for fast chat, extraction, and production assistants
mistral/mistral-small-2603
Fast Mistral production model for chat, extraction, and cost-sensitive agents
| Model | Creator | Input types | Context | Input / Output | Released | Compare |
|---|---|---|---|---|---|---|
| Mistral Medium 3.5mistral/mistral-medium-2604 | 262.144K | $1.5 / $7.5 | 2026-04-29 | |||
| Laguna M.1poolside/laguna-m.1 | 262.144K | $0.2 / $0.4 | 2026-04-28 | |||
| Laguna XS.2poolside/laguna-xs.2 | 262.144K | $0.2 / $0.4 | 2026-04-28 | |||
| Nemotron 3 Nano Omni 30B A3B Reasoningnvidia/nemotron-3-nano-omni-30b-a3b-reasoning | 256K | $0.105 / $0.42 | 2026-04-28 | |||
| Qwen3.6 Flashalibaba/qwen3.6-flash | 1M | $0.188 / $1.125 | 2026-04-27 | |||
| DeepSeek V4 Prodeepseek/deepseek-v4-pro | 1M | $0.435 / $0.87 | 2026-04-24 | |||
| DeepSeek V4 Flashdeepseek/deepseek-v4-flash | 1M | $0.14 / $0.28 | 2026-04-24 | |||
| GPT-5.5 Proopenai/gpt-5.5-pro | 1.05M | $30 / $180 | 2026-04-23 | |||
| GPT-5.5openai/gpt-5.5 | 1.05M | $5 / $30 | 2026-04-23 | |||
| MiMo-V2.5-Proxiaomi/mimo-v2.5-pro | 1.04858M | $0.435 / $0.87 | 2026-04-22 | |||
| MiMo-V2.5xiaomi/mimo-v2.5 | 1.04858M | $0.14 / $0.28 | 2026-04-22 | |||
| Qwen3.6 27Balibaba/qwen3.6-27b | 262.144K | $0.6 / $3.6 | 2026-04-22 | |||
| Gemini Embedding 2google/gemini-embedding-2 | 8.192K | $0.2 / - | 2026-04-22 | |||
| Kimi K2.6moonshotai/kimi-k2.6 | 262.144K | $0.95 / $4 | 2026-04-21 | |||
| GPT-Image-2openai/gpt-image-2 | Not documented | $5 / $30 | 2026-04-21 | |||
| Gemini Deep Research Previewgoogle/deep-research-preview-04-2026 | 1.04858M | - / - | 2026-04-21 | |||
| Deep Research Max Previewgoogle/deep-research-max-preview-04-2026 | 1.04858M | - / - | 2026-04-21 | |||
| Hy3 previewtencent/hy3-preview | 256K | $0.063 / $0.21 | 2026-04-20 | |||
| Qwen3.6 Max Previewalibaba/qwen3.6-max-preview | 262.144K | $1.3 / $7.8 | 2026-04-20 | |||
| Grok 4.3xai/grok-4.3 | 1M | $1.25 / $2.5 | 2026-04-17 | |||
| Qwen3.6 35B-A3Balibaba/qwen3.6-35b-a3b | 262.144K | $0.248 / $1.485 | 2026-04-17 | |||
| Grok Build 0.1xai/grok-build-0.1 | 256K | $1 / $2 | 2026-04-16 | |||
| Nemotron 3 Content Safetynvidia/nemotron-3-content-safety | 128K | - / - | 2026-04-16 | |||
| Claude Opus 4.7anthropic/claude-opus-4-7 | 1M | $5 / $25 | 2026-04-16 | |||
| Gemini 3.1 Flash TTS Previewgoogle/gemini-3.1-flash-tts-preview | 8.192K | $1 / $20 | 2026-04-15 | |||
| Gemini Robotics-ER 1.6 Previewgoogle/gemini-robotics-er-1.6-preview | 131.072K | $1 / $5 | 2026-04-14 | |||
| Muse Spark 1.1meta/muse-spark-1.1 | 1M | $1.25 / $4.25 | 2026-04-08 | |||
| GLM-5.1zhipuai/glm-5.1 | 200K | $1.4 / $4.4 | 2026-04-07 | |||
| Gemma 4 26B A4B ITgoogle/gemma-4-26b-a4b-it | 262.144K | $0.06 / $0.33 | 2026-04-02 | |||
| Gemma 4 31B ITgoogle/gemma-4-31b-it | 262.144K | $0.1 / $0.35 | 2026-04-02 | |||
| Gemma 4 E4B ITgoogle/gemma-4-E4B-it | 131.072K | $0.2 / $0.2 | 2026-04-02 | |||
| Gemma 4 E2B ITgoogle/gemma-4-E2B-it | 131.072K | $0.1 / $0.1 | 2026-04-02 | |||
| Step 3.5 Flash 2603stepfun/step-3.5-flash-2603 | 256K | $0.1 / $0.3 | 2026-04-02 | |||
| Qwen3.6 Plusalibaba/qwen3.6-plus | 1M | $0.5 / $3 | 2026-04-02 | |||
| GLM-5V-Turbozhipuai/glm-5v-turbo | 200K | $5 / $22 | 2026-04-01 | |||
| Llama Nemotron Rerank VL 1B v2nvidia/llama-nemotron-rerank-vl-1b-v2 | 128K | - / - | 2026-03-31 | |||
| Veo 3.1 Lite Previewgoogle/veo-3.1-lite-generate-preview | 1.024K | - / - | 2026-03-31 | |||
| Gemini 3.1 Flash Live Previewgoogle/gemini-3.1-flash-live-preview | 131.072K | $0.75 / $4.5 | 2026-03-26 | |||
| Lyria 3 Pro Previewgoogle/lyria-3-pro-preview | 131.072K | - / - | 2026-03-25 | |||
| Lyria 3 Clip Previewgoogle/lyria-3-clip-preview | 131.072K | - / - | 2026-03-25 | |||
| Nemotron Cascade 2 30B A3Bnvidia/nemotron-cascade-2-30b-a3b | 256K | - / - | 2026-03-24 | |||
| MiMo-V2-Omnixiaomi/mimo-v2-omni | 262.144K | $0.14 / $0.28 | 2026-03-18 | |||
| MiMo-V2-Proxiaomi/mimo-v2-pro | 1.04858M | $0.435 / $0.87 | 2026-03-18 | |||
| MiniMax-M2.7-highspeedminimax/MiniMax-M2.7-highspeed | 204.8K | $0.6 / $2.4 | 2026-03-18 | |||
| MiniMax-M2.7minimax/MiniMax-M2.7 | 204.8K | $0.3 / $1.2 | 2026-03-18 | |||
| GPT-5.4 nanoopenai/gpt-5.4-nano | 400K | $0.2 / $1.25 | 2026-03-17 | |||
| GPT-5.4 miniopenai/gpt-5.4-mini | 400K | $0.75 / $4.5 | 2026-03-17 | |||
| GLM-5-Turbozhipuai/glm-5-turbo | 200K | $0.9 / $3.7 | 2026-03-16 | |||
| Mistral Small (latest)mistral/mistral-small-latest | 256K | $0.15 / $0.6 | 2026-03-16 | |||
| Mistral Small 4mistral/mistral-small-2603 | 256K | $0.15 / $0.6 | 2026-03-16 |