mistral/mistral-small-2506
Efficient Mistral model for fast chat, extraction, and production assistants
- Context
- 128K
- Input
- $0.1/M
- Output
- $0.3/M
Filter 978 source-linked models by creator, price, context, modality, and published benchmark coverage.
mistral/mistral-small-2506
Efficient Mistral model for fast chat, extraction, and production assistants
google/gemini-2.5-flash-lite
Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents
google/gemini-2.5-flash
Fast Gemini workhorse for multimodal apps where latency and price matter
google/gemini-2.5-pro
Google's proven reasoning model for coding, math, and multimodal analysis
nvidia/mistral-nemotron
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
openai/o3-pro
High-effort o3 tier for difficult technical reasoning and careful answers
anthropic/claude-opus-4-0
Flagship Claude model for deep reasoning, coding, and long-horizon agents
anthropic/claude-opus-4-20250514
Flagship Claude model for deep reasoning, coding, and long-horizon agents
anthropic/claude-sonnet-4-20250514
Balanced Claude model for coding, analysis, agent workflows, and cost control
anthropic/claude-sonnet-4-0
Balanced Claude model for coding, analysis, agent workflows, and cost control
google/gemini-embedding-001
Embedding model for semantic search, retrieval, clustering, and ranking pipelines
mistral/mistral-medium-2505
Mistral model for multilingual chat, reasoning, and tool-assisted workflows
openai/gpt-image-1
OpenAI image model for production generation, edits, and brand-safe visual workflows
openai/o3
Deliberate o-series reasoner for hard math, coding, and multi-step analysis
openai/o4-mini
Fast o-series model for compact reasoning, coding, and tool use
nvidia/llama-3.1-nemotron-70b-instruct
Nemotron model for efficient reasoning, coding, and specialized AI agents
openai/gpt-4.1
Long-lived GPT workhorse for coding, instruction following, and production apps
openai/gpt-4.1-mini
Affordable GPT-4.1 lane for fast coding help and structured extraction
openai/gpt-4.1-nano
Tiny GPT-4.1 option for classification, routing, and very high-volume tasks
nvidia/llama-3.3-nemotron-super-49b-v1
Nemotron model for efficient reasoning, coding, and specialized AI agents
nvidia/llama-3.1-nemotron-ultra-253b
Flagship Nemotron model for high-throughput reasoning and complex agents
meta/llama-4-scout-17b-instruct
Open Llama with long-context vision for efficient multimodal agents
meta/llama-4-maverick-17b-instruct
Open multimodal Llama for strong reasoning with efficient everyday serving
alibaba/qwen3-coder-30b-a3b-instruct
Smaller Qwen coder for efficient local agents and repo-level fixes
alibaba/qwen3-32b
Dense open Qwen model for self-hosted chat, reasoning, and coding
alibaba/qwen3-235b-a22b
Large open Qwen MoE for multilingual reasoning, coding, and tool use
alibaba/qwen3-coder-480b-a35b-instruct
Open Qwen coding heavyweight for repository reasoning and agentic engineering
openai/o1-pro
O-series reasoning model for hard analysis, math, coding, and planning
mistral/magistral-medium-latest
Mistral reasoning model for transparent analysis, math, and complex decisions
cohere/command-a-03-2025
Cohere command model for multilingual enterprise agents, tools, and chat
alibaba/qwq-plus
Qwen reasoning model for deliberate problem solving, math, and coding
cohere/c4ai-aya-vision-32b
Open multilingual vision model for OCR, visual reasoning, and image question answering
cohere/c4ai-aya-vision-8b
Compact open multilingual vision model for OCR and visual question answering
cohere/command-r7b-arabic-02-2025
Open Command R model optimized for Arabic enterprise chat, RAG, and cultural knowledge
anthropic/claude-3-7-sonnet-20250219
Balanced Claude model for coding, analysis, agent workflows, and cost control
deepseek/deepseek-r1
Classic open reasoning model for transparent math, coding, and deliberate problem solving
alibaba/qwen-omni-turbo
Qwen omni model for text, vision, audio, and multimodal agent tasks
openai/o3-mini
Smaller o-series reasoner for economical coding, math, and planning tasks
google/gemini-2.0-flash
Earlier Gemini Flash workhorse for responsive multimodal apps and tool use
google/gemini-2.0-flash-lite
Low-latency Gemini model for high-volume multimodal and agent workloads
meta/llama-3.3-70b-instruct
Popular open Llama workhorse for multilingual chat, coding, and self-hosting
openai/o1
O-series reasoning model for hard analysis, math, coding, and planning
cohere/command-r7b-12-2024
Cohere retrieval model for long-context chat and enterprise RAG workflows
openai/gpt-4o-2024-11-20
GPT model for general reasoning, writing, coding, and tool-assisted tasks
mistral/mistral-large-2411
Flagship Mistral model for advanced reasoning, coding, and multilingual work
mistral/mistral-large-latest
Flagship Mistral model for advanced reasoning, coding, and multilingual work
mistral/pixtral-large-latest
Mistral's larger vision model for document-heavy image understanding and chat
mistral/mistral-large-2512
Mistral's largest general model for enterprise agents, coding, and multilingual reasoning
alibaba/qwen-turbo
Efficient Qwen model for fast chat, extraction, and high-volume workloads
cohere/c4ai-aya-expanse-32b
Open multilingual model optimized for generation across 23 languages
| Model | Creator | Input types | Context | Input / Output | Released | Compare |
|---|---|---|---|---|---|---|
| Mistral Small 3.2mistral/mistral-small-2506 | 128K | $0.1 / $0.3 | 2025-06-20 | |||
| Gemini 2.5 Flash-Litegoogle/gemini-2.5-flash-lite | 1.04858M | $0.1 / $0.4 | 2025-06-17 | |||
| Gemini 2.5 Flashgoogle/gemini-2.5-flash | 1.04858M | $0.3 / $2.5 | 2025-06-17 | |||
| Gemini 2.5 Progoogle/gemini-2.5-pro | 1.04858M | $1.25 / $10 | 2025-06-17 | |||
| Mistral Nemotronnvidia/mistral-nemotron | 128K | - / - | 2025-06-11 | |||
| o3-proopenai/o3-pro | 200K | $20 / $80 | 2025-06-10 | |||
| Claude Opus 4 (latest)anthropic/claude-opus-4-0 | 200K | $15 / $75 | 2025-05-22 | |||
| Claude Opus 4anthropic/claude-opus-4-20250514 | 200K | $15 / $75 | 2025-05-22 | |||
| Claude Sonnet 4anthropic/claude-sonnet-4-20250514 | 200K | $3 / $15 | 2025-05-22 | |||
| Claude Sonnet 4 (latest)anthropic/claude-sonnet-4-0 | 200K | $3 / $15 | 2025-05-22 | |||
| Gemini Embedding 001google/gemini-embedding-001 | 2.048K | $0.15 / - | 2025-05-20 | |||
| Mistral Medium 3mistral/mistral-medium-2505 | 131.072K | $0.4 / $2 | 2025-05-07 | |||
| GPT-Image-1openai/gpt-image-1 | Not documented | $5 / $40 | 2025-04-24 | |||
| o3openai/o3 | 200K | $2 / $8 | 2025-04-16 | |||
| o4-miniopenai/o4-mini | 200K | $1.1 / $4.4 | 2025-04-16 | |||
| Llama 3.1 Nemotron 70B Instructnvidia/llama-3.1-nemotron-70b-instruct | 128K | - / - | 2025-04-15 | |||
| GPT-4.1openai/gpt-4.1 | 1.04758M | $2 / $8 | 2025-04-14 | |||
| GPT-4.1 miniopenai/gpt-4.1-mini | 1.04758M | $0.4 / $1.6 | 2025-04-14 | |||
| GPT-4.1 nanoopenai/gpt-4.1-nano | 1.04758M | $0.1 / $0.4 | 2025-04-14 | |||
| Llama 3.3 Nemotron Super 49B v1nvidia/llama-3.3-nemotron-super-49b-v1 | 131.072K | - / - | 2025-04-07 | |||
| Llama 3.1 Nemotron Ultra 253Bnvidia/llama-3.1-nemotron-ultra-253b | 128K | $0.6 / $1.8 | 2025-04-07 | |||
| Llama 4 Scout 17B Instructmeta/llama-4-scout-17b-instruct | 3.5M | $0.17 / $0.66 | 2025-04-05 | |||
| Llama 4 Maverick 17B Instructmeta/llama-4-maverick-17b-instruct | 1M | $0.124 / $0.603 | 2025-04-05 | |||
| Qwen3-Coder 30B-A3B Instructalibaba/qwen3-coder-30b-a3b-instruct | 262.144K | $0.45 / $2.25 | 2025-04 | |||
| Qwen3 32Balibaba/qwen3-32b | 131.072K | $0.7 / $2.8 | 2025-04 | |||
| Qwen3 235B-A22Balibaba/qwen3-235b-a22b | 131.072K | $0.7 / $2.8 | 2025-04 | |||
| Qwen3-Coder 480B-A35B Instructalibaba/qwen3-coder-480b-a35b-instruct | 262.144K | $1.5 / $7.5 | 2025-04 | |||
| o1-proopenai/o1-pro | 200K | $150 / $600 | 2025-03-19 | |||
| Magistral Medium (latest)mistral/magistral-medium-latest | 128K | $2 / $5 | 2025-03-17 | |||
| Command Acohere/command-a-03-2025 | 256K | $2.5 / $10 | 2025-03-13 | |||
| QwQ Plusalibaba/qwq-plus | 131.072K | $0.8 / $2.4 | 2025-03-05 | |||
| Aya Vision 32Bcohere/c4ai-aya-vision-32b | 16K | - / - | 2025-03-04 | |||
| Aya Vision 8Bcohere/c4ai-aya-vision-8b | 16K | - / - | 2025-03-04 | |||
| Command R7B Arabiccohere/command-r7b-arabic-02-2025 | 128K | $0.037 / $0.15 | 2025-02-27 | |||
| Claude Sonnet 3.7anthropic/claude-3-7-sonnet-20250219 | 200K | $3 / $15 | 2025-02-19 | |||
| DeepSeek-R1deepseek/deepseek-r1 | 128K | $0.7 / $2.5 | 2025-01-20 | |||
| Qwen-Omni Turboalibaba/qwen-omni-turbo | 32.768K | $0.07 / $0.27 | 2025-01-19 | |||
| o3-miniopenai/o3-mini | 200K | $1.1 / $4.4 | 2024-12-20 | |||
| Gemini 2.0 Flashgoogle/gemini-2.0-flash | 1.04858M | $0.1 / $0.4 | 2024-12-11 | |||
| Gemini 2.0 Flash-Litegoogle/gemini-2.0-flash-lite | 1.04858M | $0.075 / $0.3 | 2024-12-11 | |||
| Llama-3.3-70B-Instructmeta/llama-3.3-70b-instruct | 128K | $0.13 / $0.4 | 2024-12-06 | |||
| o1openai/o1 | 200K | $15 / $60 | 2024-12-05 | |||
| Command R7Bcohere/command-r7b-12-2024 | 128K | $0.037 / $0.15 | 2024-12-02 | |||
| GPT-4o (2024-11-20)openai/gpt-4o-2024-11-20 | 128K | $2.5 / $10 | 2024-11-20 | |||
| Mistral Large 2.1mistral/mistral-large-2411 | 131.072K | $2 / $6 | 2024-11-18 | |||
| Mistral Large (latest)mistral/mistral-large-latest | 262.144K | $0.5 / $1.5 | 2024-11-01 | |||
| Pixtral Large (latest)mistral/pixtral-large-latest | 128K | $2 / $6 | 2024-11-01 | |||
| Mistral Large 3mistral/mistral-large-2512 | 262.144K | $0.5 / $1.5 | 2024-11-01 | |||
| Qwen Turboalibaba/qwen-turbo | 1M | $0.05 / $0.2 | 2024-11-01 | |||
| Aya Expanse 32Bcohere/c4ai-aya-expanse-32b | 128K | - / - | 2024-10-24 |