495 results Clear filters
Undated
nousresearch/hermes-3-llama-3.1-405b

Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...

T Weight access not listed
Context
131.072K
Input
$1/M
Output
$1/M
meta-llama/llama-3.1-8b-instruct

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...

T Weight access not listed
Context
131.072K
Input
$0.05/M
Output
$0.08/M
meta-llama/llama-3.1-70b-instruct

Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 70B instruct-tuned version is optimized for high quality dialogue usecases. It has demonstrated strong...

T Weight access not listed
Context
131.072K
Input
$0.4/M
Output
$0.4/M
Undated
mistralai/mistral-nemo

A 12B parameter model with a 128k token context length built by Mistral in collaboration with NVIDIA. The model is multilingual, supporting English, French, German, Spanish, Italian, Portuguese, Chinese, Japanese,...

T Weight access not listed
Context
131.072K
Input
$0.019/M
Output
$0.03/M
openai/gpt-4o-mini-2024-07-18

GPT-4o mini is OpenAI's newest model after [GPT-4 Omni](/models/openai/gpt-4o), supporting both text and image inputs with text outputs. As their most advanced small model, it is many multiples more affordable...

T Weight access not listed
Context
128K
Input
$0.15/M
Output
$0.6/M
Undated
anthropic/claude-3-haiku

Claude 3 Haiku is Anthropic's fastest and most compact model for near-instant responsiveness. Quick and accurate targeted performance. See the launch announcement and benchmark results [here](https://www.anthropic.com/news/claude-3-haiku) #multimodal

T Weight access not listed
Context
200K
Input
$0.25/M
Output
$1.25/M
Mistral Largeby Mistral AI
Undated
mistralai/mistral-large

This is Mistral AI's flagship model, Mistral Large 2 (version `mistral-large-2407`). It's a proprietary weights-available model and excels at reasoning, code, JSON, chat, and more. Read the launch announcement [here](https://mistral.ai/news/mistral-large-2407/)....

T Weight access not listed
Context
128K
Input
$2/M
Output
$6/M
Undated
openai/gpt-4-turbo-preview

The preview GPT-4 model with improved instruction following, JSON mode, reproducible outputs, parallel function calling, and more. Training data: up to Dec 2023. **Note:** heavily rate limited by OpenAI while...

T Weight access not listed
Context
128K
Input
$10/M
Output
$30/M
Auto Routerby Openrouter
Undated
openrouter/auto

Your prompt will be processed by a meta-model and routed to one of dozens of models (see below), optimizing for the best possible output. To see which model was used,...

T Weight access not listed
Context
2M
Input
-
Output
-
Undated
kwaipilot/kat-coder-air-v2.5

KAT-Coder-Air V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...

T Weight access not listed
Context
256K
Input
$0.15/M
Output
$0.6/M
Undated
kwaipilot/kat-coder-pro-v2.5

KAT-Coder-Pro V2.5 is a flagship-level Agentic Coding model that can directly hand over an entire issue or an entire business workflow to it, allowing it to autonomously locate and make...

T Weight access not listed
Context
256K
Input
$0.74/M
Output
$2.96/M
Auto Router (Beta)by Openrouter
Undated
openrouter/auto-beta

Auto Router (Beta) is a task-aware router from OpenRouter. It classifies each request, then routes it the [most popular model](/rankings#task-spend) for that task based on aggregate spend, filtered by your...

T Weight access not listed
Context
2M
Input
-
Output
-
Undated
poolside/laguna-s-2.1:free

Laguna S 2.1 is the latest coding agent model from [Poolside](<https://poolside.ai/>). Laguna S 2.1 is a 118B total parameter model with 8B active parameters, scoring 70.2% on Terminal-Bench 2.1 and...

T Weight access not listed
Context
262.144K
Input
Free
Output
Free
Ling-3.0-flash (free)by Inclusionai
Undated
inclusionai/ling-3.0-flash:free

*Ling-3.0-flash* is a *124B-parameter Mixture-of-Experts (MoE) model*, with approximately *5.1B parameters activated per token*. The model is designed with *token efficiency and production-scale agentic inference* as key priorities, enabling developers...

T Weight access not listed
Context
262.144K
Input
Free
Output
Free
Undated
anthropic/claude-opus-5-fast

Fast-mode variant of [Opus 5](/anthropic/claude-opus-5) - identical capabilities with higher output speed at 2x pricing relative to regular Opus 5. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode

T Weight access not listed
Context
1M
Input
$10/M
Output
$50/M
Qwen: Qwen3.7 Flashby Alibaba Qwen
Undated
qwen/qwen3.7-flash

Qwen3.7 Flash is a vision-language reasoning model from Alibaba. It is suited for multimodal agents, visual coding, search, and computer interaction, with strengths in object recognition, spatial understanding, and real-world...

T Weight access not listed
Context
1M
Input
$0.03/M
Output
$0.13/M
google/gemini-3.6-flash:batch

Gemini 3.6 Flash is a high-efficiency model from Google for coding, agentic workflows, and web and app development. It is designed to produce polished outputs with fewer unnecessary edits and...

T Weight access not listed
Context
1.04858M
Input
$0.75/M
Output
$3.75/M
google/gemini-3.5-flash-lite:batch

Gemini 3.5 Flash Lite is a high-efficiency model from Google with upgraded agentic capabilities. It is suited for subagents that execute focused tasks within complex, multi-agent workflows.

T Weight access not listed
Context
1.04858M
Input
$0.15/M
Output
$1.25/M
anthropic/claude-sonnet-5:batch

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

T Weight access not listed
Context
1M
Input
$1/M
Output
$5/M
Undated
anthropic/claude-fable-5:batch

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

T Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
Undated
minimax/minimax-m3:batch

MiniMax-M3 is a multimodal foundation model from MiniMax. It supports text, image, and video inputs with text output, a 1M-token context window, and is suited for long-horizon agentic work, coding,...

T Weight access not listed
Context
524.288K
Input
$0.15/M
Output
$0.6/M
anthropic/claude-opus-4.8:batch

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

T Weight access not listed
Context
1M
Input
$2.5/M
Output
$12.5/M
google/gemini-3.5-flash:batch

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

T Weight access not listed
Context
1.04858M
Input
$0.75/M
Output
$4.5/M
google/gemini-3.1-flash-lite:batch

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...

T Weight access not listed
Context
1.04858M
Input
$0.125/M
Output
$0.75/M
Undated
openai/gpt-5.5:batch

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...

T Weight access not listed
Context
1.05M
Input
$2.5/M
Output
$15/M
anthropic/claude-opus-4.7:batch

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

T Weight access not listed
Context
1M
Input
$2.5/M
Output
$12.5/M
Undated
openai/gpt-5.4-nano:batch

GPT-5.4 nano is the most lightweight and cost-efficient variant of the GPT-5.4 family, optimized for speed-critical and high-volume tasks. It supports text and image inputs and is designed for low-latency...

T Weight access not listed
Context
400K
Input
$0.1/M
Output
$0.625/M
Undated
openai/gpt-5.4-mini:batch

GPT-5.4 mini brings the core capabilities of GPT-5.4 to a faster, more efficient model optimized for high-throughput workloads. It supports text and image inputs with strong performance across reasoning, coding,...

T Weight access not listed
Context
400K
Input
$0.375/M
Output
$2.25/M
Undated
openai/gpt-5.4:batch

GPT-5.4 is OpenAI’s latest frontier model, unifying the Codex and GPT lines into a single system. It features a 1M+ token context window (922K input, 128K output) with support for...

T Weight access not listed
Context
1.05M
Input
$1.25/M
Output
$7.5/M
google/gemini-3.1-pro-preview:batch

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...

T Weight access not listed
Context
1.04858M
Input
$1/M
Output
$6/M
anthropic/claude-opus-4.6:batch

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...

T Weight access not listed
Context
1M
Input
$2.5/M
Output
$12.5/M
google/gemini-3-flash-preview:batch

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...

T Weight access not listed
Context
1.04858M
Input
$0.25/M
Output
$1.5/M
Undated
openai/gpt-5.2:batch

GPT-5.2 is the latest frontier-grade model in the GPT-5 series, offering stronger agentic and long context perfomance compared to GPT-5.1. It uses adaptive reasoning to allocate computation dynamically, responding quickly...

T Weight access not listed
Context
400K
Input
$0.875/M
Output
$7/M
anthropic/claude-opus-4.5:batch

Claude Opus 4.5 is Anthropic’s frontier reasoning model optimized for complex software engineering, agentic workflows, and long-horizon computer use. It offers strong multimodal capabilities, competitive performance across real-world coding and...

T Weight access not listed
Context
200K
Input
$2.5/M
Output
$12.5/M
Undated
openai/gpt-5.1:batch

GPT-5.1 is the latest frontier-grade model in the GPT-5 series, offering stronger general-purpose reasoning, improved instruction adherence, and a more natural conversational style compared to GPT-5. It uses adaptive reasoning...

T Weight access not listed
Context
400K
Input
$0.625/M
Output
$5/M
anthropic/claude-haiku-4.5:batch

Claude Haiku 4.5 is Anthropic’s fastest and most efficient model, delivering near-frontier intelligence at a fraction of the cost and latency of larger Claude models. Matching Claude Sonnet 4’s performance...

T Weight access not listed
Context
200K
Input
$0.5/M
Output
$2.5/M
anthropic/claude-sonnet-4.5:batch

Claude Sonnet 4.5 is Anthropic’s most advanced Sonnet model to date, optimized for real-world agents and coding workflows. It delivers state-of-the-art performance on coding benchmarks such as SWE-bench Verified, with...

T Weight access not listed
Context
1M
Input
$1.5/M
Output
$7.5/M
Undated
openai/gpt-5:batch

GPT-5 is OpenAI’s most advanced model, offering major improvements in reasoning, code quality, and user experience. It is optimized for complex tasks that require step-by-step reasoning, instruction following, and accuracy...

T Weight access not listed
Context
400K
Input
$0.625/M
Output
$5/M
Undated
openai/gpt-5-mini:batch

GPT-5 Mini is a compact version of GPT-5, designed to handle lighter-weight reasoning tasks. It provides the same instruction-following and safety-tuning benefits as GPT-5, but with reduced latency and cost....

T Weight access not listed
Context
400K
Input
$0.125/M
Output
$1/M
Undated
openai/gpt-5-nano:batch

GPT-5-Nano is the smallest and fastest variant in the GPT-5 system, optimized for developer tools, rapid interactions, and ultra-low latency environments. While limited in reasoning depth compared to its larger...

T Weight access not listed
Context
400K
Input
$0.025/M
Output
$0.2/M
anthropic/claude-opus-4.1:batch

Claude Opus 4.1 is an updated version of Anthropic’s flagship model, offering improved performance in coding, reasoning, and agentic tasks. It achieves 74.5% on SWE-bench Verified and shows notable gains...

T Weight access not listed
Context
200K
Input
$7.5/M
Output
$37.5/M
google/gemini-2.5-flash-lite:batch

Gemini 2.5 Flash-Lite is a lightweight reasoning model in the Gemini 2.5 family, optimized for ultra-low latency and cost efficiency. It offers improved throughput, faster token generation, and better performance...

T Weight access not listed
Context
1.04858M
Input
$0.05/M
Output
$0.2/M
google/gemini-2.5-flash:batch

Gemini 2.5 Flash is Google's state-of-the-art workhorse model, specifically designed for advanced reasoning, coding, mathematics, and scientific tasks. It includes built-in "thinking" capabilities, enabling it to provide responses with greater...

T Weight access not listed
Context
1.04858M
Input
$0.15/M
Output
$1.25/M
google/gemini-2.5-pro:batch

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

T Weight access not listed
Context
1.04858M
Input
$0.625/M
Output
$5/M
deepseek/deepseek-v4-flash-0731

DeepSeek V4 Flash 0731 is a sparse mixture-of-experts model from DeepSeek, with 13B active parameters out of 284B total. This re-post-trained revision is suited for coding, reasoning, and agent workflows.

T Weight access not listed
Context
1.04858M
Input
$0.14/M
Output
$0.28/M