978 results
Qwen: Qwen3.7 Maxby Alibaba Qwen
Undated
qwen/qwen3.7-max

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...

T Weight access not listed
Context
1M
Input
$1.475/M
Output
$4.425/M
Undated
x-ai/grok-build-0.1

Grok Build 0.1 is xAI’s fast coding model trained specifically for agentic software engineering workflows. It supports text and image inputs with text output, and is optimized for interactive coding...

T Weight access not listed
Context
256K
Input
$1/M
Output
$2/M
Undated
anthropic/claude-opus-4.7-fast

Fast-mode variant of [Opus 4.7](/anthropic/claude-opus-4.7) - identical capabilities with higher output speed at premium 6x pricing. Learn more in Anthropic's docs: https://platform.claude.com/docs/en/build-with-claude/fast-mode

T Weight access not listed
Context
1M
Input
$30/M
Output
$150/M
Undated
perceptron/perceptron-mk1

Perceptron Mk1 (Mark One) is Perceptron's highest-quality vision-language model for video and embodied reasoning.** It accepts image and video inputs paired with natural language queries, and produces detailed visual understanding...

T Weight access not listed
Context
32.768K
Input
$0.15/M
Output
$1.5/M
Undated
inclusionai/ring-2.6-1t

Ring-2.6-1T is a 1T-parameter-scale thinking model with 63B active parameters, built for real-world agent workflows that require both strong capability and operational efficiency. It is optimized for coding agents, tool...

T Weight access not listed
Context
262.144K
Input
$0.075/M
Output
$0.625/M
Undated
openai/gpt-chat-latest

GPT Chat Latest points to OpenAI's stable API alias `chat-latest` that always resolves to the latest Instant chat model used in ChatGPT. As OpenAI rolls out new Instant model updates...

T Weight access not listed
Context
400K
Input
$5/M
Output
$30/M
Undated
x-ai/grok-4.3

Grok 4.3 is a reasoning model from xAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

T Weight access not listed
Context
1M
Input
$1.25/M
Output
$2.5/M
Undated
ibm-granite/granite-4.1-8b

Granite 4.1 8B is a dense, decoder-only 8-billion-parameter language model from IBM, part of the Granite 4.1 family. It supports a 131K-token context window and is designed for enterprise tasks...

T Open weights
Context
131.072K
Input
$0.05/M
Output
$0.1/M
Undated
mistralai/mistral-medium-3-5

Mistral Medium 3.5 is a dense 128B instruction-following model from Mistral AI. It supports text and image inputs with text output, and is designed for agentic workflows, coding, and complex...

T Weight access not listed
Context
262.144K
Input
$1.5/M
Output
$7.5/M
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free

NVIDIA Nemotron™ 3 Nano Omni is a 30B-A3B open multimodal model designed to function as a perception and context sub-agent in enterprise agent systems. It accepts text, image, video, and...

T Weight access not listed
Context
256K
Input
Free
Output
Free
Undated
poolside/laguna-m.1:free

Laguna M.1 is the flagship coding agent model from [Poolside](https://poolside.ai/), optimized for complex software engineering tasks. Designed for agentic coding workflows, it supports tool calling and reasoning, with a 256K...

T Weight access not listed
Context
262.144K
Input
Free
Output
Free
Undated
~anthropic/claude-haiku-latest

This model always redirects to the latest model in the Anthropic Claude Haiku family.

T Weight access not listed
Context
200K
Input
$1/M
Output
$5/M
Undated
~openai/gpt-mini-latest

This model always redirects to the latest model in the OpenAI GPT Mini family.

T Weight access not listed
Context
400K
Input
$0.75/M
Output
$4.5/M
Undated
~google/gemini-pro-latest

This model always redirects to the latest model in the Google Gemini Pro family.

T Weight access not listed
Context
1.04858M
Input
$2/M
Output
$12/M
Undated
~moonshotai/kimi-latest

This model always redirects to the latest model in the MoonshotAI Kimi family.

T Weight access not listed
Context
1.04858M
Input
$2.9/M
Output
$14/M
Undated
~google/gemini-flash-latest

This model always redirects to the latest model in the Google Gemini Flash family.

T Weight access not listed
Context
1.04858M
Input
$1.5/M
Output
$7.5/M
Undated
~anthropic/claude-sonnet-latest

This model always redirects to the latest model in the Anthropic Claude Sonnet family.

T Weight access not listed
Context
1M
Input
$2/M
Output
$10/M
Undated
~openai/gpt-latest

This model always redirects to the latest model in the OpenAI GPT family.

T Weight access not listed
Context
1.05M
Input
$5/M
Output
$30/M
Undated
qwen/qwen3.5-plus-20260420

Qwen3.5 Plus (April 2026) is a large-scale multimodal language model from Alibaba. It accepts text, image, and video input and produces text output, with a 1M token context window. This...

T Weight access not listed
Context
1M
Input
$0.3/M
Output
$1.8/M
Qwen: Qwen3.6 Flashby Alibaba Qwen
Undated
qwen/qwen3.6-flash

Qwen3.6 Flash is a fast, efficient language model from Alibaba's Qwen 3.6 series. It supports text, image, and video input with a 1M token context window. Tiered pricing kicks in...

T Weight access not listed
Context
1M
Input
$0.188/M
Output
$1.125/M
Qwen: Qwen3.6 35B A3Bby Alibaba Qwen
Undated
qwen/qwen3.6-35b-a3b

Qwen3.6-35B-A3B is an open-weight multimodal model from Alibaba Cloud with 35 billion total parameters and 3 billion active parameters per token. It uses a hybrid sparse mixture-of-experts architecture combining Gated...

T Weight access not listed
Context
262.144K
Input
$0.14/M
Output
$1/M
Undated
qwen/qwen3.6-max-preview

Qwen3.6-Max-Preview is a proprietary frontier model from Alibaba Cloud built on a sparse mixture-of-experts architecture with approximately 1 trillion total parameters. It is optimized for agentic coding, tool use, and...

T Weight access not listed
Context
262.144K
Input
$1.027/M
Output
$6.162/M
Qwen: Qwen3.6 27Bby Alibaba Qwen
Undated
qwen/qwen3.6-27b

Qwen3.6 27B is a dense 27-billion-parameter language model from the Qwen Team at Alibaba, released in April 2026. It features hybrid multimodal capabilities — accepting text, image, and video inputs...

T Weight access not listed
Context
262.144K
Input
$0.3/M
Output
$2/M
Undated
inclusionai/ling-2.6-1t

Ling-2.6-1T is an instant (instruct) model from inclusionAI and the company’s trillion-parameter flagship, designed for real-world agents that require fast execution and high efficiency at scale. It uses a “fast...

T Weight access not listed
Context
262.144K
Input
$0.075/M
Output
$0.625/M
Undated
openai/gpt-5.4-image-2

[GPT-5.4](https://openrouter.ai/openai/gpt-5.4) Image 2 combines OpenAI's GPT-5.4 model with state-of-the-art image generation capabilities from GPT Image 2. It enables rich multimodal workflows, allowing users to seamlessly move between reasoning, coding, and...

T Weight access not listed
Context
272K
Input
$8/M
Output
$15/M
Undated
inclusionai/ling-2.6-flash

Ling-2.6-flash is an instant (instruct) model from inclusionAI with 104B total parameters and 7.4B active parameters, designed for real-world agents that require fast responses, strong execution, and high token efficiency....

T Weight access not listed
Context
262.144K
Input
$0.01/M
Output
$0.03/M
Undated
~anthropic/claude-opus-latest

This model always redirects to the latest model in the Claude Opus family.

T Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
Pareto Code Routerby Openrouter
Undated
openrouter/pareto-code

The Pareto Router maintains a tiered shortlist of strong coding models, ranked by [Artificial Analysis](https://artificialanalysis.ai/) coding percentiles. Set min_coding_score between 0 and 1 on the [pareto-router plugin](https://openrouter.ai/docs/guides/routing/routers/pareto-router#the-min_coding_score-parameter) to control how...

T Weight access not listed
Context
2M
Input
-
Output
-
Undated
anthropic/claude-opus-4.7

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

T Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
Undated
z-ai/glm-5.1

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...

T Weight access not listed
Context
200K
Input
$0.966/M
Output
$3.036/M
google/gemma-4-26b-a4b-it:free

Gemma 4 26B A4B IT is an instruction-tuned Mixture-of-Experts (MoE) model from Google DeepMind. Despite 25.2B total parameters, only 3.8B activate per token during inference — delivering near-31B quality at...

T Weight access not listed
Context
131.072K
Input
Free
Output
Free
Undated
google/gemma-4-31b-it:free

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

T Weight access not listed
Context
262.144K
Input
Free
Output
Free
Qwen: Qwen3.6 Plusby Alibaba Qwen
Undated
qwen/qwen3.6-plus

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...

T Weight access not listed
Context
1M
Input
$0.325/M
Output
$1.95/M
Undated
z-ai/glm-5v-turbo

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...

T Weight access not listed
Context
202.752K
Input
$1.2/M
Output
$4/M
arcee-ai/trinity-large-thinking

Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloads, and reasoning tasks. Launch video: https://youtu.be/Gc82AXLa0Rg?si=4RLn6WBz33qT--B7...

T Weight access not listed
Context
262.144K
Input
$0.22/M
Output
$0.85/M
x-ai/grok-4.20-multi-agent

Grok 4.20 Multi-Agent is a variant of xAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and synthesize information...

T Weight access not listed
Context
2M
Input
$1.25/M
Output
$2.5/M
Undated
x-ai/grok-4.20

Grok 4.20 is a reasoning model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...

T Weight access not listed
Context
2M
Input
$1.25/M
Output
$2.5/M
Undated
kwaipilot/kat-coder-pro-v2

KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engineering and SaaS integration. It builds on the agentic coding strengths of earlier versions,...

T Weight access not listed
Context
256K
Input
$0.3/M
Output
$1.2/M
Reka Edgeby Rekaai
Undated
rekaai/reka-edge

Reka Edge is an extremely efficient 7B multimodal vision-language model that accepts image/video+text inputs and generates text outputs. This model is optimized specifically to deliver industry-leading performance in image understanding,...

T Weight access not listed
Context
16.384K
Input
$0.1/M
Output
$0.1/M
Undated
minimax/minimax-m2.7

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...

T Weight access not listed
Context
196.608K
Input
$0.25/M
Output
$1/M
Undated
mistralai/mistral-small-2603

Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...

T Weight access not listed
Context
262.144K
Input
$0.15/M
Output
$0.6/M
Undated
z-ai/glm-5-turbo

GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows...

T Weight access not listed
Context
202.752K
Input
$1.2/M
Output
$4/M
nvidia/nemotron-3-super-120b-a12b:free

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

T Weight access not listed
Context
262.144K
Input
Free
Output
Free
Undated
bytedance-seed/seed-2.0-lite

Seed-2.0-Lite is a versatile, cost‑efficient enterprise workhorse that delivers strong multimodal and agent capabilities while offering noticeably lower latency, making it a practical default choice for most production workloads across...

T Weight access not listed
Context
262.144K
Input
$0.25/M
Output
$2/M
Qwen: Qwen3.5-9Bby Alibaba Qwen
Undated
qwen/qwen3.5-9b

Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...

T Weight access not listed
Context
262.144K
Input
$0.1/M
Output
$0.15/M
Undated
inception/mercury-2

Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving...

T Weight access not listed
Context
128K
Input
$0.25/M
Output
$0.75/M
Undated
openai/gpt-5.3-chat

GPT-5.3 Chat is an update to ChatGPT's most-used model that makes everyday conversations smoother, more useful, and more directly helpful. It delivers more accurate answers with better contextualization and significantly...

T Weight access not listed
Context
128K
Input
$1.75/M
Output
$14/M
Undated
bytedance-seed/seed-2.0-mini

Seed-2.0-mini targets latency-sensitive, high-concurrency, and cost-sensitive scenarios, emphasizing fast response and flexible inference deployment. It delivers performance comparable to ByteDance-Seed-1.6, supports 256k context, four reasoning effort modes (minimal/low/medium/high), multimodal understanding,...

T Weight access not listed
Context
262.144K
Input
$0.1/M
Output
$0.4/M
Qwen: Qwen3.5-35B-A3Bby Alibaba Qwen
Undated
qwen/qwen3.5-35b-a3b

The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...

T Weight access not listed
Context
262.144K
Input
$0.14/M
Output
$1/M
Qwen: Qwen3.5-27Bby Alibaba Qwen
Undated
qwen/qwen3.5-27b

The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of...

T Weight access not listed
Context
262.144K
Input
$0.195/M
Output
$1.56/M