494 results Clear filters
Undated
google/gemma-4-31b-it:free

Gemma 4 31B Instruct is Google DeepMind's 30.7B dense multimodal model supporting text and image input with text output. Features a 256K token context window, configurable thinking/reasoning mode, native function...

T Weight access not listed
Context
262.144K
Input
Free
Output
Free
Qwen: Qwen3.6 Plusby Alibaba Qwen
Undated
qwen/qwen3.6-plus

Qwen 3.6 Plus builds on a hybrid architecture that combines efficient linear attention with sparse mixture-of-experts routing, enabling strong scalability and high-performance inference. Compared to the 3.5 series, it delivers...

T Weight access not listed
Context
1M
Input
$0.325/M
Output
$1.95/M
Undated
z-ai/glm-5v-turbo

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...

T Weight access not listed
Context
202.752K
Input
$1.2/M
Output
$4/M
arcee-ai/trinity-large-thinking

Trinity Large Thinking is a powerful open source reasoning model from the team at Arcee AI. It shows strong performance in PinchBench, agentic workloads, and reasoning tasks. Launch video: https://youtu.be/Gc82AXLa0Rg?si=4RLn6WBz33qT--B7...

T Weight access not listed
Context
262.144K
Input
$0.22/M
Output
$0.85/M
x-ai/grok-4.20-multi-agent

Grok 4.20 Multi-Agent is a variant of xAI’s Grok 4.20 designed for collaborative, agent-based workflows. Multiple agents operate in parallel to conduct deep research, coordinate tool use, and synthesize information...

T Weight access not listed
Context
2M
Input
$1.25/M
Output
$2.5/M
Undated
x-ai/grok-4.20

Grok 4.20 is a reasoning model from xAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...

T Weight access not listed
Context
2M
Input
$1.25/M
Output
$2.5/M
Undated
kwaipilot/kat-coder-pro-v2

KAT-Coder-Pro V2 is the latest high-performance model in KwaiKAT’s KAT-Coder series, designed for complex enterprise-grade software engineering and SaaS integration. It builds on the agentic coding strengths of earlier versions,...

T Weight access not listed
Context
256K
Input
$0.3/M
Output
$1.2/M
Undated
minimax/minimax-m2.7

MiniMax-M2.7 is a next-generation large language model designed for autonomous, real-world productivity and continuous improvement. Built to actively participate in its own evolution, M2.7 integrates advanced agentic capabilities through multi-agent...

T Weight access not listed
Context
196.608K
Input
$0.25/M
Output
$1/M
Undated
mistralai/mistral-small-2603

Mistral Small 4 is the next major release in the Mistral Small family, unifying the capabilities of several flagship Mistral models into a single system. It combines strong reasoning from...

T Weight access not listed
Context
262.144K
Input
$0.15/M
Output
$0.6/M
Undated
z-ai/glm-5-turbo

GLM-5 Turbo is a new model from Z.ai designed for fast inference and strong performance in agent-driven environments such as OpenClaw scenarios. It is deeply optimized for real-world agent workflows...

T Weight access not listed
Context
202.752K
Input
$1.2/M
Output
$4/M
nvidia/nemotron-3-super-120b-a12b:free

NVIDIA Nemotron 3 Super is a 120B-parameter open hybrid MoE model, activating just 12B parameters for maximum compute efficiency and accuracy in complex multi-agent applications. Built on a hybrid Mamba-Transformer...

T Weight access not listed
Context
262.144K
Input
Free
Output
Free
Undated
bytedance-seed/seed-2.0-lite

Seed-2.0-Lite is a versatile, cost‑efficient enterprise workhorse that delivers strong multimodal and agent capabilities while offering noticeably lower latency, making it a practical default choice for most production workloads across...

T Weight access not listed
Context
262.144K
Input
$0.25/M
Output
$2/M
Qwen: Qwen3.5-9Bby Alibaba Qwen
Undated
qwen/qwen3.5-9b

Qwen3.5-9B is a multimodal foundation model from the Qwen3.5 family, designed to deliver strong reasoning, coding, and visual understanding in an efficient 9B-parameter architecture. It uses a unified vision-language design...

T Weight access not listed
Context
262.144K
Input
$0.1/M
Output
$0.15/M
Undated
inception/mercury-2

Mercury 2 is an extremely fast reasoning LLM, and the first reasoning diffusion LLM (dLLM). Instead of generating tokens sequentially, Mercury 2 produces and refines multiple tokens in parallel, achieving...

T Weight access not listed
Context
128K
Input
$0.25/M
Output
$0.75/M
Undated
openai/gpt-5.3-chat

GPT-5.3 Chat is an update to ChatGPT's most-used model that makes everyday conversations smoother, more useful, and more directly helpful. It delivers more accurate answers with better contextualization and significantly...

T Weight access not listed
Context
128K
Input
$1.75/M
Output
$14/M
Undated
bytedance-seed/seed-2.0-mini

Seed-2.0-mini targets latency-sensitive, high-concurrency, and cost-sensitive scenarios, emphasizing fast response and flexible inference deployment. It delivers performance comparable to ByteDance-Seed-1.6, supports 256k context, four reasoning effort modes (minimal/low/medium/high), multimodal understanding,...

T Weight access not listed
Context
262.144K
Input
$0.1/M
Output
$0.4/M
Qwen: Qwen3.5-35B-A3Bby Alibaba Qwen
Undated
qwen/qwen3.5-35b-a3b

The Qwen3.5 Series 35B-A3B is a native vision-language model designed with a hybrid architecture that integrates linear attention mechanisms and a sparse mixture-of-experts model, achieving higher inference efficiency. Its overall...

T Weight access not listed
Context
262.144K
Input
$0.14/M
Output
$1/M
Qwen: Qwen3.5-27Bby Alibaba Qwen
Undated
qwen/qwen3.5-27b

The Qwen3.5 27B native vision-language Dense model incorporates a linear attention mechanism, delivering fast response times while balancing inference speed and performance. Its overall capabilities are comparable to those of...

T Weight access not listed
Context
262.144K
Input
$0.195/M
Output
$1.56/M
Undated
qwen/qwen3.5-122b-a10b

The Qwen3.5 122B-A10B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. In terms of...

T Weight access not listed
Context
262.144K
Input
$0.26/M
Output
$2.08/M
Qwen: Qwen3.5-Flashby Alibaba Qwen
Undated
qwen/qwen3.5-flash-02-23

The Qwen3.5 native vision-language Flash models are built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. Compared to the...

T Weight access not listed
Context
1M
Input
$0.065/M
Output
$0.26/M
AionLabs: Aion-2.0by Aion Labs
Undated
aion-labs/aion-2.0

Aion-2.0 is a variant of DeepSeek V3.2 optimized for immersive roleplaying and storytelling. It is particularly strong at introducing tension, crises, and conflict into stories, making narratives feel more engaging....

T Weight access not listed
Context
131.072K
Input
$0.8/M
Output
$1.6/M
Undated
anthropic/claude-sonnet-4.6

Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, with frontier performance across coding, agents, and professional work. It excels at iterative development, complex codebase navigation, end-to-end project management with...

T Weight access not listed
Context
1M
Input
$3/M
Output
$15/M
Undated
qwen/qwen3.5-plus-02-15

The Qwen3.5 native vision-language series Plus models are built on a hybrid architecture that integrates linear attention mechanisms with sparse mixture-of-experts models, achieving higher inference efficiency. In a variety of...

T Weight access not listed
Context
1M
Input
$0.26/M
Output
$1.56/M
Undated
qwen/qwen3.5-397b-a17b

The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency. It delivers...

T Weight access not listed
Context
262.144K
Input
$0.39/M
Output
$2.34/M
Undated
minimax/minimax-m2.5

MiniMax-M2.5 is a SOTA large language model designed for real-world productivity. Trained in a diverse range of complex real-world digital working environments, M2.5 builds upon the coding expertise of M2.1...

T Weight access not listed
Context
196.608K
Input
$0.15/M
Output
$0.9/M
Undated
z-ai/glm-5

GLM-5 is Z.ai’s flagship open-source foundation model engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks, rivaling leading...

T Weight access not listed
Context
204.8K
Input
$0.95/M
Output
$2.55/M
Undated
qwen/qwen3-max-thinking

Qwen3-Max-Thinking is the flagship reasoning model in the Qwen3 series, designed for high-stakes cognitive tasks that require deep, multi-step reasoning. By significantly scaling model capacity and reinforcement learning compute, it...

T Weight access not listed
Context
262.144K
Input
$0.78/M
Output
$3.9/M
Undated
anthropic/claude-opus-4.6

Opus 4.6 is Anthropic’s strongest model for coding and long-running professional tasks. It is built for agents that operate across entire workflows rather than single prompts, making it especially effective...

T Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
Qwen: Qwen3 Coder Nextby Alibaba Qwen
Undated
qwen/qwen3-coder-next

Qwen3-Coder-Next is an open-weight causal language model optimized for coding agents and local development workflows. It uses a sparse MoE design with 80B total parameters and only 3B activated per...

T Weight access not listed
Context
262.144K
Input
$0.12/M
Output
$0.8/M
Free Models Routerby Openrouter
Undated
openrouter/free

The simplest way to get free inference. openrouter/free is a router that selects free models at random from the models available on OpenRouter. The router smartly filters for models that...

T Weight access not listed
Context
200K
Input
Free
Output
Free
Undated
upstage/solar-pro-3

Solar Pro 3 is Upstage's powerful Mixture-of-Experts (MoE) language model. With 102B total parameters and 12B active parameters per forward pass, it delivers exceptional performance while maintaining computational efficiency. Optimized...

T Weight access not listed
Context
128K
Input
$0.15/M
Output
$0.6/M
Undated
writer/palmyra-x5

Palmyra X5 is Writer's most advanced model, purpose-built for building and scaling AI agents across the enterprise. It delivers industry-leading speed and efficiency on context windows up to 1 million...

T Weight access not listed
Context
1.04M
Input
$0.6/M
Output
$6/M
Undated
openai/gpt-audio

The gpt-audio model is OpenAI's first generally available audio model. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Audio is priced...

T Weight access not listed
Context
128K
Input
$2.5/M
Output
$10/M
Undated
openai/gpt-audio-mini

A cost-efficient version of GPT Audio. The new snapshot features an upgraded decoder for more natural sounding voices and maintains better voice consistency. Input is priced at $0.60 per million...

T Weight access not listed
Context
128K
Input
$0.6/M
Output
$2.4/M
Undated
z-ai/glm-4.7-flash

As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...

T Weight access not listed
Context
202.752K
Input
$0.06/M
Output
$0.4/M
Undated
bytedance-seed/seed-1.6-flash

Seed 1.6 Flash is an ultra-fast multimodal deep thinking model by ByteDance Seed, supporting both text and visual understanding. It features a 256k context window and can generate outputs of...

T Weight access not listed
Context
262.144K
Input
$0.075/M
Output
$0.3/M
ByteDance Seed: Seed 1.6by Bytedance Seed
Undated
bytedance-seed/seed-1.6

Seed 1.6 is a general-purpose model released by the ByteDance Seed team. It incorporates multimodal capabilities and adaptive deep thinking with a 256K context window.

T Weight access not listed
Context
262.144K
Input
$0.25/M
Output
$2/M
Undated
minimax/minimax-m2.1

MiniMax-M2.1 is a lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world...

T Weight access not listed
Context
204.8K
Input
$0.3/M
Output
$1.2/M
Undated
z-ai/glm-4.7

GLM-4.7 is Z.ai’s latest flagship model, featuring upgrades in two key areas: enhanced programming capabilities and more stable multi-step reasoning/execution. It demonstrates significant improvements in executing complex agent tasks while...

T Weight access not listed
Context
202.752K
Input
$0.4/M
Output
$1.75/M
nvidia/nemotron-3-nano-30b-a3b:free

NVIDIA Nemotron 3 Nano 30B A3B is a small language MoE model with highest compute efficiency and accuracy for developers to build specialized agentic AI systems. The model is fully...

T Weight access not listed
Context
256K
Input
Free
Output
Free
Undated
openai/gpt-5.2-chat

GPT-5.2 Chat (AKA Instant) is the fast, lightweight member of the 5.2 family, optimized for low-latency chat while retaining strong general intelligence. It uses adaptive reasoning to selectively “think” on...

T Weight access not listed
Context
128K
Input
$1.75/M
Output
$14/M
Undated
mistralai/devstral-2512

Devstral 2 is a state-of-the-art open-source model by Mistral AI specializing in agentic coding. It is a 123B-parameter dense transformer model supporting a 256K context window. Devstral 2 supports exploring...

T Weight access not listed
Context
262.144K
Input
$0.4/M
Output
$2/M
Undated
relace/relace-search

The relace-search model uses 4-12 `view_file` and `grep` tools in parallel to explore a codebase and return relevant files to the user request. In contrast to RAG, relace-search performs agentic...

T Weight access not listed
Context
256K
Input
$1/M
Output
$3/M
Undated
z-ai/glm-4.6v

GLM-4.6V is a large multimodal model designed for high-fidelity visual understanding and long-context reasoning across images, documents, and mixed media. It supports up to 128K tokens, processes complex page layouts...

T Weight access not listed
Context
131.072K
Input
$0.3/M
Output
$0.9/M
Body Builder (beta)by Openrouter
Undated
openrouter/bodybuilder

Transform your natural language requests into structured OpenRouter API request objects. Describe what you want to accomplish with AI models, and Body Builder will construct the appropriate API calls. Example:...

T Weight access not listed
Context
128K
Input
-
Output
-
Undated
amazon/nova-2-lite-v1

Nova 2 Lite is a fast, cost-effective reasoning model for everyday workloads that can process text, images, and videos to generate text. Nova 2 Lite demonstrates standout capabilities in processing...

T Weight access not listed
Context
1M
Input
$0.3/M
Output
$2.5/M
Undated
mistralai/ministral-14b-2512

The largest model in the Ministral 3 family, Ministral 3 14B offers frontier capabilities and performance comparable to its larger Mistral Small 3.2 24B counterpart. A powerful and efficient language...

T Weight access not listed
Context
262.144K
Input
$0.2/M
Output
$0.2/M
Undated
mistralai/ministral-8b-2512

A balanced model in the Ministral 3 family, Ministral 3 8B is a powerful, efficient tiny language model with vision capabilities.

T Weight access not listed
Context
262.144K
Input
$0.15/M
Output
$0.15/M
Undated
mistralai/ministral-3b-2512

The smallest model in the Ministral 3 family, Ministral 3 3B is a powerful, efficient tiny language model with vision capabilities.

T Weight access not listed
Context
131.072K
Input
$0.1/M
Output
$0.1/M
Undated
mistralai/mistral-large-2512

Mistral Large 3 2512 is Mistral’s most capable model to date, featuring a sparse mixture-of-experts architecture with 41B active parameters (675B total), and released under the Apache 2.0 license.

T Weight access not listed
Context
262.144K
Input
$0.5/M
Output
$1.5/M