494 results Clear filters
Qwen3.6 Flashby Alibaba Qwen
2026-04-27
alibaba/qwen3.6-flash

Qwen vision-language model for visual reasoning, documents, and agent tasks

T Reasoning Tools Structured output Weight access not listed
Context
1M
Input
$0.188/M
Output
$1.125/M
DeepSeek V4 Proby DeepSeek
2026-04-24
deepseek/deepseek-v4-pro

Open MoE flagship with million-token context for coding and long agent runs

T Reasoning Tools Structured output Open weights
Context
1M
Input
$0.435/M
Output
$0.87/M
2026-04-24
deepseek/deepseek-v4-flash

Fast DeepSeek V4 lane for economical reasoning, coding, and long-context work

T Reasoning Tools Structured output Open weights
Context
1M
Input
$0.14/M
Output
$0.28/M
GPT-5.5 Proby OpenAI
2026-04-23
openai/gpt-5.5-pro

Highest-accuracy GPT-5.5 tier for slower, precision-heavy reasoning and coding

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$30/M
Output
$180/M
GPT-5.5by OpenAI
2026-04-23
openai/gpt-5.5

Default frontier GPT for coding, computer use, research, and knowledge work

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$5/M
Output
$30/M
MiMo-V2.5-Proby Xiaomi
2026-04-22
xiaomi/mimo-v2.5-pro

Stronger MiMo Pro tier for multimodal reasoning and coding-agent execution

T Reasoning Tools Open weights
Context
1.04858M
Input
$0.435/M
Output
$0.87/M
MiMo-V2.5by Xiaomi
2026-04-22
xiaomi/mimo-v2.5

Open MiMo model for multimodal coding agents and long-context automation

T Reasoning Tools Open weights
Context
1.04858M
Input
$0.14/M
Output
$0.28/M
Qwen3.6 27Bby Alibaba Qwen
2026-04-22
alibaba/qwen3.6-27b

Qwen vision-language model for visual reasoning, documents, and agent tasks

T Reasoning Tools Structured output Open weights
Context
262.144K
Input
$0.6/M
Output
$3.6/M
Kimi K2.6by Moonshot AI
2026-04-21
moonshotai/kimi-k2.6

Multimodal Kimi workhorse for agent loops, coding tasks, and visual context

T Reasoning Tools Structured output Open weights
Context
262.144K
Input
$0.95/M
Output
$4/M
2026-04-21
google/deep-research-preview-04-2026

Agentic model for autonomous multi-step research, synthesis, and cited reports

T Weight access not listed
Context
1.04858M
Input
-
Output
-
2026-04-21
google/deep-research-max-preview-04-2026

Maximum-comprehensiveness agentic researcher for multi-step investigation, synthesis, and cited reports

T Weight access not listed
Context
1.04858M
Input
-
Output
-
Hy3 previewby Tencent
2026-04-20
tencent/hy3-preview

Tencent Hy reasoning model for coding, instruction following, and agent tasks

T Reasoning Tools Open weights
Context
256K
Input
$0.063/M
Output
$0.21/M
Qwen3.6 Max Previewby Alibaba Qwen
2026-04-20
alibaba/qwen3.6-max-preview

Flagship Qwen model for complex reasoning, coding, and agentic workflows

T Reasoning Tools Weight access not listed
Context
262.144K
Input
$1.3/M
Output
$7.8/M
Grok 4.3by xAI
2026-04-17
xai/grok-4.3

xAI's default Grok for chat, coding, agentic tools, and lower hallucination risk

T Reasoning Tools Structured output Weight access not listed
Context
1M
Input
$1.25/M
Output
$2.5/M
Qwen3.6 35B-A3Bby Alibaba Qwen
2026-04-17
alibaba/qwen3.6-35b-a3b

Open multimodal Qwen MoE for local agents that need vision, audio, and code

T Reasoning Tools Structured output Open weights
Context
262.144K
Input
$0.248/M
Output
$1.485/M
2026-04-16
xai/grok-build-0.1

Fast Grok coding model tuned for agentic engineering and iterative edits

T Reasoning Tools Structured output Weight access not listed
Context
256K
Input
$1/M
Output
$2/M
2026-04-16
nvidia/nemotron-3-content-safety

Safety model for policy screening, moderation, and risk-aware routing workflows

T Open weights
Context
128K
Input
-
Output
-
Claude Opus 4.7by Anthropic
2026-04-16
anthropic/claude-opus-4-7

Stronger Opus tier for advanced software work and high-stakes reasoning

T Reasoning Tools Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
2026-04-14
google/gemini-robotics-er-1.6-preview

Vision-language model for embodied reasoning: spatial understanding, task planning, and physical-world agentic robotics

T Reasoning Tools Structured output Weight access not listed
Context
131.072K
Input
$1/M
Output
$5/M
2026-04-08
meta/muse-spark-1.1

Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.

T Reasoning Tools Structured output Weight access not listed
Context
1M
Input
$1.25/M
Output
$4.25/M
GLM-5.1by Zhipu AI
2026-04-07
zhipuai/glm-5.1

Strong GLM coding model for agentic engineering, terminals, and repository generation

T Reasoning Tools Structured output Open weights
Context
200K
Input
$1.4/M
Output
$4.4/M
2026-04-02
google/gemma-4-26b-a4b-it

Open Gemma instruction model for efficient chat and self-hosted deployments

T Reasoning Tools Structured output Open weights
Context
262.144K
Input
$0.06/M
Output
$0.33/M
2026-04-02
google/gemma-4-31b-it

Largest Gemma 4 instruction model for open, self-hosted chat and reasoning

T Reasoning Tools Open weights
Context
262.144K
Input
$0.1/M
Output
$0.35/M
2026-04-02
google/gemma-4-E4B-it

Open Gemma instruction model for efficient chat and self-hosted deployments

T Reasoning Tools Structured output Open weights
Context
131.072K
Input
$0.2/M
Output
$0.2/M
2026-04-02
google/gemma-4-E2B-it

Open Gemma instruction model for efficient chat and self-hosted deployments

T Reasoning Tools Structured output Open weights
Context
131.072K
Input
$0.1/M
Output
$0.1/M
2026-04-02
stepfun/step-3.5-flash-2603

StepFun flash model for efficient multimodal reasoning, coding, and tool use

T Reasoning Tools Open weights
Context
256K
Input
$0.1/M
Output
$0.3/M
Qwen3.6 Plusby Alibaba Qwen
2026-04-02
alibaba/qwen3.6-plus

Earlier Qwen multimodal workhorse for million-token agent and document tasks

T Reasoning Tools Weight access not listed
Context
1M
Input
$0.5/M
Output
$3/M
GLM-5V-Turboby Zhipu AI
2026-04-01
zhipuai/glm-5v-turbo

Fast GLM vision model for screenshots, documents, and multimodal agent tasks

T Reasoning Tools Weight access not listed
Context
200K
Input
$5/M
Output
$22/M
2026-03-31
nvidia/llama-nemotron-rerank-vl-1b-v2

Reranking model for improving retrieval quality in search and recommendation systems

T Open weights
Context
128K
Input
-
Output
-
2026-03-26
google/gemini-3.1-flash-live-preview

High-quality, low-latency Live API model for real-time dialogue and voice-first AI applications

T Reasoning Tools Weight access not listed
Context
131.072K
Input
$0.75/M
Output
$4.5/M
2026-03-25
google/lyria-3-pro-preview

Music generation model for full-length songs from text or images with vocals and structure

T Weight access not listed
Context
131.072K
Input
-
Output
-
2026-03-25
google/lyria-3-clip-preview

Music generation model for short 30-second clips, loops, and previews from text or image prompts

T Weight access not listed
Context
131.072K
Input
-
Output
-
2026-03-24
nvidia/nemotron-cascade-2-30b-a3b

Nemotron model for efficient reasoning, coding, and specialized AI agents

T Open weights
Context
256K
Input
-
Output
-
MiMo-V2-Omniby Xiaomi
2026-03-18
xiaomi/mimo-v2-omni

MiMo omni model for text, image, video, audio, and agents

T Reasoning Tools Weight access not listed
Context
262.144K
Input
$0.14/M
Output
$0.28/M
MiMo-V2-Proby Xiaomi
2026-03-18
xiaomi/mimo-v2-pro

Earlier MiMo Pro model for multimodal agents, reasoning, and code tasks

T Reasoning Tools Structured output Weight access not listed
Context
1.04858M
Input
$0.435/M
Output
$0.87/M
2026-03-18
minimax/MiniMax-M2.7-highspeed

Low-latency M2.7 variant for interactive coding plans and agent loops

T Reasoning Tools Open weights
Context
204.8K
Input
$0.6/M
Output
$2.4/M
MiniMax-M2.7by MiniMax
2026-03-18
minimax/MiniMax-M2.7

Open MiniMax flagship for coding agents, office automation, and complex environments

T Reasoning Tools Open weights
Context
204.8K
Input
$0.3/M
Output
$1.2/M
GPT-5.4 nanoby OpenAI
2026-03-17
openai/gpt-5.4-nano

Cheapest GPT-5.4 lane for simple routing, extraction, and bulk automation

T Reasoning Tools Structured output Weight access not listed
Context
400K
Input
$0.2/M
Output
$1.25/M
GPT-5.4 miniby OpenAI
2026-03-17
openai/gpt-5.4-mini

Strong small GPT for coding subagents, quick tool use, and high-volume work

T Reasoning Tools Structured output Weight access not listed
Context
400K
Input
$0.75/M
Output
$4.5/M
GLM-5-Turboby Zhipu AI
2026-03-16
zhipuai/glm-5-turbo

Faster GLM-5 lane for coding agents that need lower latency

T Reasoning Tools Structured output Weight access not listed
Context
200K
Input
$0.9/M
Output
$3.7/M
2026-03-16
mistral/mistral-small-latest

Efficient Mistral model for fast chat, extraction, and production assistants

T Reasoning Tools Open weights
Context
256K
Input
$0.15/M
Output
$0.6/M
Mistral Small 4by Mistral AI
2026-03-16
mistral/mistral-small-2603

Fast Mistral production model for chat, extraction, and cost-sensitive agents

T Reasoning Tools Open weights
Context
256K
Input
$0.15/M
Output
$0.6/M
2026-03-16
nvidia/nemotron-voicechat

Nemotron multimodal model for visual reasoning and agentic AI workflows

T Tools Open weights
Context
128K
Input
-
Output
-
2026-03-11
nvidia/nemotron-3-super-120b-a12b

Nemotron middle tier for collaborative agents and high-volume reasoning workloads

T Reasoning Tools Open weights
Context
262.144K
Input
$0.2/M
Output
$0.8/M
2026-03-09
xai/grok-4.20-0309-non-reasoning

Grok model for agentic tool use, reasoning, coding, and live assistance

T Tools Structured output Weight access not listed
Context
1M
Input
$1.25/M
Output
$2.5/M
2026-03-09
xai/grok-4.20-0309-reasoning

Reasoning Grok for document-heavy analysis and long-horizon tool use

T Reasoning Tools Structured output Weight access not listed
Context
1M
Input
$1.25/M
Output
$2.5/M
GPT-5.4 Proby OpenAI
2026-03-05
openai/gpt-5.4-pro

More exact GPT-5.4 tier for demanding professional reasoning and agent tasks

T Reasoning Tools Weight access not listed
Context
1.05M
Input
$30/M
Output
$180/M
GPT-5.4by OpenAI
2026-03-05
openai/gpt-5.4

Agent-ready GPT for coding and computer-use workflows at a lower cost

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$2.5/M
Output
$15/M
2026-03-03
google/gemini-3.1-flash-lite-preview

Low-latency Gemini model for high-volume multimodal and agent workloads

T Reasoning Tools Structured output Weight access not listed
Context
1.04858M
Input
$0.25/M
Output
$1.5/M
2026-03-03
openai/gpt-5.3-chat-latest

Chat-tuned GPT model for conversational assistance, writing, and tool workflows

T Tools Structured output Weight access not listed
Context
128K
Input
$1.75/M
Output
$14/M