494 results Clear filters
Undated
qwen/qwen3-235b-a22b-thinking-2507

Qwen3-235B-A22B-Thinking-2507 is a high-performance, open-weight Mixture-of-Experts (MoE) language model optimized for complex reasoning tasks. It activates 22B of its 235B parameters per forward pass and natively supports up to 262,144...

T Weight access not listed
Context
131.072K
Input
$0.23/M
Output
$2.3/M
Undated
qwen/qwen3-coder:free

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...

T Weight access not listed
Context
262K
Input
Free
Output
Free
Undated
qwen/qwen3-coder

Qwen3-Coder-480B-A35B-Instruct is a Mixture-of-Experts (MoE) code generation model developed by the Qwen team. It is optimized for agentic coding tasks such as function calling, tool use, and long-context reasoning over...

T Weight access not listed
Context
262.144K
Input
$0.3/M
Output
$1/M
Undated
bytedance/ui-tars-1.5-7b

UI-TARS-1.5 is a multimodal vision-language agent optimized for GUI-based environments, including desktop interfaces, web browsers, mobile systems, and games. Built by ByteDance, it builds upon the UI-TARS framework with reinforcement...

T Weight access not listed
Context
128K
Input
$0.1/M
Output
$0.2/M
Undated
qwen/qwen3-235b-a22b-2507

Qwen3-235B-A22B-Instruct-2507 is a multilingual, instruction-tuned mixture-of-experts language model based on the Qwen3-235B architecture, with 22B active parameters per forward pass. It is optimized for general-purpose text generation, including instruction following,...

T Weight access not listed
Context
262.144K
Input
$0.09/M
Output
$0.55/M
Undated
moonshotai/kimi-k2

Kimi K2 Instruct is a large-scale Mixture-of-Experts (MoE) language model developed by Moonshot AI, featuring 1 trillion total parameters with 32 billion active per forward pass. It is optimized for...

T Weight access not listed
Context
131.072K
Input
$0.57/M
Output
$2.3/M
Venice: Uncensoredby Cognitivecomputations
Undated
cognitivecomputations/dolphin-mistral-24b-venice-edition

Venice Uncensored Dolphin Mistral 24B Venice Edition is a fine-tuned variant of Mistral-Small-24B-Instruct-2501, developed by dphn.ai in collaboration with Venice.ai. This model is designed as an “uncensored” instruct-tuned LLM, preserving...

T Weight access not listed
Context
128K
Input
$0.2/M
Output
$0.9/M
Undated
tencent/hunyuan-a13b-instruct

Hunyuan-A13B is a 13B active parameter Mixture-of-Experts (MoE) language model developed by Tencent, with a total parameter count of 80B and support for reasoning via Chain-of-Thought. It offers competitive benchmark...

T Weight access not listed
Context
131.072K
Input
$0.14/M
Output
$0.57/M
Undated
morph/morph-v3-large

Morph's high-accuracy apply model for complex code edits. ~4,500 tokens/sec with 98% accuracy for precise code transformations. The model requires the prompt to be in the following format: <instruction>{instruction}</instruction> <code>{initial_code}</code>...

T Weight access not listed
Context
262.144K
Input
$0.9/M
Output
$1.9/M
Undated
mistralai/mistral-small-3.2-24b-instruct

Mistral-Small-3.2-24B-Instruct-2506 is an updated 24B parameter model from Mistral optimized for instruction following, repetition reduction, and improved function calling. Compared to the 3.1 release, version 3.2 significantly improves accuracy on...

T Weight access not listed
Context
131.072K
Input
$0.1/M
Output
$0.3/M
Undated
minimax/minimax-m1

MiniMax-M1 is a large-scale, open-weight reasoning model designed for extended context and high-efficiency inference. It leverages a hybrid Mixture-of-Experts (MoE) architecture paired with a custom "lightning attention" mechanism, allowing it...

T Weight access not listed
Context
1M
Input
$0.55/M
Output
$2.2/M
google/gemini-2.5-pro-preview

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

T Weight access not listed
Context
1.04858M
Input
$1.25/M
Output
$10/M
Undated
deepseek/deepseek-r1-0528

May 28th update to the [original DeepSeek R1](/deepseek/deepseek-r1) Performance on par with [OpenAI o1](/openai/o1), but open-sourced and with fully open reasoning tokens. It's 671B parameters in size, with 37B active...

T Weight access not listed
Context
163.84K
Input
$0.5/M
Output
$2.15/M
Undated
anthropic/claude-opus-4

Claude Opus 4 is benchmarked as the world’s best coding model, at time of release, bringing sustained performance on complex, long-running tasks and agent workflows. It sets new benchmarks in...

T Weight access not listed
Context
200K
Input
$15/M
Output
$75/M
Undated
anthropic/claude-sonnet-4

Claude Sonnet 4 significantly enhances the capabilities of its predecessor, Sonnet 3.7, excelling in both coding and reasoning tasks with improved precision and controllability. Achieving state-of-the-art performance on SWE-bench (72.7%),...

T Weight access not listed
Context
200K
Input
$3/M
Output
$15/M
Undated
mistralai/mistral-medium-3

Mistral Medium 3 is a high-performance enterprise-grade language model designed to deliver frontier-level capabilities at significantly reduced operational cost. It balances state-of-the-art reasoning and multimodal performance with 8× lower cost...

T Weight access not listed
Context
131.072K
Input
$0.4/M
Output
$2/M
google/gemini-2.5-pro-preview-05-06

Gemini 2.5 Pro is Google’s state-of-the-art AI model designed for advanced reasoning, coding, mathematics, and scientific tasks. It employs “thinking” capabilities, enabling it to reason through responses with enhanced accuracy...

T Weight access not listed
Context
1.04858M
Input
$1.25/M
Output
$10/M
Undated
arcee-ai/virtuoso-large

Virtuoso‑Large is Arcee's top‑tier general‑purpose LLM at 72 B parameters, tuned to tackle cross‑domain reasoning, creative writing and enterprise QA. Unlike many 70 B peers, it retains the 128 k...

T Weight access not listed
Context
131.072K
Input
$0.75/M
Output
$1.2/M
Undated
meta-llama/llama-guard-4-12b

Llama Guard 4 is a Llama 4 Scout-derived multimodal pretrained model, fine-tuned for content safety classification. Similar to previous versions, it can be used to classify content in both LLM...

T Weight access not listed
Context
163.84K
Input
$0.18/M
Output
$0.18/M
Qwen: Qwen3 8Bby Alibaba Qwen
Undated
qwen/qwen3-8b

Qwen3-8B is a dense 8.2B parameter causal language model from the Qwen3 series, designed for both reasoning-heavy tasks and efficient dialogue. It supports seamless switching between "thinking" mode for math,...

T Weight access not listed
Context
131.072K
Input
$0.117/M
Output
$0.455/M
Qwen: Qwen3 14Bby Alibaba Qwen
Undated
qwen/qwen3-14b

Qwen3-14B is a dense 14.8B parameter causal language model from the Qwen3 series, designed for both complex reasoning and efficient dialogue. It supports seamless switching between a "thinking" mode for...

T Weight access not listed
Context
131.072K
Input
$0.227/M
Output
$0.91/M
Qwen: Qwen3 235B A22Bby Alibaba Qwen
Undated
qwen/qwen3-235b-a22b

Qwen3-235B-A22B is a 235B parameter mixture-of-experts (MoE) model developed by Qwen, activating 22B parameters per forward pass. It supports seamless switching between a "thinking" mode for complex reasoning, math, and...

T Weight access not listed
Context
131.072K
Input
$0.455/M
Output
$1.82/M
Undated
openai/o4-mini-high

OpenAI o4-mini-high is the same model as [o4-mini](/openai/o4-mini) with reasoning_effort set to high. OpenAI o4-mini is a compact reasoning model in the o-series, optimized for fast, cost-efficient performance while retaining...

T Weight access not listed
Context
200K
Input
$1.1/M
Output
$4.4/M
Undated
meta-llama/llama-4-maverick

Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...

T Weight access not listed
Context
1.04858M
Input
$0.2/M
Output
$0.8/M
Undated
meta-llama/llama-4-scout

Llama 4 Scout 17B Instruct (16E) is a mixture-of-experts (MoE) language model developed by Meta, activating 17 billion parameters out of a total of 109B. It supports native multimodal input...

T Weight access not listed
Context
327.68K
Input
$0.1/M
Output
$0.3/M
Undated
deepseek/deepseek-chat-v3-0324

DeepSeek V3, a 685B-parameter, mixture-of-experts model, is the latest iteration of the flagship chat model family from the DeepSeek team. It succeeds the [DeepSeek V3](/deepseek/deepseek-chat-v3) model and performs really well...

T Weight access not listed
Context
163.84K
Input
$0.27/M
Output
$1.12/M
Undated
mistralai/mistral-small-3.1-24b-instruct

Mistral Small 3.1 24B Instruct is an upgraded variant of Mistral Small 3 (2501), featuring 24 billion parameters with advanced multimodal capabilities. It provides state-of-the-art performance in text-based reasoning and...

T Weight access not listed
Context
128K
Input
$0.351/M
Output
$0.555/M
Undated
google/gemma-3-4b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...

T Weight access not listed
Context
131.072K
Input
$0.05/M
Output
$0.1/M
Undated
google/gemma-3-12b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...

T Weight access not listed
Context
131.072K
Input
$0.05/M
Output
$0.15/M
Undated
cohere/command-a

Command A is an open-weights 111B parameter model with a 256k context window focused on delivering great performance across agentic, multilingual, and coding use cases. Compared to other leading proprietary...

T Weight access not listed
Context
256K
Input
$2.5/M
Output
$10/M
openai/gpt-4o-mini-search-preview

GPT-4o mini Search Preview is a specialized model for web search in Chat Completions. It is trained to understand and execute web search queries.

T Weight access not listed
Context
128K
Input
$0.15/M
Output
$0.6/M
openai/gpt-4o-search-preview

GPT-4o Search Previewis a specialized model for web search in Chat Completions. It is trained to understand and execute web search queries.

T Weight access not listed
Context
128K
Input
$2.5/M
Output
$10/M
Undated
google/gemma-3-27b-it

Gemma 3 introduces multimodality, supporting vision-language input and text outputs. It handles context windows up to 128k tokens, understands over 140 languages, and offers improved math, reasoning, and chat capabilities,...

T Weight access not listed
Context
131.072K
Input
$0.08/M
Output
$0.45/M
Undated
perplexity/sonar-deep-research

Sonar Deep Research is a research-focused model designed for multi-step retrieval, synthesis, and reasoning across complex topics. It autonomously searches, reads, and evaluates sources, refining its approach as it gathers...

T Weight access not listed
Context
128K
Input
$2/M
Output
$8/M
Undated
openai/o3-mini-high

OpenAI o3-mini-high is the same model as [o3-mini](/openai/o3-mini) with reasoning_effort set to high. o3-mini is a cost-efficient language model optimized for STEM reasoning tasks, particularly excelling in science, mathematics, and...

T Weight access not listed
Context
200K
Input
$1.1/M
Output
$4.4/M
Undated
qwen/qwen2.5-vl-72b-instruct

Qwen2.5-VL is proficient in recognizing common objects such as flowers, birds, fish, and insects. It is also highly capable of analyzing texts, charts, icons, graphics, and layouts within images.

T Weight access not listed
Context
128K
Input
$0.8/M
Output
$1/M
Qwen: Qwen-Plusby Alibaba Qwen
Undated
qwen/qwen-plus

Qwen-Plus, based on the Qwen2.5 foundation model, is a 131K context model with a balanced performance, speed, and cost combination.

T Weight access not listed
Context
1M
Input
$0.26/M
Output
$0.78/M
Undated
minimax/minimax-01

MiniMax-01 is a combines MiniMax-Text-01 for text generation and MiniMax-VL-01 for image understanding. It has 456 billion parameters, with 45.9 billion parameters activated per inference, and can handle a context...

T Weight access not listed
Context
1.00019M
Input
$0.2/M
Output
$1.1/M
sao10k/l3.3-euryale-70b

Euryale L3.3 70B is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.2](/models/sao10k/l3-euryale-70b).

T Weight access not listed
Context
131.072K
Input
$0.65/M
Output
$0.75/M
meta-llama/llama-3.3-70b-instruct

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...

T Weight access not listed
Context
131.072K
Input
$0.13/M
Output
$0.4/M
Undated
amazon/nova-lite-v1

Amazon Nova Lite 1.0 is a very low-cost multimodal model from Amazon that focused on fast processing of image, video, and text inputs to generate text output. Amazon Nova Lite...

T Weight access not listed
Context
300K
Input
$0.06/M
Output
$0.24/M
Undated
amazon/nova-micro-v1

Amazon Nova Micro 1.0 is a text-only model that delivers the lowest latency responses in the Amazon Nova family of models at a very low cost. With a context length...

T Weight access not listed
Context
128K
Input
$0.035/M
Output
$0.14/M
Undated
amazon/nova-pro-v1

Amazon Nova Pro 1.0 is a capable multimodal model from Amazon focused on providing a combination of accuracy, speed, and cost for a wide range of tasks. As of December...

T Weight access not listed
Context
300K
Input
$0.8/M
Output
$3.2/M
Mistral Large 2407by Mistral AI
Undated
mistralai/mistral-large-2407

This is Mistral AI's flagship model, Mistral Large 2 (version mistral-large-2407). It's a proprietary weights-available model and excels at reasoning, code, JSON, chat, and more. Read the launch announcement [here](https://mistral.ai/news/mistral-large-2407/)....

T Weight access not listed
Context
131.072K
Input
$2/M
Output
$6/M
meta-llama/llama-3.2-11b-vision-instruct

Llama 3.2 11B Vision is a multimodal model with 11 billion parameters, designed to handle tasks combining visual and textual data. It excels in tasks such as image captioning and...

T Weight access not listed
Context
131.072K
Input
$0.345/M
Output
$0.345/M
meta-llama/llama-3.2-3b-instruct:free

Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization. Designed with the latest transformer architecture, it...

T Weight access not listed
Context
131.072K
Input
Free
Output
Free
meta-llama/llama-3.2-3b-instruct

Llama 3.2 3B is a 3-billion-parameter multilingual large language model, optimized for advanced natural language processing tasks like dialogue generation, reasoning, and summarization. Designed with the latest transformer architecture, it...

T Weight access not listed
Context
131.072K
Input
$0.05/M
Output
$0.33/M
sao10k/l3.1-euryale-70b

Euryale L3.1 70B v2.2 is a model focused on creative roleplay from [Sao10k](https://ko-fi.com/sao10k). It is the successor of [Euryale L3 70B v2.1](/models/sao10k/l3-euryale-70b).

T Weight access not listed
Context
131.072K
Input
$0.85/M
Output
$0.85/M
Undated
nousresearch/hermes-3-llama-3.1-70b

Hermes 3 is a generalist language model with many improvements over [Hermes 2](/models/nousresearch/nous-hermes-2-mistral-7b-dpo), including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...

T Weight access not listed
Context
131.072K
Input
$0.7/M
Output
$0.7/M
Undated
nousresearch/hermes-3-llama-3.1-405b:free

Hermes 3 is a generalist language model with many improvements over Hermes 2, including advanced agentic capabilities, much better roleplaying, reasoning, multi-turn conversation, long context coherence, and improvements across the...

T Weight access not listed
Context
131.072K
Input
Free
Output
Free