26 results Clear filters
Ranked by raw score for Design Arena: python-pptxslides - . Versions are kept separate, and incompatible results are never combined.
Kimi K3by Moonshot AI
2026-07-16
moonshotai/kimi-k3

Multimodal Kimi model with 1M context and toggleable max-effort thinking for long-horizon agent work

T Reasoning Tools Structured output Open weights
Context
1.04858M
Input
$3/M
Output
$15/M
Undated
anthropic/claude-opus-4.7

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

T Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
anthropic/claude-opus-4.7:batch

Opus 4.7 is the next generation of Anthropic's Opus family, built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on...

T Weight access not listed
Context
1M
Input
$2.5/M
Output
$12.5/M
Claude Fable 5by Anthropic
2026-06-09
anthropic/claude-fable-5

Claude model for creative writing, analysis, and controlled agent workflows

T Reasoning Tools Weight access not listed
Context
1M
Input
$10/M
Output
$50/M
Undated
anthropic/claude-fable-5:batch

Claude Fable 5 is a Mythos-class model from Anthropic, built for autonomous knowledge work and coding. It supports text, image, and file inputs with text output, with reasoning support and...

T Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
Undated
anthropic/claude-opus-4.8

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

T Weight access not listed
Context
1M
Input
$5/M
Output
$25/M
anthropic/claude-opus-4.8:batch

Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family. It supports text, image, and file inputs with text output, with reasoning support and a 1M-token...

T Weight access not listed
Context
1M
Input
$2.5/M
Output
$12.5/M
Undated
z-ai/glm-5.1

GLM-5.1 delivers a major leap in coding capability, with particularly significant gains in handling long-horizon tasks. Unlike previous models built around minute-level interactions, GLM-5.1 can work independently and continuously on...

T Weight access not listed
Context
200K
Input
$0.966/M
Output
$3.036/M
2026-05-19
google/gemini-3.5-flash

Fast Gemini model balancing multimodal reasoning, tool use, and cost

T Reasoning Tools Structured output Weight access not listed
Context
1.04858M
Input
$1.5/M
Output
$9/M
google/gemini-3.5-flash:batch

Gemini 3.5 Flash is Google's high-efficiency multimodal model, bringing near-Pro level coding and reasoning at Flash-tier cost and speed. It is highly optimized for coding proficiency and parallel agentic execution...

T Weight access not listed
Context
1.04858M
Input
$0.75/M
Output
$4.5/M
Claude Sonnet 5by Anthropic
2026-06-30
anthropic/claude-sonnet-5

Everyday Claude agent model for coding, planning, browsing, and general work

T Reasoning Tools Weight access not listed
Context
1M
Input
$2/M
Output
$10/M
anthropic/claude-sonnet-5:batch

Sonnet 5 is Anthropic's most capable Sonnet-class model, with frontier performance across coding, agents, and professional work. It supports adaptive thinking with selectable reasoning effort levels (low, medium, high, max,...

T Weight access not listed
Context
1M
Input
$1/M
Output
$5/M
Qwen: Qwen3.7 Maxby Alibaba Qwen
Undated
qwen/qwen3.7-max

Qwen3.7-Max is the flagship model in Alibaba's Qwen3.7 series. It supports text input and output and is designed for agent-centric workloads, with particular strengths in coding, office and productivity tasks,...

T Weight access not listed
Context
1M
Input
$1.475/M
Output
$4.425/M
Undated
x-ai/grok-4.5

Grok 4.5 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

T Weight access not listed
Context
500K
Input
$2/M
Output
$6/M
Undated
z-ai/glm-5.2

GLM 5.2 is a large-scale reasoning model from Z.ai. It supports text input and output with a 1M-token context window, and is suited for long-horizon agent workflows, project-level software engineering,...

T Weight access not listed
Context
1.024M
Input
$0.966/M
Output
$3.036/M
Kimi K2.6by Moonshot AI
2026-04-21
moonshotai/kimi-k2.6

Multimodal Kimi workhorse for agent loops, coding tasks, and visual context

T Reasoning Tools Structured output Open weights
Context
262.144K
Input
$0.95/M
Output
$4/M
2026-04-08
meta/muse-spark-1.1

Muse Spark is a natively multimodal reasoning model with support for tool-use, visual chain of thought, and multi-agent orchestration.

T Reasoning Tools Structured output Weight access not listed
Context
1M
Input
$1.25/M
Output
$4.25/M
Undated
z-ai/glm-5v-turbo

GLM-5V-Turbo is Z.ai’s first native multimodal agent foundation model, built for vision-based coding and agent-driven tasks. It natively handles image, video, and text inputs, excels at long-horizon planning, complex coding,...

T Weight access not listed
Context
202.752K
Input
$1.2/M
Output
$4/M
GPT-5.5by OpenAI
2026-04-23
openai/gpt-5.5

Default frontier GPT for coding, computer use, research, and knowledge work

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$5/M
Output
$30/M
Undated
openai/gpt-5.5:batch

GPT-5.5 is OpenAI’s frontier model designed for complex professional workloads, building on GPT-5.4 with stronger reasoning, higher reliability, and improved token efficiency on hard tasks. It features a 1M+ token...

T Weight access not listed
Context
1.05M
Input
$2.5/M
Output
$15/M
Kimi K2.7 Codeby Moonshot AI
2026-06-12
moonshotai/kimi-k2.7-code

Coding-focused Kimi model, stronger on long-horizon repo work with less overthinking

T Reasoning Tools Open weights
Context
262.144K
Input
$0.95/M
Output
$4/M
2026-02-19
google/gemini-3.1-pro-preview

Reasoning-first Gemini preview for agentic coding and complex problem solving

T Reasoning Tools Structured output Weight access not listed
Context
1.04858M
Input
$2/M
Output
$12/M
google/gemini-3.1-pro-preview:batch

Gemini 3.1 Pro Preview is Google’s frontier reasoning model, delivering enhanced software engineering performance, improved agentic reliability, and more efficient token usage across complex workflows. Building on the multimodal foundation...

T Weight access not listed
Context
1.04858M
Input
$1/M
Output
$6/M
Undated
x-ai/grok-4.3

Grok 4.3 is a reasoning model from xAI. It accepts text and image inputs with text output, and is suited for agentic workflows, instruction-following tasks, and applications requiring high factual...

T Weight access not listed
Context
1M
Input
$1.25/M
Output
$2.5/M
2025-12-17
google/gemini-3-flash-preview

New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs

T Reasoning Tools Structured output Weight access not listed
Context
1.04858M
Input
$0.5/M
Output
$3/M
google/gemini-3-flash-preview:batch

Gemini 3 Flash Preview is a high speed, high value thinking model designed for agentic workflows, multi turn chat, and coding assistance. It delivers near Pro level reasoning and tool...

T Weight access not listed
Context
1.04858M
Input
$0.25/M
Output
$1.5/M