Thinkingmachines logo

Inkling

Multimodal MoE reasoning model (975B total, 41B active) for text, image, and audio

Source-linked ling Open weights Released 2026-07-15
API data Report
Input modalitiesT
Output modalitiesT
Input price$1.87/1M tokens
Output price$4.68/1M tokens
Context window1.04858M
Max output1.04858M
Overview

Capabilities and specifications

Reasoning Yes
Tool calling Yes
Structured output Yes
Attachments Yes
Vision input Yes
Open weights Yes
Model family
ling
Knowledge cutoff
Not documented
License
Apache-2.0
Release date
2026-07-15
Model ID
thinkingmachines/inkling
Inference availability

Provider pricing and availability

Provider-specific identifiers, limits, and prices per million tokens.

Report a price
Providers offering Inkling
ProviderProvider model IDContextInputOutputCache readCapabilitiesSource
Nvidia thinkingmachines/inkling 1.04858M Not listed Not listed - ReasoningTools Docs
Baseten thinkingmachines/inkling 1.04858M $1 $4.05 - ReasoningToolsJSON Docs
Vercel AI Gateway thinkingmachines/inkling 256K $1 $4.05 $0.17 ReasoningTools Docs
OpenRouter thinkingmachines/inkling 1.04858M $1 $4.05 $0.17 ReasoningTools Docs
Together AI thinkingmachines/Inkling 524.288K $1 $4.05 $0.17 ReasoningToolsJSON Docs
Hugging Face thinkingmachines/Inkling 1.04858M $1 $4.05 - ReasoningToolsJSON Docs
Modal thinkingmachines/Inkling-NVFP4 1.04858M $1.2 $5 $0.27 ReasoningTools Docs
Venice AI inkling 1M $1.25 $5.062 $0.212 ReasoningToolsJSON Docs
Thinking Machines thinkingmachines/Inkling 65.536K $1.87 $4.68 $0.374 ReasoningTools Docs
Thinking Machines inkling 256K $3.74 $9.36 $0.748 ReasoningTools Docs
Thinking Machines thinkingmachines/Inkling:peft:262144 262.144K $3.74 $9.36 $0.748 ReasoningTools Docs

Capability badges appear only when the provider catalog explicitly lists support.

Serving quality

Performance and reliability

Observed serving measurements stay separate from model specifications and are never estimated from price or model size.

Throughput Not reported

Generation speed; higher is faster.

Latency / TTFT Not reported

Time until the first output token arrives.

Uptime Not reported

Successful requests over a measured period.

Performance varies by inference provider, region, and load. Provider-level measurements appear only when a source reports them.

Published evaluations

Benchmarks

Raw results remain attached to their original source and version.

Benchmark registry
Agentic Index - Unversioned - indexTop 8 reported models
Claude Opus 5
55.3
GPT-5.6 Sol
54.0
Anthropic: Claude Fable 5 (batch)
52.8
Claude Fable 5
52.8
Kimi K3
50.1
GPT-5.6 Terra
47.4
Anthropic: Claude Opus 4.8 (batch)
47.2
Inkling
32.3
Coding Index - Unversioned - indexTop 8 reported models
Claude Opus 5
78.0
GPT-5.6 Sol
77.4
GPT-5.6 Terra
76.7
Anthropic: Claude Fable 5 (batch)
76.5
Claude Fable 5
76.5
Kimi K3
76.2
OpenAI: GPT-5.5 (batch)
74.9
Inkling
52.1
Intelligence Index - Unversioned - indexTop 8 reported models
Claude Opus 5
60.7
Anthropic: Claude Fable 5 (batch)
59.9
Claude Fable 5
59.9
GPT-5.6 Sol
58.9
Kimi K3
57.1
Anthropic: Claude Opus 4.8 (batch)
55.7
Anthropic: Claude Opus 4.8
55.7
Inkling
40.7
Catalog activity

Recent record changes

Detected changes from successful source imports.

Full change log
Context Length524288 → 1048576
Price Completion4.05 → 4.68
Price Prompt1.0 → 1.87
Context Length1048576 → 524288
Price Completion4.68 → 4.05
Price Prompt1.87 → 1.0
Context Length524288 → 1048576
Price Completion4.05 → 4.68
Public API

Use this record

Fetch the complete source-linked model record without an API key.

API documentation
Endpoint
GET https://www.modelbench.lol/api/v1/models/thinkingmachines/inkling
curl
curl "https://www.modelbench.lol/api/v1/models/thinkingmachines/inkling"