Side-by-side model data
3 of 6 selected
Compare models
Align pricing, capabilities, performance, and published benchmark results without hiding missing data.
Model comparison matrix
| Metric |
openai/gpt-5.6-sol
|
anthropic/claude-sonnet-5
|
google/gemini-3.1-pro-preview
|
|---|---|---|---|
| Overview | |||
| Creator | OpenAI | Anthropic | |
| Context window | 1.05M | 1M | 1.04858M |
| Input modalities | |||
| Output modalities | |||
| Released | 2026-07-09 | 2026-06-30 | 2026-02-19 |
| Knowledge cutoff | 2026-02-16 | 2026-01-31 | 2025-01 |
| Open source | Not reported | Not reported | Not reported |
| Verification | Source-linkedMedium confidence | Source-linkedMedium confidence | Source-linkedMedium confidence |
| Inference providers | OpenAIabacusai-routeraihubmixamazon-bedrockazureazure-cognitive-servicescloudflare-ai-gatewaycrossmodeldatabricksgithub-copilotkenarillmgatewaymerge-gatewayofoxopencodeopenrouterrouting-runsnowflake-cortexvenicevercelxpersonazenmux | Anthropicabacusaihubmixamazon-bedrockazurecloudflare-ai-gatewaycrossmodelgithub-copilotgoogle-vertexgoogle-vertex-anthropickenarillmgatewaymerge-gatewaymodel-oracle-aiofoxopencodeopenrouterunoroutervenicevercelzenmux | Googleabacusaihubmixaurikocrossmodeldaoxefastrouterfrogbotgithub-copilotgoogle-vertexkilollmgatewaymerge-gatewaynano-gptofoxopencodeopenrouterorcarouterperplexity-agentpioneersnowflake-cortexvercelvivgridzenmux |
| Pricing | |||
| Input | $5 / 1M tokens | $2 / 1M tokens Lowest | $2 / 1M tokens Lowest |
| Output | $30 / 1M tokens | $10 / 1M tokens Lowest | $12 / 1M tokens |
| Cache read | $0.5 / 1M tokens | $0.2 / 1M tokens | $0.2 / 1M tokens |
| Cache write | $6.25 / 1M tokens | $2.5 / 1M tokens | Not listed |
| Capabilities | |||
| Reasoning | Yes | Yes | Yes |
| Tool use | Yes | Yes | Yes |
| Structured output | Yes | Not reported | Yes |
| Attachments | Yes | Yes | Yes |
| Vision | Yes | Yes | Yes |
| Provider offers | 23listed endpoints | 21listed endpoints | 24listed endpoints |
| Performance | |||
| Maximum output | 128K | 128K | 65.536K |
| Throughput | Not reported | Not reported | Not reported |
| Benchmarks | |||
| ARC-AGI-2accuracy - Unversioned | Not available | Not available | 77.1 |
| Agentic Indexindex - Unversioned | 54.0 | 46.7 | 21.4 |
| Agents' Last Examscore - Unversioned | 52.7 | Not available | Not available |
| Artificial Analysis Coding Agent Indexaverage pass@1 - Unversioned - Gemini CLI | Not available | Not available | 43.0 |
| Artificial Analysis Coding Agent Indexindex score - 1.1 - Codex | 80.0 | Not available | Not available |
| Artificial Analysis Intelligence Indexindex score - 4.1 | 58.9 | Not available | Not available |
| BrowseCompaccuracy - Unversioned | 90.4 | 84.7 | Not available |
| CharXiv Reasoningaccuracy - Unversioned | Not available | Not available | 83.3 |
| Coding Indexindex - Unversioned | 77.4 | 71.5 | 68.8 |
| DeepSWEresolve rate - 1.1 | 72.7 | Not available | Not available |
| Design Arena: 3delo - Unversioned | Not available | 1312.0 | 1290.0 |
| Design Arena: agenticgamedevelo - Unversioned | Not available | 1238.0 | 1123.0 |
| Design Arena: agentichtmlslideselo - Unversioned | Not available | Not available | 1226.0 |
| Design Arena: agenticslideselo - Unversioned | Not available | Not available | 1112.0 |
| Design Arena: agenticslides(html)elo - Unversioned | Not available | Not available | 1219.0 |
| Design Arena: agenticslides(python-pptx)elo - Unversioned | Not available | Not available | 1107.0 |
| Design Arena: androidnativeelo - Unversioned | Not available | 1232.0 | 1032.0 |
| Design Arena: asciiartelo - Unversioned | Not available | 1240.0 | 1309.0 |
| Design Arena: codecategorieselo - Unversioned | Not available | 1303.0 | 1272.0 |
| Design Arena: datavizelo - Unversioned | Not available | 1267.0 | 1255.0 |
| Design Arena: fullstackelo - Unversioned | Not available | 1269.0 | 1108.0 |
| Design Arena: gamedevelo - Unversioned | Not available | 1352.0 | 1255.0 |
| Design Arena: godotgamedevelo - Unversioned | Not available | 1269.0 | 1236.0 |
| Design Arena: htmlslideselo - Unversioned | Not available | 1234.0 | 1204.0 |
| Design Arena: mobileappselo - Unversioned | Not available | 1230.0 | 1163.0 |
| Design Arena: pptxslideselo - Unversioned | Not available | Not available | 1110.0 |
| Design Arena: python-pptxslideselo - Unversioned | Not available | 1246.0 | 1109.0 |
| Design Arena: svgelo - Unversioned | Not available | 1240.0 | 1335.0 |
| Design Arena: uicomponentelo - Unversioned | Not available | 1308.0 | 1303.0 |
| Design Arena: webappselo - Unversioned | Not available | 1304.0 | 1178.0 |
| Design Arena: websiteelo - Unversioned | Not available | 1300.0 | 1276.0 |
| FrontierCodepass rate - v1 | Not available | 38.8 | Not available |
| FrontierMathaccuracy - v2 | 89.0 | Not available | Not available |
| GDPval-AAElo - Unversioned | Not available | Not available | 1314.0 |
| GPQA Diamondaccuracy - Unversioned | 94.6 | Not available | 94.3 |
| Humanity's Last Examaccuracy - Unversioned | Not available | Not available | 44.4 |
| Intelligence Indexindex - Unversioned | 58.9 | 53.4 | 46.5 |
| MCP Atlassuccess rate - Unversioned | Not available | Not available | 78.2 |
| MMMU Proaccuracy - Unversioned | 83.0 | Not available | 80.5 |
| OSWorldsuccess rate - 2.0 | 62.6 | Not available | Not available |
| OSWorld-Verifiedsuccess rate - Unversioned | Not available | 81.2 | 76.2 |
| SWE-Atlas Codebase QnApass@1 - Unversioned - Gemini CLI | Not available | Not available | 45.6 |
| SWE-Atlas Codebase QnAscore - Unversioned - Mini-SWE-Agent | Not available | Not available | 13.5 |
| SWE-Atlas Refactoringscore - Unversioned - Gemini CLI | Not available | Not available | 33.8 |
| SWE-Atlas Test Writingscore - Unversioned - Mini-SWE-Agent | Not available | Not available | 29.8 |
| SWE-Bench Multilingualresolve rate - Unversioned | Not available | 78.3 | Not available |
| SWE-Bench Proresolve rate - Unversioned | 64.6 | 63.2 | 54.2 |
| SWE-Bench Propass@1 - Unversioned - Gemini CLI | Not available | Not available | 15.1 |
| SWE-Bench Verifiedresolved - Unversioned | Not available | 85.2 | Not available |
| Terminal-Benchsuccess rate - 2.1 | 88.8 | Not available | Not available |
| Terminal-Benchpass@1 - 2.1 - Gemini CLI | Not available | Not available | 68.3 |
| Terminal-Benchsuccess rate - 2.1 - Terminus-2 | Not available | 80.4 | 70.3 |
| Toolathlonsuccess rate - Unversioned | 58.0 | Not available | Not available |