3 results Clear filters
Ranked by raw score for τ²-Bench Telecom - . Versions are kept separate, and incompatible results are never combined.
GPT-5.5by OpenAI
2026-04-23
openai/gpt-5.5

Default frontier GPT for coding, computer use, research, and knowledge work

T Reasoning Tools Structured output Weight access not listed
Context
1.05M
Input
$5/M
Output
$30/M
Grok 4.3by xAI
2026-04-17
xai/grok-4.3

xAI's default Grok for chat, coding, agentic tools, and lower hallucination risk

T Reasoning Tools Structured output Weight access not listed
Context
1M
Input
$1.25/M
Output
$2.5/M
2026-06-09
cohere/north-mini-code-1-0

Cohere coding model for practical software engineering and agentic edits

T Reasoning Tools Structured output Open weights
Context
256K
Input
-
Output
-