2 results Clear filters
Ranked by raw score for IFBench - . Versions are kept separate, and incompatible results are never combined.
2026-06-04
nvidia/nemotron-3-ultra-550b-a55b

Largest Nemotron 3 model for maximum open-weight reasoning and agent accuracy

T Reasoning Tools Structured output Open weights
Context
1M
Input
$0.5/M
Output
$2.5/M
Grok 4.3by xAI
2026-04-17
xai/grok-4.3

xAI's default Grok for chat, coding, agentic tools, and lower hallucination risk

T Reasoning Tools Structured output Weight access not listed
Context
1M
Input
$1.25/M
Output
$2.5/M