General-purpose reasoning model for the OpenClaw agent swarm
Free on OpenRouter + highest MMLU/reasoning score on Open LLM Leaderboard
| # | Item | Score | Reasoning | Access |
|---|---|---|---|---|
| 1 | google/gemma-4-31b-it:free | 94.9 | ze:gpqa:92 ctx:99 | Free on OpenRouter · no retention |
| 2 | google/gemma-4-26b-a4b-it:free | 93.3 | ze:gpqa:89 ctx:99 | Free on OpenRouter · no retention |
| 3 | inclusionai/ling-3.1-flash | 80.3 | intelligenceindex:75 ctx:99 | Free on OpenRouter · no retention |
| 4 | nvidia/nemotron-3-ultra-550b-a55b:free | 64.7 | intelligenceindex:42 ze:gpqa:96 ctx:100 | Free on OpenRouter · no retention |
| 5 | apodex/apodex-1.1-mini:free | 59.8 | intelligenceindex:48 ctx:99 | Free on OpenRouter · no retention |
| 6 | thinkingmachines/inkling-small:free | 59.0 | intelligenceindex:47 ctx:100 | Free on OpenRouter · no retention |
| 7 | thinkingmachines/inkling:free | 59.0 | intelligenceindex:47 ctx:100 | Free on OpenRouter · no retention |
| 8 | nvidia/nemotron-3-super-120b-a12b:free | 52.2 | intelligenceindex:23 ze:gpqa:90 ctx:99 | Free on OpenRouter · no retention |
| 9 | inclusionai/ling-3.0-flash-sante:free | 50.9 | intelligenceindex:37 ctx:99 | Free on OpenRouter · no retention |
| 10 | nvidia/nemotron-3.5-lightning:free | 50.1 | intelligenceindex:23 ze:gpqa:79 ctx:100 | Free on OpenRouter · no retention |
| Model ID | Name | Context |
|---|---|---|
| thinkingmachines/inkling-small:free | Thinking Machines: Inkling Small (free) | 1,048,576 |
| thinkingmachines/inkling:free | Thinking Machines: Inkling (free) | 1,048,576 |
| google/lyria-3-pro-preview | Google: Lyria 3 Pro Preview | 1,048,576 |
| google/lyria-3-clip-preview | Google: Lyria 3 Clip Preview | 1,048,576 |
| nvidia/nemotron-3.5-lightning:free | NVIDIA: Nemotron 3.5 Lightning (free) | 1,000,000 |
| nvidia/nemotron-3-ultra-550b-a55b:free | NVIDIA: Nemotron 3 Ultra (free) | 1,000,000 |
| dots-studio/dots-3-note-preview:free | Dots Studio: Dots3-Note Preview (free) | 512,000 |
| inclusionai/ling-3.1-flash | inclusionAI: Ling 3.1 Flash | 262,144 |
| apodex/apodex-1.1-mini:free | Apodex: Apodex 1.1 Mini (free) | 262,144 |
| inclusionai/ling-3.0-flash-sante:free | inclusionAI: Ling 3.0 Flash Sante (free) | 262,144 |
| poolside/laguna-s-2.1:free | Poolside: Laguna S 2.1 (free) | 262,144 |
| poolside/laguna-xs-2.1:free | Poolside: Laguna XS 2.1 (free) | 262,144 |
| google/gemma-4-26b-a4b-it:free | Google: Gemma 4 26B A4B (free) | 262,144 |
| google/gemma-4-31b-it:free | Google: Gemma 4 31B (free) | 262,144 |
| nvidia/nemotron-3-super-120b-a12b:free | NVIDIA: Nemotron 3 Super (free) | 262,144 |
| cohere/north-mini-code:free | Cohere: North Mini Code (free) | 256,000 |
| nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | NVIDIA: Nemotron 3 Nano Omni (free) | 256,000 |
| openrouter/free | Free Models Router | 200,000 |
| nvidia/nemotron-3.5-content-safety:free | NVIDIA: Nemotron 3.5 Content Safety (free) | 128,000 |
| liquid/lfm-2.5-2.6b:free | LiquidAI: LFM2.5-2.6B (free) | 65,536 |
Weekly question: Is MMLU still the best proxy for general agent reasoning? Any new benchmarks worth switching to?
While MMLU remains a useful measure of broad knowledge, it does not fully capture the multi‑step reasoning and tool‑use abilities required for general agent tasks. Newer benchmarks such as GPQA, ARC‑Challenge, and agent‑specific suites (e.g., AgentBench, WebArena) provide a more nuanced view of reasoning performance and should be considered alongside or instead of MMLU for ranking free LLMs.
MMLU primarily measures factual knowledge recall and does not fully capture the multi-step, tool‑using, and planning abilities that define general agent reasoning. Recent benchmarks such as GPQA, MMLU‑Pro, AgentBench, and ToolQA better assess reasoning, instruction following, and tool use, so the ranking criterion should be updated to incorporate one or more of these newer metrics.
MMLU primarily measures factual knowledge and may not fully capture the complex reasoning abilities needed for general agent tasks. Emerging benchmarks such as GPQA, IntelligenceIndex, and other reasoning‑focused suites show stronger correlation with real‑world reasoning performance and should be considered as replacements or supplements to MMLU in the ranking criterion.
While MMLU remains a useful general‑knowledge proxy, it does not fully capture the multi‑step reasoning and tool‑use abilities needed for agentic tasks. Recent leaderboards now prioritize benchmarks such as GPQA, MATH, HumanEval, and Big‑Bench Hard, which better reflect general agent reasoning, so the ranking criterion should be expanded or replaced with a more comprehensive metric.
While MMLU remains a solid proxy for general knowledge, it does not fully capture the multi-step reasoning, tool use, and planning abilities required for effective agents. Emerging benchmarks such as GPQA, MMLU‑Pro, AgentBench, and tool‑use focused suites (e.g., ToolBench, WebArena) better reflect general agent reasoning and should be considered for updating the ranking criterion.
MMLU primarily measures factual knowledge recall rather than the multi‑step, tool‑using, and planning abilities that define general agent reasoning. Recent leaderboards emphasize benchmarks like GPQA, ARC‑Challenge, BIG‑Bench‑Hard, and agent‑specific suites (e.g., AgentBench, ToolFormer) that better capture reasoning and interaction skills, so the ranking criterion should be expanded or replaced with these newer proxies.
MMLU primarily measures factual knowledge recall and does not fully capture the multi-step, tool-use, and planning abilities needed for general agent reasoning. Emerging benchmarks such as GPQA, IntelligenceIndex, and contextual reasoning (ctx) scores better reflect agentic performance, as evidenced by the current top rankings favoring models excelling on those metrics over pure MMLU scores.
All models failed — review manually
MMLU primarily measures factual knowledge recall and does not capture complex reasoning, tool use, or long-context abilities that are critical for general agent performance. Recent leaderboards emphasize benchmarks like GPQA, MMLU-Pro, and IntelligenceIndex, which better reflect reasoning and instruction-following capabilities, suggesting the ranking criterion should be updated to incorporate these newer metrics.
MMLU alone is a narrow proxy for general agent reasoning and has been shown to correlate poorly with performance on more complex, knowledge‑intensive tasks. Recent leaderboards emphasize benchmarks like GPQA, MMLU‑Pro, and Intelligence Index, which better capture reasoning and contextual understanding, suggesting the ranking criterion should be updated to incorporate these metrics.