← Back to dashboard

Best free LLM for OpenClaw (general reasoning)

General-purpose reasoning model for the OpenClaw agent swarm

Current pick: nvidia/nemotron-3-super-120b-a12b:free

Ranking Criterion

Free on OpenRouter + highest MMLU/reasoning score on Open LLM Leaderboard

Rankings

# Item Score Reasoning Access
1 google/gemma-4-31b-it:free 94.9 ze:gpqa:92 ctx:99 Free on OpenRouter · no retention
2 google/gemma-4-26b-a4b-it:free 93.3 ze:gpqa:89 ctx:99 Free on OpenRouter · no retention
3 inclusionai/ling-3.1-flash 80.3 intelligenceindex:75 ctx:99 Free on OpenRouter · no retention
4 nvidia/nemotron-3-ultra-550b-a55b:free 64.7 intelligenceindex:42 ze:gpqa:96 ctx:100 Free on OpenRouter · no retention
5 apodex/apodex-1.1-mini:free 59.8 intelligenceindex:48 ctx:99 Free on OpenRouter · no retention
6 thinkingmachines/inkling-small:free 59.0 intelligenceindex:47 ctx:100 Free on OpenRouter · no retention
7 thinkingmachines/inkling:free 59.0 intelligenceindex:47 ctx:100 Free on OpenRouter · no retention
8 nvidia/nemotron-3-super-120b-a12b:free 52.2 intelligenceindex:23 ze:gpqa:90 ctx:99 Free on OpenRouter · no retention
9 inclusionai/ling-3.0-flash-sante:free 50.9 intelligenceindex:37 ctx:99 Free on OpenRouter · no retention
10 nvidia/nemotron-3.5-lightning:free 50.1 intelligenceindex:23 ze:gpqa:79 ctx:100 Free on OpenRouter · no retention

Free OpenRouter Models (Latest Fetch)

Model ID Name Context
thinkingmachines/inkling-small:free Thinking Machines: Inkling Small (free) 1,048,576
thinkingmachines/inkling:free Thinking Machines: Inkling (free) 1,048,576
google/lyria-3-pro-preview Google: Lyria 3 Pro Preview 1,048,576
google/lyria-3-clip-preview Google: Lyria 3 Clip Preview 1,048,576
nvidia/nemotron-3.5-lightning:free NVIDIA: Nemotron 3.5 Lightning (free) 1,000,000
nvidia/nemotron-3-ultra-550b-a55b:free NVIDIA: Nemotron 3 Ultra (free) 1,000,000
dots-studio/dots-3-note-preview:free Dots Studio: Dots3-Note Preview (free) 512,000
inclusionai/ling-3.1-flash inclusionAI: Ling 3.1 Flash 262,144
apodex/apodex-1.1-mini:free Apodex: Apodex 1.1 Mini (free) 262,144
inclusionai/ling-3.0-flash-sante:free inclusionAI: Ling 3.0 Flash Sante (free) 262,144
poolside/laguna-s-2.1:free Poolside: Laguna S 2.1 (free) 262,144
poolside/laguna-xs-2.1:free Poolside: Laguna XS 2.1 (free) 262,144
google/gemma-4-26b-a4b-it:free Google: Gemma 4 26B A4B (free) 262,144
google/gemma-4-31b-it:free Google: Gemma 4 31B (free) 262,144
nvidia/nemotron-3-super-120b-a12b:free NVIDIA: Nemotron 3 Super (free) 262,144
cohere/north-mini-code:free Cohere: North Mini Code (free) 256,000
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free NVIDIA: Nemotron 3 Nano Omni (free) 256,000
openrouter/free Free Models Router 200,000
nvidia/nemotron-3.5-content-safety:free NVIDIA: Nemotron 3.5 Content Safety (free) 128,000
liquid/lfm-2.5-2.6b:free LiquidAI: LFM2.5-2.6B (free) 65,536

AI Methodology Reviews

Weekly question: Is MMLU still the best proxy for general agent reasoning? Any new benchmarks worth switching to?

Needs Attention 2026-10-04 via nvidia/nemotron-3-super-120b-a12b:free

While MMLU remains a useful measure of broad knowledge, it does not fully capture the multi‑step reasoning and tool‑use abilities required for general agent tasks. Newer benchmarks such as GPQA, ARC‑Challenge, and agent‑specific suites (e.g., AgentBench, WebArena) provide a more nuanced view of reasoning performance and should be considered alongside or instead of MMLU for ranking free LLMs.

Needs Attention 2026-09-27 via nvidia/nemotron-3-super-120b-a12b:free

MMLU primarily measures factual knowledge recall and does not fully capture the multi-step, tool‑using, and planning abilities that define general agent reasoning. Recent benchmarks such as GPQA, MMLU‑Pro, AgentBench, and ToolQA better assess reasoning, instruction following, and tool use, so the ranking criterion should be updated to incorporate one or more of these newer metrics.

Needs Attention 2026-09-20 via nvidia/nemotron-3-super-120b-a12b:free

MMLU primarily measures factual knowledge and may not fully capture the complex reasoning abilities needed for general agent tasks. Emerging benchmarks such as GPQA, IntelligenceIndex, and other reasoning‑focused suites show stronger correlation with real‑world reasoning performance and should be considered as replacements or supplements to MMLU in the ranking criterion.

Needs Attention 2026-09-13 via nvidia/nemotron-3-super-120b-a12b:free

While MMLU remains a useful general‑knowledge proxy, it does not fully capture the multi‑step reasoning and tool‑use abilities needed for agentic tasks. Recent leaderboards now prioritize benchmarks such as GPQA, MATH, HumanEval, and Big‑Bench Hard, which better reflect general agent reasoning, so the ranking criterion should be expanded or replaced with a more comprehensive metric.

Needs Attention 2026-09-06 via nvidia/nemotron-3-super-120b-a12b:free

While MMLU remains a solid proxy for general knowledge, it does not fully capture the multi-step reasoning, tool use, and planning abilities required for effective agents. Emerging benchmarks such as GPQA, MMLU‑Pro, AgentBench, and tool‑use focused suites (e.g., ToolBench, WebArena) better reflect general agent reasoning and should be considered for updating the ranking criterion.

Needs Attention 2026-08-30 via nvidia/nemotron-3-super-120b-a12b:free

MMLU primarily measures factual knowledge recall rather than the multi‑step, tool‑using, and planning abilities that define general agent reasoning. Recent leaderboards emphasize benchmarks like GPQA, ARC‑Challenge, BIG‑Bench‑Hard, and agent‑specific suites (e.g., AgentBench, ToolFormer) that better capture reasoning and interaction skills, so the ranking criterion should be expanded or replaced with these newer proxies.

Needs Attention 2026-08-23 via nvidia/nemotron-3-super-120b-a12b:free

MMLU primarily measures factual knowledge recall and does not fully capture the multi-step, tool-use, and planning abilities needed for general agent reasoning. Emerging benchmarks such as GPQA, IntelligenceIndex, and contextual reasoning (ctx) scores better reflect agentic performance, as evidenced by the current top rankings favoring models excelling on those metrics over pure MMLU scores.

Needs Attention 2026-08-16 via none

All models failed — review manually

Needs Attention 2026-08-09 via nvidia/nemotron-3-super-120b-a12b:free

MMLU primarily measures factual knowledge recall and does not capture complex reasoning, tool use, or long-context abilities that are critical for general agent performance. Recent leaderboards emphasize benchmarks like GPQA, MMLU-Pro, and IntelligenceIndex, which better reflect reasoning and instruction-following capabilities, suggesting the ranking criterion should be updated to incorporate these newer metrics.

Needs Attention 2026-08-02 via nvidia/nemotron-3-super-120b-a12b:free

MMLU alone is a narrow proxy for general agent reasoning and has been shown to correlate poorly with performance on more complex, knowledge‑intensive tasks. Recent leaderboards emphasize benchmarks like GPQA, MMLU‑Pro, and Intelligence Index, which better capture reasoning and contextual understanding, suggesting the ranking criterion should be updated to incorporate these metrics.