← Back to dashboard

Best local LLM for general reasoning

Local model for OpenClaw agent reasoning (ollama / llama.cpp)

Current pick: Qwen 3.5 27B

Ranking Criterion

Reasoning benchmarks + fits consumer GPU VRAM

Rankings

# Item Score Reasoning Access
1 QwQ-32B 69.5 ze:gpqa:65 elo:72 Free from Huggingface · no retention
2 Gemma-3-27B-it 61.7 ze:gpqa:32 elo:76 Free from Huggingface · no retention
3 Nemotron-3-Nano-30B-A3B 60.6 intelligenceindex:37 mmlupro:94 ze:gpqa:79 Free from Huggingface · no retention
4 Gemma-3-12B-it 58.5 ze:gpqa:30 elo:73 Free from Huggingface · no retention
5 Qwen3.5-27B 56.5 intelligenceindex:42 ze:gpqa:94 Free from Huggingface · no retention

Free OpenRouter Models (Latest Fetch)

Model ID Name Context
thinkingmachines/inkling-small:free Thinking Machines: Inkling Small (free) 1,048,576
thinkingmachines/inkling:free Thinking Machines: Inkling (free) 1,048,576
google/lyria-3-pro-preview Google: Lyria 3 Pro Preview 1,048,576
google/lyria-3-clip-preview Google: Lyria 3 Clip Preview 1,048,576
nvidia/nemotron-3.5-lightning:free NVIDIA: Nemotron 3.5 Lightning (free) 1,000,000
nvidia/nemotron-3-ultra-550b-a55b:free NVIDIA: Nemotron 3 Ultra (free) 1,000,000
dots-studio/dots-3-note-preview:free Dots Studio: Dots3-Note Preview (free) 512,000
inclusionai/ling-3.1-flash inclusionAI: Ling 3.1 Flash 262,144
apodex/apodex-1.1-mini:free Apodex: Apodex 1.1 Mini (free) 262,144
inclusionai/ling-3.0-flash-sante:free inclusionAI: Ling 3.0 Flash Sante (free) 262,144
poolside/laguna-s-2.1:free Poolside: Laguna S 2.1 (free) 262,144
poolside/laguna-xs-2.1:free Poolside: Laguna XS 2.1 (free) 262,144
google/gemma-4-26b-a4b-it:free Google: Gemma 4 26B A4B (free) 262,144
google/gemma-4-31b-it:free Google: Gemma 4 31B (free) 262,144
nvidia/nemotron-3-super-120b-a12b:free NVIDIA: Nemotron 3 Super (free) 262,144
cohere/north-mini-code:free Cohere: North Mini Code (free) 256,000
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free NVIDIA: Nemotron 3 Nano Omni (free) 256,000
openrouter/free Free Models Router 200,000
nvidia/nemotron-3.5-content-safety:free NVIDIA: Nemotron 3.5 Content Safety (free) 128,000
liquid/lfm-2.5-2.6b:free LiquidAI: LFM2.5-2.6B (free) 65,536

AI Methodology Reviews

Weekly question: Any new small models that outperform current picks on reasoning tasks?

Needs Attention 2026-10-04 via nvidia/nemotron-3-super-120b-a12b:free

The ranking shows QwQ-32B scoring 69.5, which exceeds the current pick Qwen 3.5 27B (56.5) on reasoning benchmarks. However, QwQ-32B’s 32B size likely exceeds typical consumer GPU VRAM limits, indicating the methodology may need to re‑evaluate the VRAM constraint or update the top pick.

Needs Attention 2026-09-27 via nvidia/nemotron-3-super-120b-a12b:free

The ranking shows QwQ-32B with a higher reasoning score (69.5) than the current pick Qwen 3.5 27B (56.5), indicating a newer small model surpasses the existing choice on reasoning benchmarks. Unless QwQ-32B exceeds consumer GPU VRAM limits, the methodology should be revisited to incorporate this new model.

Needs Attention 2026-09-20 via nvidia/nemotron-3-super-120b-a12b:free

Gemma-3-27B-it achieves a higher reasoning score (61.7) than the current pick Qwen 3.5 27B (56.5) while having the same 27B parameter size, thus fitting consumer GPU VRAM. This indicates the ranking methodology should be updated to reflect the better-performing model.

Needs Attention 2026-09-13 via nvidia/nemotron-3-super-120b-a12b:free

The methodology (reasoning benchmarks + consumer GPU VRAM fit) remains valid, but the current pick (Qwen 3.5 27B) is now outperformed by newer models such as QwQ-32B, which scores higher on reasoning benchmarks while still being feasible for consumer GPUs with quantization. The ranking should be updated to reflect this new top performer.

Valid 2026-09-06 via nvidia/nemotron-3-super-120b-a12b:free

The ranking methodology prioritizes both reasoning performance and fit within consumer GPU VRAM limits. While newer models like QwQ-32B achieve higher reasoning scores, their 32B size likely exceeds typical consumer VRAM (e.g., 24 GB on an RTX 4090) without aggressive quantization, so they do not satisfy the criterion. Thus the current pick (Qwen 3.5 27B) remains the best trade‑off, and the methodology is still sound.

Valid 2026-08-30 via nvidia/nemotron-3-super-120b-a12b:free

The current top-ranked model Qwen3.5-27B leads the list with a score of 71.7; the next best, Nemotron-3-Nano-30B-A3B, scores 71.1 and is slightly larger, while all other listed models fall below. No new small model exceeds the current pick on reasoning benchmarks, so the criterion remains appropriate.

Valid 2026-08-23 via nvidia/nemotron-3-super-120b-a12b:free

The current top-ranked model, Qwen 3.5 27B, maintains the highest reasoning score (71.7) among the listed options, and no newer small model exceeds it on the reasoning benchmarks while fitting consumer GPU VRAM limits. Therefore, the existing methodology remains appropriate for ranking the best local LLM for general reasoning.

Needs Attention 2026-08-16 via none

All models failed — review manually

Valid 2026-08-09 via nvidia/nemotron-3-super-120b-a12b:free

The provided rankings show Qwen3.5-27B leading with a score of 71.7, and no newer small model is listed that exceeds this score on the reasoning benchmarks used. Therefore, the existing criterion (reasoning performance + consumer‑GPU VRAM fit) remains appropriate and the methodology is still sound.

Valid 2026-08-02 via nvidia/nemotron-3-super-120b-a12b:free

The higher‑scoring models (Nemotron‑3‑Nano‑30B‑A3B and QwQ‑32B) likely exceed typical consumer GPU VRAM limits, so they do not satisfy the 'fits consumer GPU VRAM' part of the criterion. Since no new small model is shown to both outperform Qwen 3.5 27B on reasoning benchmarks and fit within consumer VRAM, the existing methodology remains appropriate.