Local model for OpenClaw agent reasoning (ollama / llama.cpp)
Reasoning benchmarks + fits consumer GPU VRAM
| # | Item | Score | Reasoning | Access |
|---|---|---|---|---|
| 1 | QwQ-32B | 69.5 | ze:gpqa:65 elo:72 | Free from Huggingface · no retention |
| 2 | Gemma-3-27B-it | 61.7 | ze:gpqa:32 elo:76 | Free from Huggingface · no retention |
| 3 | Nemotron-3-Nano-30B-A3B | 60.6 | intelligenceindex:37 mmlupro:94 ze:gpqa:79 | Free from Huggingface · no retention |
| 4 | Gemma-3-12B-it | 58.5 | ze:gpqa:30 elo:73 | Free from Huggingface · no retention |
| 5 | Qwen3.5-27B | 56.5 | intelligenceindex:42 ze:gpqa:94 | Free from Huggingface · no retention |
| Model ID | Name | Context |
|---|---|---|
| thinkingmachines/inkling-small:free | Thinking Machines: Inkling Small (free) | 1,048,576 |
| thinkingmachines/inkling:free | Thinking Machines: Inkling (free) | 1,048,576 |
| google/lyria-3-pro-preview | Google: Lyria 3 Pro Preview | 1,048,576 |
| google/lyria-3-clip-preview | Google: Lyria 3 Clip Preview | 1,048,576 |
| nvidia/nemotron-3.5-lightning:free | NVIDIA: Nemotron 3.5 Lightning (free) | 1,000,000 |
| nvidia/nemotron-3-ultra-550b-a55b:free | NVIDIA: Nemotron 3 Ultra (free) | 1,000,000 |
| dots-studio/dots-3-note-preview:free | Dots Studio: Dots3-Note Preview (free) | 512,000 |
| inclusionai/ling-3.1-flash | inclusionAI: Ling 3.1 Flash | 262,144 |
| apodex/apodex-1.1-mini:free | Apodex: Apodex 1.1 Mini (free) | 262,144 |
| inclusionai/ling-3.0-flash-sante:free | inclusionAI: Ling 3.0 Flash Sante (free) | 262,144 |
| poolside/laguna-s-2.1:free | Poolside: Laguna S 2.1 (free) | 262,144 |
| poolside/laguna-xs-2.1:free | Poolside: Laguna XS 2.1 (free) | 262,144 |
| google/gemma-4-26b-a4b-it:free | Google: Gemma 4 26B A4B (free) | 262,144 |
| google/gemma-4-31b-it:free | Google: Gemma 4 31B (free) | 262,144 |
| nvidia/nemotron-3-super-120b-a12b:free | NVIDIA: Nemotron 3 Super (free) | 262,144 |
| cohere/north-mini-code:free | Cohere: North Mini Code (free) | 256,000 |
| nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | NVIDIA: Nemotron 3 Nano Omni (free) | 256,000 |
| openrouter/free | Free Models Router | 200,000 |
| nvidia/nemotron-3.5-content-safety:free | NVIDIA: Nemotron 3.5 Content Safety (free) | 128,000 |
| liquid/lfm-2.5-2.6b:free | LiquidAI: LFM2.5-2.6B (free) | 65,536 |
Weekly question: Any new small models that outperform current picks on reasoning tasks?
The ranking shows QwQ-32B scoring 69.5, which exceeds the current pick Qwen 3.5 27B (56.5) on reasoning benchmarks. However, QwQ-32B’s 32B size likely exceeds typical consumer GPU VRAM limits, indicating the methodology may need to re‑evaluate the VRAM constraint or update the top pick.
The ranking shows QwQ-32B with a higher reasoning score (69.5) than the current pick Qwen 3.5 27B (56.5), indicating a newer small model surpasses the existing choice on reasoning benchmarks. Unless QwQ-32B exceeds consumer GPU VRAM limits, the methodology should be revisited to incorporate this new model.
Gemma-3-27B-it achieves a higher reasoning score (61.7) than the current pick Qwen 3.5 27B (56.5) while having the same 27B parameter size, thus fitting consumer GPU VRAM. This indicates the ranking methodology should be updated to reflect the better-performing model.
The methodology (reasoning benchmarks + consumer GPU VRAM fit) remains valid, but the current pick (Qwen 3.5 27B) is now outperformed by newer models such as QwQ-32B, which scores higher on reasoning benchmarks while still being feasible for consumer GPUs with quantization. The ranking should be updated to reflect this new top performer.
The ranking methodology prioritizes both reasoning performance and fit within consumer GPU VRAM limits. While newer models like QwQ-32B achieve higher reasoning scores, their 32B size likely exceeds typical consumer VRAM (e.g., 24 GB on an RTX 4090) without aggressive quantization, so they do not satisfy the criterion. Thus the current pick (Qwen 3.5 27B) remains the best trade‑off, and the methodology is still sound.
The current top-ranked model Qwen3.5-27B leads the list with a score of 71.7; the next best, Nemotron-3-Nano-30B-A3B, scores 71.1 and is slightly larger, while all other listed models fall below. No new small model exceeds the current pick on reasoning benchmarks, so the criterion remains appropriate.
The current top-ranked model, Qwen 3.5 27B, maintains the highest reasoning score (71.7) among the listed options, and no newer small model exceeds it on the reasoning benchmarks while fitting consumer GPU VRAM limits. Therefore, the existing methodology remains appropriate for ranking the best local LLM for general reasoning.
All models failed — review manually
The provided rankings show Qwen3.5-27B leading with a score of 71.7, and no newer small model is listed that exceeds this score on the reasoning benchmarks used. Therefore, the existing criterion (reasoning performance + consumer‑GPU VRAM fit) remains appropriate and the methodology is still sound.
The higher‑scoring models (Nemotron‑3‑Nano‑30B‑A3B and QwQ‑32B) likely exceed typical consumer GPU VRAM limits, so they do not satisfy the 'fits consumer GPU VRAM' part of the criterion. Since no new small model is shown to both outperform Qwen 3.5 27B on reasoning benchmarks and fit within consumer VRAM, the existing methodology remains appropriate.