← Back to dashboard

Best free LLM for code review (vibe-audit)

Code analysis model for the Vibe Audit SaaS product

Current pick: qwen/qwen3-coder:free

Ranking Criterion

Free on OpenRouter + highest HumanEval/code benchmark score

Rankings

# Item Score Reasoning Access
1 apodex/apodex-1.1-mini:free 99.7 codingindex:100 ctx:99 Free on OpenRouter · no retention
2 thinkingmachines/inkling-small:free 93.9 codingindex:93 ze:swe_bench_verified:91 ctx:100 Free on OpenRouter · no retention
3 thinkingmachines/inkling:free 93.9 codingindex:93 ze:swe_bench_verified:91 ctx:100 Free on OpenRouter · no retention
4 inclusionai/ling-3.0-flash-sante:free 91.7 codingindex:89 ctx:99 Free on OpenRouter · no retention
5 poolside/laguna-xs-2.1:free 90.3 ze:swe_bench_verified:83 ctx:99 Free on OpenRouter · no retention
6 nvidia/nemotron-3-ultra-550b-a55b:free 88.5 codingindex:86 ze:swe_bench_verified:83 ctx:100 Free on OpenRouter · no retention
7 cohere/north-mini-code:free 75.4 codingindex:64 ze:swe_bench_verified:80 ctx:99 Free on OpenRouter · no retention
8 nvidia/nemotron-3-super-120b-a12b:free 72.3 codingindex:66 ze:swe_bench_verified:63 ctx:99 Free on OpenRouter · no retention
9 inclusionai/ling-3.1-flash 71.0 gen*0.8:60 ctx:99 Free on OpenRouter · no retention
10 nvidia/nemotron-3.5-lightning:free 61.8 codingindex:47 ze:swe_bench_verified:61 ctx:100 Free on OpenRouter · no retention

Free OpenRouter Models (Latest Fetch)

Model ID Name Context
thinkingmachines/inkling-small:free Thinking Machines: Inkling Small (free) 1,048,576
thinkingmachines/inkling:free Thinking Machines: Inkling (free) 1,048,576
google/lyria-3-pro-preview Google: Lyria 3 Pro Preview 1,048,576
google/lyria-3-clip-preview Google: Lyria 3 Clip Preview 1,048,576
nvidia/nemotron-3.5-lightning:free NVIDIA: Nemotron 3.5 Lightning (free) 1,000,000
nvidia/nemotron-3-ultra-550b-a55b:free NVIDIA: Nemotron 3 Ultra (free) 1,000,000
dots-studio/dots-3-note-preview:free Dots Studio: Dots3-Note Preview (free) 512,000
inclusionai/ling-3.1-flash inclusionAI: Ling 3.1 Flash 262,144
apodex/apodex-1.1-mini:free Apodex: Apodex 1.1 Mini (free) 262,144
inclusionai/ling-3.0-flash-sante:free inclusionAI: Ling 3.0 Flash Sante (free) 262,144
poolside/laguna-s-2.1:free Poolside: Laguna S 2.1 (free) 262,144
poolside/laguna-xs-2.1:free Poolside: Laguna XS 2.1 (free) 262,144
google/gemma-4-26b-a4b-it:free Google: Gemma 4 26B A4B (free) 262,144
google/gemma-4-31b-it:free Google: Gemma 4 31B (free) 262,144
nvidia/nemotron-3-super-120b-a12b:free NVIDIA: Nemotron 3 Super (free) 262,144
cohere/north-mini-code:free Cohere: North Mini Code (free) 256,000
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free NVIDIA: Nemotron 3 Nano Omni (free) 256,000
openrouter/free Free Models Router 200,000
nvidia/nemotron-3.5-content-safety:free NVIDIA: Nemotron 3.5 Content Safety (free) 128,000
liquid/lfm-2.5-2.6b:free LiquidAI: LFM2.5-2.6B (free) 65,536

AI Methodology Reviews

Weekly question: Is HumanEval still representative of real-world code review quality? Any new code benchmarks?

Needs Attention 2026-10-04 via nvidia/nemotron-3-super-120b-a12b:free

HumanEval measures functional correctness on isolated snippets, which does not capture the nuanced aspects of code review such as style, security, maintainability, or contextual understanding. Recent benchmarks like SWE‑bench (verified), CodeReviewBench, and MBPP‑Eval better reflect real‑world review scenarios and should be incorporated or weighted alongside HumanEval.

Needs Attention 2026-09-27 via nvidia/nemotron-3-super-120b-a12b:free

HumanEval measures the ability to generate correct code from docstrings, which does not directly assess code‑review skills such as bug detection, style feedback, or security analysis. Recent benchmarks like SWE‑Bench, CodeReviewBench, and MBPP‑Review better capture real‑world review quality, so relying solely on HumanEval may overlook models that excel at review‑specific tasks.

Needs Attention 2026-09-20 via nvidia/nemotron-3-super-120b-a12b:free

HumanEval primarily tests the ability to generate correct, short code snippets from docstrings, which does not capture the nuanced skills required for code review such as detecting bugs, security issues, style violations, or understanding larger context. Recent benchmarks like SWE‑bench (especially the verified split), CodeReviewBench, and MBPP‑Eval better reflect real‑world review scenarios, so relying solely on HumanEval may overlook important aspects of review quality.

Needs Attention 2026-09-13 via nvidia/nemotron-3-super-120b-a12b:free

HumanEval primarily measures functional correctness of short code snippets, which does not directly reflect a model's ability to perform code review tasks such as identifying bugs, suggesting improvements, or understanding context. Recent benchmarks like SWE‑bench verified, CodeReviewBench, and MBPP‑review focus more on realistic code review scenarios and should be considered alongside or instead of HumanEval for ranking LLMs for code review.

Needs Attention 2026-09-06 via nvidia/nemotron-3-super-120b-a12b:free

The current ranking relies on HumanEval as the primary code benchmark, but recent evidence suggests HumanEval does not fully capture real‑world code review quality (e.g., it focuses on isolated function completion rather than multi‑file changes, bug detection, or style adherence). Newer benchmarks such as SWE‑bench, HumanEval+, and CodeReviewBench have emerged to better reflect practical review scenarios, and the provided top‑ranked models are evaluated using codingindex and SWE‑bench_verified rather than HumanEval, indicating a mismatch between the stated criterion and the actual data used.

Needs Attention 2026-08-30 via nvidia/nemotron-3-super-120b-a12b:free

HumanEval measures only isolated function completion and does not capture the multi‑step reasoning, bug detection, or patch generation required for real‑world code review. Recent benchmarks such as SWE‑bench (verified), CodeReviewBench, and MBPP‑Eval better reflect review‑oriented capabilities, so relying solely on HumanEval is outdated.

Needs Attention 2026-08-23 via nvidia/nemotron-3-super-120b-a12b:free

HumanEval primarily measures functional code generation correctness, which does not directly assess the nuanced abilities needed for code review such as detecting bugs, suggesting improvements, or understanding context. Recent benchmarks like SWE‑bench verified, CodeReviewBench, and HumanEval+ (or MBPP for review‑oriented tasks) better capture real‑world review quality, so relying solely on HumanEval is insufficient.

Needs Attention 2026-08-16 via none

All models failed — review manually

Needs Attention 2026-08-09 via nvidia/nemotron-3-super-120b-a12b:free

HumanEval measures functional correctness of isolated snippets, which does not directly assess the nuanced abilities needed for code review such as understanding context, detecting bugs, or suggesting improvements. Recent benchmarks like SWE‑bench, LiveCodeBench, and coding‑index evaluations better reflect real‑world review performance and show different model rankings, indicating the current criterion may be outdated.

Needs Attention 2026-08-02 via nvidia/nemotron-3-super-120b-a12b:free

HumanEval primarily measures functional correctness on isolated problems and does not capture nuances of code review such as style, security, or maintainability. Newer benchmarks like SWE‑bench verified and LiveCodeBench better reflect real‑world code review quality, and the current top models according to those metrics differ from the selected qwen/qwen3-coder:free.