Code analysis model for the Vibe Audit SaaS product
Free on OpenRouter + highest HumanEval/code benchmark score
| # | Item | Score | Reasoning | Access |
|---|---|---|---|---|
| 1 | apodex/apodex-1.1-mini:free | 99.7 | codingindex:100 ctx:99 | Free on OpenRouter · no retention |
| 2 | thinkingmachines/inkling-small:free | 93.9 | codingindex:93 ze:swe_bench_verified:91 ctx:100 | Free on OpenRouter · no retention |
| 3 | thinkingmachines/inkling:free | 93.9 | codingindex:93 ze:swe_bench_verified:91 ctx:100 | Free on OpenRouter · no retention |
| 4 | inclusionai/ling-3.0-flash-sante:free | 91.7 | codingindex:89 ctx:99 | Free on OpenRouter · no retention |
| 5 | poolside/laguna-xs-2.1:free | 90.3 | ze:swe_bench_verified:83 ctx:99 | Free on OpenRouter · no retention |
| 6 | nvidia/nemotron-3-ultra-550b-a55b:free | 88.5 | codingindex:86 ze:swe_bench_verified:83 ctx:100 | Free on OpenRouter · no retention |
| 7 | cohere/north-mini-code:free | 75.4 | codingindex:64 ze:swe_bench_verified:80 ctx:99 | Free on OpenRouter · no retention |
| 8 | nvidia/nemotron-3-super-120b-a12b:free | 72.3 | codingindex:66 ze:swe_bench_verified:63 ctx:99 | Free on OpenRouter · no retention |
| 9 | inclusionai/ling-3.1-flash | 71.0 | gen*0.8:60 ctx:99 | Free on OpenRouter · no retention |
| 10 | nvidia/nemotron-3.5-lightning:free | 61.8 | codingindex:47 ze:swe_bench_verified:61 ctx:100 | Free on OpenRouter · no retention |
| Model ID | Name | Context |
|---|---|---|
| thinkingmachines/inkling-small:free | Thinking Machines: Inkling Small (free) | 1,048,576 |
| thinkingmachines/inkling:free | Thinking Machines: Inkling (free) | 1,048,576 |
| google/lyria-3-pro-preview | Google: Lyria 3 Pro Preview | 1,048,576 |
| google/lyria-3-clip-preview | Google: Lyria 3 Clip Preview | 1,048,576 |
| nvidia/nemotron-3.5-lightning:free | NVIDIA: Nemotron 3.5 Lightning (free) | 1,000,000 |
| nvidia/nemotron-3-ultra-550b-a55b:free | NVIDIA: Nemotron 3 Ultra (free) | 1,000,000 |
| dots-studio/dots-3-note-preview:free | Dots Studio: Dots3-Note Preview (free) | 512,000 |
| inclusionai/ling-3.1-flash | inclusionAI: Ling 3.1 Flash | 262,144 |
| apodex/apodex-1.1-mini:free | Apodex: Apodex 1.1 Mini (free) | 262,144 |
| inclusionai/ling-3.0-flash-sante:free | inclusionAI: Ling 3.0 Flash Sante (free) | 262,144 |
| poolside/laguna-s-2.1:free | Poolside: Laguna S 2.1 (free) | 262,144 |
| poolside/laguna-xs-2.1:free | Poolside: Laguna XS 2.1 (free) | 262,144 |
| google/gemma-4-26b-a4b-it:free | Google: Gemma 4 26B A4B (free) | 262,144 |
| google/gemma-4-31b-it:free | Google: Gemma 4 31B (free) | 262,144 |
| nvidia/nemotron-3-super-120b-a12b:free | NVIDIA: Nemotron 3 Super (free) | 262,144 |
| cohere/north-mini-code:free | Cohere: North Mini Code (free) | 256,000 |
| nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | NVIDIA: Nemotron 3 Nano Omni (free) | 256,000 |
| openrouter/free | Free Models Router | 200,000 |
| nvidia/nemotron-3.5-content-safety:free | NVIDIA: Nemotron 3.5 Content Safety (free) | 128,000 |
| liquid/lfm-2.5-2.6b:free | LiquidAI: LFM2.5-2.6B (free) | 65,536 |
Weekly question: Is HumanEval still representative of real-world code review quality? Any new code benchmarks?
HumanEval measures functional correctness on isolated snippets, which does not capture the nuanced aspects of code review such as style, security, maintainability, or contextual understanding. Recent benchmarks like SWE‑bench (verified), CodeReviewBench, and MBPP‑Eval better reflect real‑world review scenarios and should be incorporated or weighted alongside HumanEval.
HumanEval measures the ability to generate correct code from docstrings, which does not directly assess code‑review skills such as bug detection, style feedback, or security analysis. Recent benchmarks like SWE‑Bench, CodeReviewBench, and MBPP‑Review better capture real‑world review quality, so relying solely on HumanEval may overlook models that excel at review‑specific tasks.
HumanEval primarily tests the ability to generate correct, short code snippets from docstrings, which does not capture the nuanced skills required for code review such as detecting bugs, security issues, style violations, or understanding larger context. Recent benchmarks like SWE‑bench (especially the verified split), CodeReviewBench, and MBPP‑Eval better reflect real‑world review scenarios, so relying solely on HumanEval may overlook important aspects of review quality.
HumanEval primarily measures functional correctness of short code snippets, which does not directly reflect a model's ability to perform code review tasks such as identifying bugs, suggesting improvements, or understanding context. Recent benchmarks like SWE‑bench verified, CodeReviewBench, and MBPP‑review focus more on realistic code review scenarios and should be considered alongside or instead of HumanEval for ranking LLMs for code review.
The current ranking relies on HumanEval as the primary code benchmark, but recent evidence suggests HumanEval does not fully capture real‑world code review quality (e.g., it focuses on isolated function completion rather than multi‑file changes, bug detection, or style adherence). Newer benchmarks such as SWE‑bench, HumanEval+, and CodeReviewBench have emerged to better reflect practical review scenarios, and the provided top‑ranked models are evaluated using codingindex and SWE‑bench_verified rather than HumanEval, indicating a mismatch between the stated criterion and the actual data used.
HumanEval measures only isolated function completion and does not capture the multi‑step reasoning, bug detection, or patch generation required for real‑world code review. Recent benchmarks such as SWE‑bench (verified), CodeReviewBench, and MBPP‑Eval better reflect review‑oriented capabilities, so relying solely on HumanEval is outdated.
HumanEval primarily measures functional code generation correctness, which does not directly assess the nuanced abilities needed for code review such as detecting bugs, suggesting improvements, or understanding context. Recent benchmarks like SWE‑bench verified, CodeReviewBench, and HumanEval+ (or MBPP for review‑oriented tasks) better capture real‑world review quality, so relying solely on HumanEval is insufficient.
All models failed — review manually
HumanEval measures functional correctness of isolated snippets, which does not directly assess the nuanced abilities needed for code review such as understanding context, detecting bugs, or suggesting improvements. Recent benchmarks like SWE‑bench, LiveCodeBench, and coding‑index evaluations better reflect real‑world review performance and show different model rankings, indicating the current criterion may be outdated.
HumanEval primarily measures functional correctness on isolated problems and does not capture nuances of code review such as style, security, or maintainability. Newer benchmarks like SWE‑bench verified and LiveCodeBench better reflect real‑world code review quality, and the current top models according to those metrics differ from the selected qwen/qwen3-coder:free.