← Back to dashboard

Best free LLM for chat analysis (alphalens)

Summarization and extraction model for Telegram chat analysis

Current pick: qwen/qwen3-coder:free

Ranking Criterion

Free on OpenRouter + best summarization/extraction benchmark score

Rankings

# Item Score Reasoning Access
1 inclusionai/ling-3.1-flash 84.8 intelligenceindex:75 ctx:99 Free on OpenRouter · no retention
2 nvidia/nemotron-3-ultra-550b-a55b:free 76.6 intelligenceindex:42 ifbench:98 ctx:100 Free on OpenRouter · no retention
3 apodex/apodex-1.1-mini:free 69.2 intelligenceindex:48 ctx:99 Free on OpenRouter · no retention
4 thinkingmachines/inkling-small:free 68.9 intelligenceindex:47 ctx:100 Free on OpenRouter · no retention
5 thinkingmachines/inkling:free 68.9 intelligenceindex:47 ctx:100 Free on OpenRouter · no retention
6 nvidia/nemotron-3-super-120b-a12b:free 64.6 intelligenceindex:23 ifbench:84 ctx:99 Free on OpenRouter · no retention
7 inclusionai/ling-3.0-flash-sante:free 62.5 intelligenceindex:37 ctx:99 Free on OpenRouter · no retention
8 nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free 59.0 intelligenceindex:19 ifbench:73 ctx:99 Free on OpenRouter · no retention
9 cohere/north-mini-code:free 56.1 intelligenceindex:18 ifbench:65 ctx:99 Free on OpenRouter · no retention
10 nvidia/nemotron-3.5-lightning:free 55.3 intelligenceindex:23 ctx:100 Free on OpenRouter · no retention

Free OpenRouter Models (Latest Fetch)

Model ID Name Context
thinkingmachines/inkling-small:free Thinking Machines: Inkling Small (free) 1,048,576
thinkingmachines/inkling:free Thinking Machines: Inkling (free) 1,048,576
google/lyria-3-pro-preview Google: Lyria 3 Pro Preview 1,048,576
google/lyria-3-clip-preview Google: Lyria 3 Clip Preview 1,048,576
nvidia/nemotron-3.5-lightning:free NVIDIA: Nemotron 3.5 Lightning (free) 1,000,000
nvidia/nemotron-3-ultra-550b-a55b:free NVIDIA: Nemotron 3 Ultra (free) 1,000,000
dots-studio/dots-3-note-preview:free Dots Studio: Dots3-Note Preview (free) 512,000
inclusionai/ling-3.1-flash inclusionAI: Ling 3.1 Flash 262,144
apodex/apodex-1.1-mini:free Apodex: Apodex 1.1 Mini (free) 262,144
inclusionai/ling-3.0-flash-sante:free inclusionAI: Ling 3.0 Flash Sante (free) 262,144
poolside/laguna-s-2.1:free Poolside: Laguna S 2.1 (free) 262,144
poolside/laguna-xs-2.1:free Poolside: Laguna XS 2.1 (free) 262,144
google/gemma-4-26b-a4b-it:free Google: Gemma 4 26B A4B (free) 262,144
google/gemma-4-31b-it:free Google: Gemma 4 31B (free) 262,144
nvidia/nemotron-3-super-120b-a12b:free NVIDIA: Nemotron 3 Super (free) 262,144
cohere/north-mini-code:free Cohere: North Mini Code (free) 256,000
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free NVIDIA: Nemotron 3 Nano Omni (free) 256,000
openrouter/free Free Models Router 200,000
nvidia/nemotron-3.5-content-safety:free NVIDIA: Nemotron 3.5 Content Safety (free) 128,000
liquid/lfm-2.5-2.6b:free LiquidAI: LFM2.5-2.6B (free) 65,536

AI Methodology Reviews

Weekly question: Are we measuring the right capability for chat analysis tasks?

Needs Attention 2026-10-04 via nvidia/nemotron-3-super-120b-a12b:free

The current methodology prioritizes free availability on OpenRouter and a generic summarization/extraction benchmark, which may not fully capture the nuances required for chat analysis such as dialogue coherence, turn‑taking understanding, and sentiment detection. Since chat‑specific performance can diverge from pure summarization scores, the criterion should be revisited to include a chat‑oriented evaluation metric.

Needs Attention 2026-09-27 via nvidia/nemotron-3-super-120b-a12b:free

The current ranking relies on intelligence index and context length rather than a dedicated summarization/extraction benchmark, which does not directly assess the capability needed for chat analysis. To align with the stated criterion, the methodology should incorporate a summarization/extraction performance metric.

Needs Attention 2026-09-20 via none

All models failed — review manually

Needs Attention 2026-09-13 via nvidia/nemotron-3-super-120b-a12b:free

The current ranking uses intelligenceindex, IFBench, and context length as proxies for summarization/extraction ability, but these metrics do not directly measure chat‑analysis‑specific skills like extracting key points, sentiment, or actionable insights from conversational text. To align the criterion with the intended use case, the benchmark should incorporate a dedicated summarization/extraction evaluation (e.g., on chat logs or dialogue datasets).

Needs Attention 2026-09-06 via none

All models failed — review manually

Needs Attention 2026-08-30 via nvidia/nemotron-3-super-120b-a12b:free

The ranking relies on a summarization/extraction benchmark score, but the displayed top models are ordered by composite metrics (intelligenceindex, ifbench, ctx) that do not directly reflect summarization/extraction performance. This mismatch suggests the current criterion may not be capturing the intended capability for chat analysis, warranting a review of the measurement approach.

Needs Attention 2026-08-23 via nvidia/nemotron-3-super-120b-a12b:free

The ranking relies on generic intelligence, IFBench, and context length metrics rather than a dedicated summarization/extraction benchmark, which may not accurately reflect chat‑analysis performance. Consequently, the current top pick (qwen/qwen3-coder:free) does not align with the highest‑scoring models under the presented criteria, suggesting the methodology needs revision to measure the intended capability.

Needs Attention 2026-08-16 via none

All models failed — review manually

Needs Attention 2026-08-09 via nvidia/nemotron-3-super-120b-a12b:free

The criterion specifies using a summarization/extraction benchmark score, but the provided rankings are based on intelligenceindex, IFBench, and context length metrics, which do not directly measure summarization or extraction ability. Therefore the current methodology does not align with the stated capability for chat analysis tasks.

Needs Attention 2026-08-02 via nvidia/nemotron-3-super-120b-a12b:free

The ranking relies on intelligenceindex, IFBench, and context length metrics rather than a direct summarization/extraction benchmark, which may not accurately reflect chat‑analysis performance. Consequently, the current top pick (qwen/qwen3-coder:free) may not be the best model for the intended capability.