Summarization and extraction model for Telegram chat analysis
Free on OpenRouter + best summarization/extraction benchmark score
| # | Item | Score | Reasoning | Access |
|---|---|---|---|---|
| 1 | inclusionai/ling-3.1-flash | 84.8 | intelligenceindex:75 ctx:99 | Free on OpenRouter · no retention |
| 2 | nvidia/nemotron-3-ultra-550b-a55b:free | 76.6 | intelligenceindex:42 ifbench:98 ctx:100 | Free on OpenRouter · no retention |
| 3 | apodex/apodex-1.1-mini:free | 69.2 | intelligenceindex:48 ctx:99 | Free on OpenRouter · no retention |
| 4 | thinkingmachines/inkling-small:free | 68.9 | intelligenceindex:47 ctx:100 | Free on OpenRouter · no retention |
| 5 | thinkingmachines/inkling:free | 68.9 | intelligenceindex:47 ctx:100 | Free on OpenRouter · no retention |
| 6 | nvidia/nemotron-3-super-120b-a12b:free | 64.6 | intelligenceindex:23 ifbench:84 ctx:99 | Free on OpenRouter · no retention |
| 7 | inclusionai/ling-3.0-flash-sante:free | 62.5 | intelligenceindex:37 ctx:99 | Free on OpenRouter · no retention |
| 8 | nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | 59.0 | intelligenceindex:19 ifbench:73 ctx:99 | Free on OpenRouter · no retention |
| 9 | cohere/north-mini-code:free | 56.1 | intelligenceindex:18 ifbench:65 ctx:99 | Free on OpenRouter · no retention |
| 10 | nvidia/nemotron-3.5-lightning:free | 55.3 | intelligenceindex:23 ctx:100 | Free on OpenRouter · no retention |
| Model ID | Name | Context |
|---|---|---|
| thinkingmachines/inkling-small:free | Thinking Machines: Inkling Small (free) | 1,048,576 |
| thinkingmachines/inkling:free | Thinking Machines: Inkling (free) | 1,048,576 |
| google/lyria-3-pro-preview | Google: Lyria 3 Pro Preview | 1,048,576 |
| google/lyria-3-clip-preview | Google: Lyria 3 Clip Preview | 1,048,576 |
| nvidia/nemotron-3.5-lightning:free | NVIDIA: Nemotron 3.5 Lightning (free) | 1,000,000 |
| nvidia/nemotron-3-ultra-550b-a55b:free | NVIDIA: Nemotron 3 Ultra (free) | 1,000,000 |
| dots-studio/dots-3-note-preview:free | Dots Studio: Dots3-Note Preview (free) | 512,000 |
| inclusionai/ling-3.1-flash | inclusionAI: Ling 3.1 Flash | 262,144 |
| apodex/apodex-1.1-mini:free | Apodex: Apodex 1.1 Mini (free) | 262,144 |
| inclusionai/ling-3.0-flash-sante:free | inclusionAI: Ling 3.0 Flash Sante (free) | 262,144 |
| poolside/laguna-s-2.1:free | Poolside: Laguna S 2.1 (free) | 262,144 |
| poolside/laguna-xs-2.1:free | Poolside: Laguna XS 2.1 (free) | 262,144 |
| google/gemma-4-26b-a4b-it:free | Google: Gemma 4 26B A4B (free) | 262,144 |
| google/gemma-4-31b-it:free | Google: Gemma 4 31B (free) | 262,144 |
| nvidia/nemotron-3-super-120b-a12b:free | NVIDIA: Nemotron 3 Super (free) | 262,144 |
| cohere/north-mini-code:free | Cohere: North Mini Code (free) | 256,000 |
| nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:free | NVIDIA: Nemotron 3 Nano Omni (free) | 256,000 |
| openrouter/free | Free Models Router | 200,000 |
| nvidia/nemotron-3.5-content-safety:free | NVIDIA: Nemotron 3.5 Content Safety (free) | 128,000 |
| liquid/lfm-2.5-2.6b:free | LiquidAI: LFM2.5-2.6B (free) | 65,536 |
Weekly question: Are we measuring the right capability for chat analysis tasks?
The current methodology prioritizes free availability on OpenRouter and a generic summarization/extraction benchmark, which may not fully capture the nuances required for chat analysis such as dialogue coherence, turn‑taking understanding, and sentiment detection. Since chat‑specific performance can diverge from pure summarization scores, the criterion should be revisited to include a chat‑oriented evaluation metric.
The current ranking relies on intelligence index and context length rather than a dedicated summarization/extraction benchmark, which does not directly assess the capability needed for chat analysis. To align with the stated criterion, the methodology should incorporate a summarization/extraction performance metric.
All models failed — review manually
The current ranking uses intelligenceindex, IFBench, and context length as proxies for summarization/extraction ability, but these metrics do not directly measure chat‑analysis‑specific skills like extracting key points, sentiment, or actionable insights from conversational text. To align the criterion with the intended use case, the benchmark should incorporate a dedicated summarization/extraction evaluation (e.g., on chat logs or dialogue datasets).
All models failed — review manually
The ranking relies on a summarization/extraction benchmark score, but the displayed top models are ordered by composite metrics (intelligenceindex, ifbench, ctx) that do not directly reflect summarization/extraction performance. This mismatch suggests the current criterion may not be capturing the intended capability for chat analysis, warranting a review of the measurement approach.
The ranking relies on generic intelligence, IFBench, and context length metrics rather than a dedicated summarization/extraction benchmark, which may not accurately reflect chat‑analysis performance. Consequently, the current top pick (qwen/qwen3-coder:free) does not align with the highest‑scoring models under the presented criteria, suggesting the methodology needs revision to measure the intended capability.
All models failed — review manually
The criterion specifies using a summarization/extraction benchmark score, but the provided rankings are based on intelligenceindex, IFBench, and context length metrics, which do not directly measure summarization or extraction ability. Therefore the current methodology does not align with the stated capability for chat analysis tasks.
The ranking relies on intelligenceindex, IFBench, and context length metrics rather than a direct summarization/extraction benchmark, which may not accurately reflect chat‑analysis performance. Consequently, the current top pick (qwen/qwen3-coder:free) may not be the best model for the intended capability.