← Back to dashboard

CEO / Orchestrator (Agent Framework + LLM)

Top-level orchestration model for the Green Wick AI agent swarm. Ranked by agentic benchmarks (SWE-bench, GAIA, AgentBench, BFCL) — not just raw LLM reasoning. Framework scaffolding matters as much as the model.

Current pick: Claude Code (Opus 4.8) + native tool-use

Ranking Criterion

SWE-bench resolve rate + GAIA multi-step score + BFCL function-calling accuracy + tool-use reliability. Framework+LLM combo ranked together because scaffolding (prompt design, error recovery, memory management) affects results as much as model intelligence.

Rankings

# Item Score Reasoning Access
1 Claude Code (Opus 4.8) + native tool-use 98.5 SWE:100 GPQA:100 scaffold:95
2 Claude Code (Sonnet 4.6) + native tool-use 95.3 SWE:94 GPQA:100 scaffold:95
3 Codex CLI + GPT-5.5 (xhigh) 94.4 TB:100 CI:100 IF:90 GPQA:100 scaffold:82
4 Cursor + GPT-5.5 (xhigh) 93.0 TB:100 CI:100 IF:90 GPQA:100 scaffold:75
5 Aider + Claude Opus 4.8 89.5 SWE:100 GPQA:100 scaffold:65
6 Antigravity + Gemini 3 Pro 83.8 TB:72 IF:83 SWE:90 GPQA:100 scaffold:80
7 OpenHands + Claude Sonnet 4.6 83.3 SWE:94 GPQA:100 scaffold:55

AI Methodology Reviews

Weekly question: Has any Framework+LLM combo surpassed the current pick on SWE-bench, GAIA, or AgentBench? Any new agentic benchmarks (WebArena, OSWorld) worth adding to the ranking criterion?

Needs Attention 2026-10-04 via nvidia/nemotron-3-super-120b-a12b:free

As of the latest data, no Framework+LLM combination has exceeded Claude Code (Opus 4.8) + native tool-use on SWE‑bench, GAIA, or AgentBench scores. However, newer agent‑centric benchmarks such as WebArena and OSWorld have gained traction and should be evaluated for inclusion to ensure the ranking reflects current tool‑use and multi‑step reasoning capabilities.

Needs Attention 2026-09-27 via nvidia/nemotron-3-super-120b-a12b:free

No Framework+LLM combination has exceeded Claude Code (Opus 4.8) + native tool-use on SWE-bench, GAIA, or AgentBench, so the existing criterion still ranks the current pick correctly. However, emerging agentic benchmarks such as WebArena and OSWorld evaluate real‑world tool use, multi‑step planning, and environment interaction—core aspects of the current methodology—suggesting the ranking should be updated to incorporate these metrics for continued validity.

Valid 2026-09-20 via nvidia/nemotron-3-super-120b-a12b:free

None of the listed Framework+LLM combos exceed Claude Code (Opus 4.8) + native tool-use on SWE-bench, GAIA, or AgentBench scores, and no newer agentic benchmarks (e.g., WebArena, OSWorld) have been introduced that would necessitate criterion changes. Therefore the existing ranking methodology remains sound.

Valid 2026-09-13 via nvidia/nemotron-3-super-120b-a12b:free

The current top-ranked combo (Claude Code Opus 4.8 + native tool‑use) remains ahead of all listed alternatives on the existing metrics, and no newer agent‑centric benchmarks (e.g., WebArena, OSWorld) have been reported to show a different ordering. Thus the existing criterion continues to capture the relative performance of framework+LLM pairs.

Needs Attention 2026-09-06 via nvidia/nemotron-3-super-120b-a12b:free

No publicly reported Framework+LLM combination has exceeded Claude Code (Opus 4.8) on SWE‑bench, GAIA, or AgentBench, so the core criterion remains valid. However, emerging agent‑centric benchmarks such as WebArena and OSWorld reveal notable performance differences across scaffolds and LLMs, suggesting the ranking should be expanded to capture these dimensions for a more up‑to‑date assessment.

Needs Attention 2026-08-30 via nvidia/nemotron-3-super-120b-a12b:free

No publicly reported Framework+LLM combination has exceeded Claude Code (Opus 4.8) + native tool-use on SWE‑bench, GAIA, or AgentBench as of the latest data, so the core criterion remains valid. However, emerging agent‑centric benchmarks such as WebArena and OSWorld have gained traction and should be evaluated for inclusion to ensure the ranking reflects the full spectrum of tool‑use and reasoning capabilities.

Valid 2026-08-23 via nvidia/nemotron-3-super-120b-a12b:free

The current top-ranked Claude Code (Opus 4.8) + native tool‑use remains ahead of all listed alternatives on the existing criteria (SWE‑bench, GAIA, BFCL, tool‑use reliability), and no newer agentic benchmarks such as WebArena or OSWorld have been reported to shift the rankings. Therefore the methodology continues to reflect the relative performance of framework+LLM combos accurately.

Needs Attention 2026-08-16 via none

All models failed — review manually

Valid 2026-08-09 via nvidia/nemotron-3-super-120b-a12b:free

The current top entry (Claude Code + Opus 4.8) maintains the highest composite score, and none of the listed alternatives exceed it on SWE‑bench, GAIA, or AgentBench. No new agentic benchmark results (e.g., WebArena, OSWorld) have been presented that would overturn this ordering, so the existing criterion remains appropriate.

Valid 2026-08-02 via nvidia/nemotron-3-super-120b-a12b:free

None of the listed Framework+LLM combos exceed the current pick's score, indicating no surpassing performance on SWE-bench, GAIA, or AgentBench. There is also no provided evidence that newer agentic benchmarks like WebArena or OSWorld warrant changes to the criterion at this time.