Top-level orchestration model for the Green Wick AI agent swarm. Ranked by agentic benchmarks (SWE-bench, GAIA, AgentBench, BFCL) — not just raw LLM reasoning. Framework scaffolding matters as much as the model.
SWE-bench resolve rate + GAIA multi-step score + BFCL function-calling accuracy + tool-use reliability. Framework+LLM combo ranked together because scaffolding (prompt design, error recovery, memory management) affects results as much as model intelligence.
| # | Item | Score | Reasoning | Access |
|---|---|---|---|---|
| 1 | Claude Code (Opus 4.8) + native tool-use | 98.5 | SWE:100 GPQA:100 scaffold:95 | |
| 2 | Claude Code (Sonnet 4.6) + native tool-use | 95.3 | SWE:94 GPQA:100 scaffold:95 | |
| 3 | Codex CLI + GPT-5.5 (xhigh) | 94.4 | TB:100 CI:100 IF:90 GPQA:100 scaffold:82 | |
| 4 | Cursor + GPT-5.5 (xhigh) | 93.0 | TB:100 CI:100 IF:90 GPQA:100 scaffold:75 | |
| 5 | Aider + Claude Opus 4.8 | 89.5 | SWE:100 GPQA:100 scaffold:65 | |
| 6 | Antigravity + Gemini 3 Pro | 83.8 | TB:72 IF:83 SWE:90 GPQA:100 scaffold:80 | |
| 7 | OpenHands + Claude Sonnet 4.6 | 83.3 | SWE:94 GPQA:100 scaffold:55 |
Weekly question: Has any Framework+LLM combo surpassed the current pick on SWE-bench, GAIA, or AgentBench? Any new agentic benchmarks (WebArena, OSWorld) worth adding to the ranking criterion?
As of the latest data, no Framework+LLM combination has exceeded Claude Code (Opus 4.8) + native tool-use on SWE‑bench, GAIA, or AgentBench scores. However, newer agent‑centric benchmarks such as WebArena and OSWorld have gained traction and should be evaluated for inclusion to ensure the ranking reflects current tool‑use and multi‑step reasoning capabilities.
No Framework+LLM combination has exceeded Claude Code (Opus 4.8) + native tool-use on SWE-bench, GAIA, or AgentBench, so the existing criterion still ranks the current pick correctly. However, emerging agentic benchmarks such as WebArena and OSWorld evaluate real‑world tool use, multi‑step planning, and environment interaction—core aspects of the current methodology—suggesting the ranking should be updated to incorporate these metrics for continued validity.
None of the listed Framework+LLM combos exceed Claude Code (Opus 4.8) + native tool-use on SWE-bench, GAIA, or AgentBench scores, and no newer agentic benchmarks (e.g., WebArena, OSWorld) have been introduced that would necessitate criterion changes. Therefore the existing ranking methodology remains sound.
The current top-ranked combo (Claude Code Opus 4.8 + native tool‑use) remains ahead of all listed alternatives on the existing metrics, and no newer agent‑centric benchmarks (e.g., WebArena, OSWorld) have been reported to show a different ordering. Thus the existing criterion continues to capture the relative performance of framework+LLM pairs.
No publicly reported Framework+LLM combination has exceeded Claude Code (Opus 4.8) on SWE‑bench, GAIA, or AgentBench, so the core criterion remains valid. However, emerging agent‑centric benchmarks such as WebArena and OSWorld reveal notable performance differences across scaffolds and LLMs, suggesting the ranking should be expanded to capture these dimensions for a more up‑to‑date assessment.
No publicly reported Framework+LLM combination has exceeded Claude Code (Opus 4.8) + native tool-use on SWE‑bench, GAIA, or AgentBench as of the latest data, so the core criterion remains valid. However, emerging agent‑centric benchmarks such as WebArena and OSWorld have gained traction and should be evaluated for inclusion to ensure the ranking reflects the full spectrum of tool‑use and reasoning capabilities.
The current top-ranked Claude Code (Opus 4.8) + native tool‑use remains ahead of all listed alternatives on the existing criteria (SWE‑bench, GAIA, BFCL, tool‑use reliability), and no newer agentic benchmarks such as WebArena or OSWorld have been reported to shift the rankings. Therefore the methodology continues to reflect the relative performance of framework+LLM combos accurately.
All models failed — review manually
The current top entry (Claude Code + Opus 4.8) maintains the highest composite score, and none of the listed alternatives exceed it on SWE‑bench, GAIA, or AgentBench. No new agentic benchmark results (e.g., WebArena, OSWorld) have been presented that would overturn this ordering, so the existing criterion remains appropriate.
None of the listed Framework+LLM combos exceed the current pick's score, indicating no surpassing performance on SWE-bench, GAIA, or AgentBench. There is also no provided evidence that newer agentic benchmarks like WebArena or OSWorld warrant changes to the criterion at this time.