Claude Code CLI
Claude Opus 4.8
Latest benchmark: task-set v0.13.0
A compact benchmark for comparing agent harnesses on file edits, data transforms, pytest bug fixes, memory discipline, synthetic benchmark-like workflows, version-control work, skill discovery, and adversarial / hostile-environment handling.
Current results only
Results below use the full 351-task set (v0.13.0). One run per harness + model setup is listed. Filter by model group or a single harness; steps and tokens are shown only when the harness exposes them.
Claude Opus 4.8
grok-4.5
Claude Sonnet 4.6
Claude Haiku 4.5
GLM-5.2
DeepSeek V4 Pro
GLM-5.1
GPT-5.6 Luna
Claude Haiku 4.5
Qwen 3.7 Max
DeepSeek V3.2
Qwen 3.6 Flash
GigaChat 3.5
DeepSeek V4 Flash
GigaChat 3 Ultra
GPT-4.1
GigaChat 3.5
GigaChat 2 Max
Qwen 3.5 Flash
GigaChat 3.5
MiniMax M2.7
Qwen3-Coder-30B-A3B
GigaChat 3 Pro
yandex/gpt5.1-pro
GPT-OSS-120B
yandex/gpt5-pro
GigaChat 3 Lightning
Llama 4 Maverick
yandex/gpt5-lite
No runs match these filters.
Verification
Each task is checked mechanically: exact files, JSON shape and values, pytest outcomes, SQLite/XLSX checks, regex matches, and deterministic command outputs.