Latest benchmark: task-set v0.16.0

391 mechanically verified coding-agent tasks

A compact benchmark for comparing agent harnesses on file edits, data transforms, pytest bug fixes, memory discipline, synthetic benchmark-like workflows, version-control work, skill discovery, adversarial / hostile-environment handling, and CLI composition with bespoke tools whose --help must actually be read.

391 tasks
14 runs listed
12 models
99.7% top score
Kimi CLI / Kimi K3 99.7%
Pi / DeepSeek V4 Flash 0731 99.0%
Claude Code CLI / Haiku 4.5 97.2%
1-253 File, code, data, diagnostic, and memory tasks
254-313 Synthetic agentic (Terminal-Bench / tau / SWE-bench-like) and Git / version-control workflows
314-351 Skill-discovery tasks and adversarial / hostile-environment handling (broken versions, unreadable files, contradictions, a ~100 MB log)
352-391 Calibrated Terminal-Bench-inspired workflows and CLI composition: bespoke tools (logq, pktool, xtab, …) with unguessable flags, plus POSIX pipelines whose solve.sh the verifier executes

Current results only

Leaderboard

Results below use the full 391-task set (task-set v0.16.0); the v0.16.0 audit changed task semantics, so earlier scores are not comparable. One full run per harness + model setup is listed. Filter by model group or a single harness; steps and tokens are shown only when the harness exposes them. GigaChat rows are the IFT stand, build 32.9.23.6; GigaChat-profile rows were measured with deepagents-gigachat 0.0.3 (see the README for the result-affecting 0.0.4 pin change). Rows within ~3 points of each other are a tie: a single mid-scale run carries ±1-2 pp of sampling noise.

14 of 14 shown

1

Kimi CLI

Kimi K3

390/391
99.7%
Steps — · Tokens —
2

Pi

DeepSeek V4 Flash 0731 (high)

387/391
99.0%
Steps 1,694 · Tokens 10,004,355
3

Claude Code CLI

Claude Haiku 4.5

380/391
97.2%
Steps 1,645 · Tokens 176,430,286
4

opencode

GLM-5.2 (self-hosted)

362/391
92.6%
Steps — · Tokens —
5

deepagents + GigaChat profile

GigaChat 3.5

340/391
87.0%
Steps 3,316 · Tokens 4,319,421
5

deepagents + GigaChat profile

GigaChat 3 Ultra

340/391
87.0%
Steps 3,258 · Tokens 4,415,043
7

deepagents

DeepSeek V4 Flash

320/391
81.8%
Steps 5,048 · Tokens 61,715,547
8

deepagents, no profile

GigaChat 3 Ultra

312/391
79.8%
Steps 3,591 · Tokens 7,994,643
9

deepagents, no profile

GigaChat 3.5

302/391
77.2%
Steps 3,542 · Tokens 7,371,576
10

deepagents

Qwen3 Coder 30B-A3B

284/391
72.6%
Steps 5,467 · Tokens 90,234,558
11

deepagents + GigaChat profile

GigaChat 3 Pro

241/391
61.6%
Steps 2,993 · Tokens 3,975,672
12

deepagents

GPT-OSS-20B

193/391
49.4%
Steps 2,727 · Tokens 29,137,983
13

deepagents

GPT-OSS-120B

186/391
47.6%
Steps 2,193 · Tokens 23,831,283
14

deepagents + GigaChat profile

GigaChat 3 Lightning

178/391
45.5%
Steps 2,520 · Tokens 2,275,821

Verification

No LLM judge

Each task is checked mechanically: exact files, JSON shape and values, pytest outcomes, SQLite/XLSX checks, regex matches, and deterministic command outputs.