Kimi CLI
Kimi K3
Latest benchmark: task-set v0.16.0
A compact benchmark for comparing agent harnesses on file edits,
data transforms, pytest bug fixes, memory discipline, synthetic
benchmark-like workflows, version-control work, skill discovery,
adversarial / hostile-environment handling, and CLI composition
with bespoke tools whose --help must actually be read.
logq, pktool, xtab, …) with unguessable flags, plus POSIX pipelines whose solve.sh the verifier executes
Current results only
Results below use the full 391-task set (task-set v0.16.0); the
v0.16.0 audit changed task semantics, so earlier scores are not
comparable. One full run per harness + model setup is listed.
Filter by model group or a single harness; steps and tokens are
shown only when the harness exposes them. GigaChat rows are the
IFT stand, build 32.9.23.6; GigaChat-profile rows were measured
with deepagents-gigachat 0.0.3 (see the README for
the result-affecting 0.0.4 pin change). Rows within ~3 points of
each other are a tie: a single mid-scale run carries ±1-2 pp of
sampling noise.
Kimi K3
DeepSeek V4 Flash 0731 (high)
Claude Haiku 4.5
GLM-5.2 (self-hosted)
GigaChat 3.5
GigaChat 3 Ultra
DeepSeek V4 Flash
GigaChat 3 Ultra
GigaChat 3.5
Qwen3 Coder 30B-A3B
GigaChat 3 Pro
GPT-OSS-20B
GPT-OSS-120B
GigaChat 3 Lightning
No runs match these filters.
Verification
Each task is checked mechanically: exact files, JSON shape and values, pytest outcomes, SQLite/XLSX checks, regex matches, and deterministic command outputs.