Diagnostic Benchmark 2026

Harness Bench

A benchmark for measuring model-harness configuration effects across 300 sandboxed agent tasks in realistic workflows, scored by task-specific oracles without an LLM judge.

300
Total Tasks
7
Native Harnesses
10
Model Backends
20,454
Evaluated Runs
Scroll
Task Suite

Task Suite Overview

274 general and 26 multimodal tasks across six work categories and 17 subcategories, with explicit input resources, deliverables, and operational constraints. Scores below are from the general suite.

View All Tasks
Implementation and Repair · Repository Maintenance
81tasks
Software Engineering
Implementing features, repairing defects, and maintaining repositories: tests, migrations, dependency upgrades, and build or CI fixes inside realistic project workspaces.
2
Subcategories
57.3
Mean Score
39.4%
Score ≥ 0.80
22.8%
Full Credit
Insights

Key Findings

What paired comparisons of seven native harnesses on shared tasks reveal about delivery quality and cost.

FINDING 01
⚖️
No Harness Is Best on Every Model
Six of the seven harnesses lead at least one of the ten model backends on the general suite. Within nine models, score gaps across harnesses range from 5.0 to 9.5 points, and reach 39.2 points for Qwen3.8-27B.
FINDING 02
🎯
Full-Credit Delivery Is More Harness-Sensitive Than Mean Score
GPT-5.6 Sol varies by only 1.93 points (SD) in mean score across harnesses, yet its full-credit rate varies by 9.12 points. For both GLM-5.3 variants, the full-credit gap reaches 13.5 percentage points.
FINDING 03
💸
Higher Spending Does Not Ensure Better Delivery
With GPT-5.6 Sol, Nanobot incurs 2.97× OpenClaw's API cost while achieving a lower full-credit rate (24.82% vs. 31.02%).
FINDING 04
🧭
The Delivery Objective Changes the Best Choice
Configurations that lead mean score need not lead full-credit attainment. Selecting a harness per model improves held-out full-credit attainment by 2.53–2.73 points over one shared harness, while mean-score gains stay within ±0.14 points.
FINDING 05
📊
High Partial Credit Can Conceal Incomplete Work
On the 274 general tasks, about half of all attempts score at least 0.80 (R0.8 = 50.81%), but only about a quarter receive full credit (R1 = 26.22%). Paired trajectories trace the gap to feedback that does not change the next action, evidence lost during context management, and decisions that never become checkable work.
R0.8 vs. full credit by harness (all evaluations)
Claude
Codex
OpenCode
OpenClaw
Hermes
Goose
Nanobot
Score ≥ 0.80 (R0.8)
Full credit (R1)
Browse by Category

Explore Every Category

Open each category to browse its subcategories and every task title, with category-level results from the general suite.

How It Works

Evaluation Methodology

Harness Bench fixes external task conditions and scoring rules, preserves native harness behavior, and records token usage and API expenditure.

01
Step 01
Task Construction & Sandboxing
Work objectives come from de-identified user logs, public workflows, and resource collections, and are reconstructed as sandboxed workspaces of files, repositories, records, and local services.
Log-based and domain-based synthesis
Explicit resources, deliverables, and constraints
Workflow refinement of business dependencies
Human review; reference solutions earn full credit
02
Step 02
Native Harness Execution
Each model-harness configuration receives identical task inputs in an isolated workspace and keeps its own prompting, tools, context management, and recovery behavior.
7 harnesses × 10 models (general), × 7 (multimodal)
200K-token context threshold, 150-turn limit
Token usage, requests, and API cost recorded
Traces kept for trajectory analysis
03
Step 03
Module-Level Oracle Scoring
Hidden task-specific oracles check only final artifacts and workspace state. A module earns its weight only when all of its checks pass; safety and risk gates zero the score on violations.
1,408 scored modules, 3–7 per task
No LLM judge in scoring
Gates for protected inputs, sensitive data, prohibited actions
Reported as mean score, R0.8, and full credit (R1)
Citation

Cite This Work

If you use Harness Bench in your research, please cite our paper using the BibTeX entry below.

@misc{harnessbench2026,
  title     = {Harness Bench: Measuring Harness Effects in Realistic Agent Workflows},
  author    = {Harness Bench Team},
  year      = {2026},
  url       = {https://arxiv.org/abs/2605.27922},
  note      = {300 sandboxed agent tasks across 6 categories, 7 harnesses, and 10 model backends}
}