A benchmark for measuring model-harness configuration effects across 300 sandboxed agent tasks in realistic workflows, scored by task-specific oracles without an LLM judge.
274 general and 26 multimodal tasks across six work categories and 17 subcategories, with explicit input resources, deliverables, and operational constraints. Scores below are from the general suite.
What paired comparisons of seven native harnesses on shared tasks reveal about delivery quality and cost.
Open each category to browse its subcategories and every task title, with category-level results from the general suite.
Harness Bench fixes external task conditions and scoring rules, preserves native harness behavior, and records token usage and API expenditure.
If you use Harness Bench in your research, please cite our paper using the BibTeX entry below.
@misc{harnessbench2026, title = {Harness Bench: Measuring Harness Effects in Realistic Agent Workflows}, author = {Harness Bench Team}, year = {2026}, url = {https://arxiv.org/abs/2605.27922}, note = {300 sandboxed agent tasks across 6 categories, 7 harnesses, and 10 model backends} }