Toolbox/Evals & benchmarks/QoderAI/better-harness

QoderAI/better-harness

An open-source Harness Engineering platform for coding agents—define harnesses as code, run controlled experiments, inspect evidence, and compare outcomes. Turn task evidence into actionable team and organization insights.

View on GitHub →
JavaScriptMITagent-pluginclaude-codecodexcursorharness-design
Stars
2.3k
Δ 7d
+96
Δ 30d
+438
Age
1 mo
Last push
2h ago
2.4k2.0k1.5k1.1k633Jul 29Aug 20Sep 12
90-day star history

Where this sits in the catalog

Most-wanted open issues

Ranked by 👍 reactions. 4 open issues total, checked 2026-09-09.

Head to head with its peers

The same six axes and the same star history, two projects at a time.

Alternatives in Evals & benchmarks

非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-…

6.4k
+147d
3 yrs

A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)

3.7k
+147d
Python
3 yrs

A model-mediated harness for reliable agentic software development.

2.8k
+17d
Python
7 yrs

Measuring frontier coding agents on original, long-horizon engineering tasks

1.7k
+537d
Python
4 mo

Skills that guide AI coding agents to help you build product-specific AI evals.

580
+427d
2 mo

Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 1…

554
+87d
Python
9 mo