QoderAI/better-harness
An open-source Harness Engineering platform for coding agents—define harnesses as code, run controlled experiments, inspect evidence, and compare outcomes. Turn task evidence into actionable team and organization insights.
View on GitHub →Where this sits in the catalog
Most-wanted open issues
Ranked by 👍 reactions. 4 open issues total, checked 2026-09-09.
| Issue | 👍 | Comments | Opened |
|---|---|---|---|
| 0 | 1 | 1 mo | |
[Feature]:OpenClaw Host Support enhancement | 0 | 1 | 2 wk |
| 0 | 1 | 1 mo | |
| 0 | 1 | 2 wk |
Head to head with its peers
The same six axes and the same star history, two projects at a time.
Alternatives in Evals & benchmarks
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-…
A model-mediated harness for reliable agentic software development.
Measuring frontier coding agents on original, long-horizon engineering tasks
Skills that guide AI coding agents to help you build product-specific AI evals.
Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 1…