Toolbox/Evals & benchmarks/QoderAI/better-harness

QoderAI/better-harness

Help your coding agents get better at getting better. Better Harness evaluates how Claude Code, Codex, Cursor, and other coding agents understand tasks, execute changes, verify results, and learn from each run—then helps you build a more reliable, repeatable, and continuously improving agent workflow.

View on GitHub →
JavaScriptMITagent-pluginclaude-codecodexcursorharness-design
Stars
782
Δ 7d
Δ 30d
Age
1 wk
Last push
5h ago
789785781777773Jul 29
90-day star history
#5 of 10 in Evals & benchmarks

Alternatives in Evals & benchmarks

非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-…

6.3k
+257d
3 yrs

A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)

3.6k
+277d
Python
3 yrs

Deterministic execution for non-deterministic AI.

2.8k
+17d
Python
7 yrs

Measuring frontier coding agents on original, long-horizon engineering tasks

1.3k
+717d
Python
2 mo

Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 1…

475
+127d
Python
8 mo

GateMem: a benchmark and evaluation toolkit for memory governance in multi-principal shared-memory LLM agents.

192
-27d
Python
1 mo