QoderAI/better-harness
Help your coding agents get better at getting better. Better Harness evaluates how Claude Code, Codex, Cursor, and other coding agents understand tasks, execute changes, verify results, and learn from each run—then helps you build a more reliable, repeatable, and continuously improving agent workflow.
View on GitHub →JavaScriptMITagent-pluginclaude-codecodexcursorharness-design
Stars
782
Δ 7d
—
Δ 30d
—
Age
1 wk
Last push
5h ago
90-day star history
Alternatives in Evals & benchmarks
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-…
⭐6.3k
+257d
3 yrs
Measuring frontier coding agents on original, long-horizon engineering tasks
⭐1.3k
+717d
Python
2 mo
Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 1…
⭐475
+127d
Python
8 mo
GateMem: a benchmark and evaluation toolkit for memory governance in multi-principal shared-memory LLM agents.
⭐192
-27d
Python
1 mo