better-harness vs chinese-llm-benchmark
chinese-llm-benchmark is much bigger: 6.4k stars against 2.2k. Over the days we have tracked them better-harness moved +177.5% and chinese-llm-benchmark +5%, so better-harness is growing faster right now.
better-harness leads on context depth. chinese-llm-benchmark does not take any axis by a clear margin.
Stars and commit dates come from our own daily tracking. The six axes are read off each project's documentation by our review pipeline, so they describe what a project says about itself, not what we measured in its code.
Where they stand today
Reviews AI coding workflows end-to-end and produces evidence-bounded, prioritized findings across five Agent Work Loop dimensions rather than only analyzing final diffs.
- Stars
- 2.2k
- Tracked growth
- +177.5%
- Maturity
- ●●●●●
- Last commit
- 1d ago
- Language
- JavaScript
- License
- MIT
- Cost to run
- Depends on host and agent (may use paid services)
Chinese-first, large-scale benchmark that covers hundreds of commercial and open-source models and ships a >2M-entry defect database for model diagnosis.
- Stars
- 6.4k
- Tracked growth
- +5%
- Maturity
- ●●●●●
- Last commit
- 1d ago
- Language
- mixed
- License
- none declared
- Cost to run
- Requires Nonelinear API key (cloud service)
Six axes, head to head
Each axis runs 0 to 5. The label under a score is what that project's own docs claim, not a category average.
| Axis | better-harness | chinese-llm-benchmark |
|---|---|---|
Context depth How much of your codebase it sees before it answers: the open diff, the diff plus related files, or the whole repository. | ●●●●● Whole-repo analysis | ●●●●● Diff only |
Noise control How it keeps output volume down — severity thresholds, deduplication, incremental runs over new commits only. | ●●●●● Evidence-backed findings | ●●●●● Basic scoring filters |
Customization How far it bends to your team: custom rules, prompts, style guides, per-path config. | ●●●●● Extensible configs & plugins | ●●●●● Config & custom data |
Privacy Whether your code stays on your own infrastructure: fully local, self-hostable, or cloud API only. | ●●●●● Self-hostable | ●●●●● Cloud API (your key) |
Model freedom Whether you can point it at any provider, or it is wired to one. | ●●●●● Multiple host providers | ●●●●● Gateway to many models |
Setup ease What it takes to get a first useful run out of it. | ●●●●● Host-specific quick start | ●●●●● API key required |
Which one to pick
Pick better-harness if…
Workflow-focused — pick Better Harness when you want honest, evidence-backed reviews of your agents' task understanding, execution, validation, delivery, and learning capture.
- Context depth: Whole-repo analysis (5/5 against 1/5)
Pick chinese-llm-benchmark if…
Chinese-focused — pick this when you need the most extensive, regularly updated multi-domain Chinese LLM benchmark and a huge defect corpus for analysis.
What people want from each one
QoderAI/better-harness
jeinlee1991/chinese-llm-benchmark
Questions people ask
Is better-harness better than chinese-llm-benchmark?
better-harness leads on context depth. chinese-llm-benchmark does not take any axis by a clear margin. better-harness is worth picking when workflow-focused — pick Better Harness when you want honest, evidence-backed reviews of your agents' task understanding, execution, validation, delivery, and learning capture.
Which of better-harness and chinese-llm-benchmark keeps my code private?
better-harness: Self-hostable (4/5). chinese-llm-benchmark: Cloud API (your key) (3/5).
What does each one cost to run?
better-harness: Depends on host and agent (may use paid services). chinese-llm-benchmark: Requires Nonelinear API key (cloud service).
Full profiles: QoderAI/better-harness and jeinlee1991/chinese-llm-benchmark. Everything else in Evals & benchmarks.