better-harness vs chinese-llm-benchmark

chinese-llm-benchmark is much bigger: 6.4k stars against 2.2k. Over the days we have tracked them better-harness moved +177.5% and chinese-llm-benchmark +5%, so better-harness is growing faster right now.

better-harness leads on context depth. chinese-llm-benchmark does not take any axis by a clear margin.

Stars and commit dates come from our own daily tracking. The six axes are read off each project's documentation by our review pipeline, so they describe what a project says about itself, not what we measured in its code.

Where they stand today

Reviews AI coding workflows end-to-end and produces evidence-bounded, prioritized findings across five Agent Work Loop dimensions rather than only analyzing final diffs.

Stars
2.2k
Tracked growth
+177.5%
Maturity
Last commit
1d ago
Language
JavaScript
License
MIT
Cost to run
Depends on host and agent (may use paid services)

Chinese-first, large-scale benchmark that covers hundreds of commercial and open-source models and ships a >2M-entry defect database for model diagnosis.

Stars
6.4k
Tracked growth
+5%
Maturity
Last commit
1d ago
Language
mixed
License
none declared
Cost to run
Requires Nonelinear API key (cloud service)
0%+177%86 tracked days
QoderAI/better-harnessjeinlee1991/chinese-llm-benchmark

Six axes, head to head

Each axis runs 0 to 5. The label under a score is what that project's own docs claim, not a category average.

Axisbetter-harnesschinese-llm-benchmark
Context depth
How much of your codebase it sees before it answers: the open diff, the diff plus related files, or the whole repository.
Whole-repo analysis
Diff only
Noise control
How it keeps output volume down — severity thresholds, deduplication, incremental runs over new commits only.
Evidence-backed findings
Basic scoring filters
Customization
How far it bends to your team: custom rules, prompts, style guides, per-path config.
Extensible configs & plugins
Config & custom data
Privacy
Whether your code stays on your own infrastructure: fully local, self-hostable, or cloud API only.
Self-hostable
Cloud API (your key)
Model freedom
Whether you can point it at any provider, or it is wired to one.
Multiple host providers
Gateway to many models
Setup ease
What it takes to get a first useful run out of it.
Host-specific quick start
API key required

Which one to pick

Pick better-harness if…

Workflow-focused — pick Better Harness when you want honest, evidence-backed reviews of your agents' task understanding, execution, validation, delivery, and learning capture.

  • Context depth: Whole-repo analysis (5/5 against 1/5)
Runs in cli, ide, coding-agent-plugin. Works with other-fixed.

Pick chinese-llm-benchmark if…

Chinese-focused — pick this when you need the most extensive, regularly updated multi-domain Chinese LLM benchmark and a huge defect corpus for analysis.

Runs in web-app. Works with other-fixed.

What people want from each one

Questions people ask

Is better-harness better than chinese-llm-benchmark?

better-harness leads on context depth. chinese-llm-benchmark does not take any axis by a clear margin. better-harness is worth picking when workflow-focused — pick Better Harness when you want honest, evidence-backed reviews of your agents' task understanding, execution, validation, delivery, and learning capture.

Which of better-harness and chinese-llm-benchmark keeps my code private?

better-harness: Self-hostable (4/5). chinese-llm-benchmark: Cloud API (your key) (3/5).

What does each one cost to run?

better-harness: Depends on host and agent (may use paid services). chinese-llm-benchmark: Requires Nonelinear API key (cloud service).

Full profiles: QoderAI/better-harness and jeinlee1991/chinese-llm-benchmark. Everything else in Evals & benchmarks.