chinese-llm-benchmark vs teaql-agent-kit
chinese-llm-benchmark is much bigger: 6.4k stars against 2.8k. Over the days we have tracked them chinese-llm-benchmark moved +5% and teaql-agent-kit +1.2%, so chinese-llm-benchmark is growing faster right now.
They split the axes: chinese-llm-benchmark leads on model freedom, teaql-agent-kit on context depth.
Stars and commit dates come from our own daily tracking. The six axes are read off each project's documentation by our review pipeline, so they describe what a project says about itself, not what we measured in its code.
Where they stand today
Chinese-first, large-scale benchmark that covers hundreds of commercial and open-source models and ships a >2M-entry defect database for model diagnosis.
- Stars
- 6.4k
- Tracked growth
- +5%
- Maturity
- ●●●●●
- Last commit
- 1d ago
- Language
- mixed
- License
- none declared
- Cost to run
- Requires Nonelinear API key (cloud service)
Provides a TEAQL-focused, auditable evaluation harness that measures software-engineering discipline and token-efficiency for coding agents rather than offering general-purpose agent automation.
- Stars
- 2.8k
- Tracked growth
- +1.2%
- Maturity
- ●●●●●
- Last commit
- 2d ago
- Language
- Python
- License
- MIT
- Cost to run
- Your API key or model costs may apply
Six axes, head to head
Each axis runs 0 to 5. The label under a score is what that project's own docs claim, not a category average.
| Axis | chinese-llm-benchmark | teaql-agent-kit |
|---|---|---|
Context depth How much of your codebase it sees before it answers: the open diff, the diff plus related files, or the whole repository. | ●●●●● Diff only | ●●●●● Diff + related files |
Noise control How it keeps output volume down — severity thresholds, deduplication, incremental runs over new commits only. | ●●●●● Basic scoring filters | ●●●●● Guides & checkpoints |
Customization How far it bends to your team: custom rules, prompts, style guides, per-path config. | ●●●●● Config & custom data | ●●●●● Config & prompts |
Privacy Whether your code stays on your own infrastructure: fully local, self-hostable, or cloud API only. | ●●●●● Cloud API (your key) | ●●●●● Cloud via API key |
Model freedom Whether you can point it at any provider, or it is wired to one. | ●●●●● Gateway to many models | ●●●●● Single fixed provider |
Setup ease What it takes to get a first useful run out of it. | ●●●●● API key required | ●●●●● Manual setup |
Which one to pick
Pick chinese-llm-benchmark if…
Chinese-focused — pick this when you need the most extensive, regularly updated multi-domain Chinese LLM benchmark and a huge defect corpus for analysis.
- Model freedom: Gateway to many models (3/5 against 1/5)
Pick teaql-agent-kit if…
Evaluation-first — pick this when you need reproducible, auditable benchmarks of coding agents working with TEAQL contracts, explicit guardrails, and token-efficiency measurement.
- Context depth: Diff + related files (3/5 against 1/5)
What people want from each one
jeinlee1991/chinese-llm-benchmark
Questions people ask
Is chinese-llm-benchmark better than teaql-agent-kit?
They split the axes: chinese-llm-benchmark leads on model freedom, teaql-agent-kit on context depth. chinese-llm-benchmark is worth picking when chinese-focused — pick this when you need the most extensive, regularly updated multi-domain Chinese LLM benchmark and a huge defect corpus for analysis.
Which of chinese-llm-benchmark and teaql-agent-kit keeps my code private?
chinese-llm-benchmark: Cloud API (your key) (3/5). teaql-agent-kit: Cloud via API key (3/5).
What does each one cost to run?
chinese-llm-benchmark: Requires Nonelinear API key (cloud service). teaql-agent-kit: Your API key or model costs may apply.
Full profiles: jeinlee1991/chinese-llm-benchmark and teaql/teaql-agent-kit. Everything else in Evals & benchmarks.