Evals & benchmarks
evaluation harnesses and benchmarks
jeinlee1991/chinese-llm-benchmark jeinlee1991/chinese-llm-benchmark leads the category by adoption and scale—it’s a Chinese‑first, large benchmark with a >2M‑entry defect database; the single most useful thing to know is to pick a suite that matches your target workload (language, interactive agents, coding/auditability, or privacy/self‑hosting) because no one project covers all needs.
Which one matches your setup?
Answer any of the questions — the shortlist updates as you go. Recommendations come from the capability passports below, nothing else.
Comparison matrix
Axes are extracted from each project's docs by our review pipeline; the maturity score is computed from stars, growth and commit activity — not an opinion. Click a column to sort.
| Runs in | Models | Context | Cost to run | |||||
|---|---|---|---|---|---|---|---|---|
⭐ 2.2k +131/7d | CLIIDECoding-agent plugin | Fixed provider | Whole-repo analysis | Evidence-backed findings | ●●●●● | Self-hostable | ●●●●● | Depends on host and agent (may use paid services) |
⭐ 6.4k +14/7d | Web app | Fixed provider | Diff only | Basic scoring filters | ●●●●● | Cloud API (your key) | ●●●●● | Requires Nonelinear API key (cloud service) |
⭐ 2.8k +4/7d | CLICI | Fixed provider | Diff + related files | Guides & checkpoints | ●●●●● | Cloud via API key | ●●●●● | Your API key or model costs may apply |
⭐ 1.6k +85/7d | CLICI | BYOKOpenAIAnthropicGemini | Whole-repo analysis | Behavioral verification | ●●●●● | Cloud via API key | ●●●●● | Your API key (OpenAI/Anthropic), pay-per-run model costs |
⭐ 546 +9/7d | CLICI | BYOKOpenAIAnthropicGeminiFixed provider | Whole-repo access | Judge + rules | ●●●●● | APIs with your key | ●●●●● | Your API keys + H100 GPU (cluster or rented) |
⭐ 534 +57/7d | CLICoding-agent plugin | Fixed provider | Single-file analysis | Severity prioritization | ●●●●● | Fully local | ●●●●● | Free, local |
⭐ 181 +11/7d | CLICI | BYOKOpenAIAnthropicGeminiFixed provider | Task-level analysis | Basic review tooling | ●●●●● | Self-hostable | ●●●●● | Your API key + upstream inference costs |
⭐ 3.7k +12/7d | CLI | OpenAIFixed provider | Env/task only | None mentioned | ●●●●● | API key cloud | ●●●●● | Your OpenAI API key required |
⭐ 536 | CLICI | Fixed provider | Whole-repo analysis | Verifier & tests | ●●●●● | Self-hostable | ●●●●● | Free, self-hosted (Docker required). |
⭐ 140 +1/7d | CLICI | BYOKOpenAIAnthropicGeminiFixed provider | Episode + related files | Basic scoring | ●●●●● | Cloud API keys | ●●●●● | Your API key for cloud models |
⭐ 53 +1/7d | CLICI | OpenAIAnthropicFixed provider | Task-only | Pytest verifier | ●●●●● | Cloud APIs (your key) | ●●●●● | Your API key (OpenAI/Anthropic/OpenRouter); local Docker image required |
⭐ 44 | CLI | BYOKOpenAI | Whole sandbox | Evaluation-only checks | ●●●●● | API-key cloud | ●●●●● | Your API key; large dataset download (~43 GB) |
Capability profiles
Six axes, 0–5 each. The shape tells you the strategy: a wide hexagon is a generalist, a spike is a specialist. Showing the 8 most established — the rest are in the full catalog.
Reviews AI coding workflows end-to-end and produces evidence-bounded, prioritized findings across five Agent Work Loop dimensions rather than only analyzing final diffs.
Chinese-first, large-scale benchmark that covers hundreds of commercial and open-source models and ships a >2M-entry defect database for model diagnosis.
Provides a TEAQL-focused, auditable evaluation harness that measures software-engineering discipline and token-efficiency for coding agents rather than offering general-purpose agent automation.
Provides long-horizon, behaviorally-graded software-engineering tasks with isolated sandbox execution and separate verifier environments (via Pier).
Measures autonomous CLI agents' ability to post-train base LLMs within a 10‑hour H100 budget, evaluating agent-driven R&D rather than only inference.
Provides modular, reusable agent 'skills' (not just benchmarks), including an error-discovery skill that builds a dependency-free single-file review app for guided failure-mode analysis.
Adds explicit support for installed agents in air‑gapped sandboxes and an augmented ATIF v1.7 to produce more faithful, consistent agent trajectories.
Provides an end-to-end, multi-environment benchmark specifically for evaluating LLMs as interactive agents with containerized task workers and function-calling support.
All repositories (12)
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-…
A model-mediated harness for reliable agentic software development.
An open-source Harness Engineering platform for coding agents—define harnesses as code, run controlled experim…
Measuring frontier coding agents on original, long-horizon engineering tasks
Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 1…
An open-source game coding agent environment and benchmark for Godot.
Skills that guide AI coding agents to help you build product-specific AI evals.
Pier is a Harbor fork built for DeepSWE, with stronger support for CLI agents in air-gapped (no-internet) task…
GateMem: a benchmark and evaluation toolkit for memory governance in multi-principal shared-memory LLM agents.
The benchmark tasks and evaluation harness for "PhysicianBench: Evaluating LLM Agents in Real-World EHR Enviro…
CoDA-Bench is a benchmark for code agents on data-intensive tasks. 🎈代码智能体能搞定数据密集型任务吗?
Hosted alternatives
If running your own reviewer is more ops than you want, these managed services cover the same job.
A hosted evals platform for LLM apps: run scored test suites, track regressions and compare prompts in the cloud rather than maintaining a benchmark harness of your own.
Try Braintrust →LangChain's managed platform for evaluation and tracing, so datasets, graders and run comparisons are hosted instead of assembled and operated by you.
Try LangSmith →The hosted side of open-source Langfuse: evals, datasets and scoring for agents run for you, an alternative to self-hosting the same stack.
Try Langfuse Cloud →