Toolbox/Running agents in production

Evals & benchmarks

evaluation harnesses and benchmarks

12 repositories, most starred first
State of the category

jeinlee1991/chinese-llm-benchmark jeinlee1991/chinese-llm-benchmark leads the category by adoption and scale—it’s a Chinese‑first, large benchmark with a >2M‑entry defect database; the single most useful thing to know is to pick a suite that matches your target workload (language, interactive agents, coding/auditability, or privacy/self‑hosting) because no one project covers all needs.

Interactive environments
Teams building or evaluating agents that must act in realistic, instrumented environments need containerized tasks, function-calling and runtime feedback—capabilities many benchmarks still implement only partially.
Coding & auditability
Organizations that care about developer workflows, reproducible evidence, token-efficiency and software‑engineering metrics need evaluation harnesses focused on end‑to‑end coding loops and auditable findings rather than simple final-diff scoring.
Long‑horizon training
Groups researching agent-driven model development or long experiments need suites that support post‑training, extended compute budgets and separate verifier sandboxes rather than single-shot inference tests.
Domain & safety specialization
Regulated or safety‑critical workloads require domain‑specific scenarios, realistic data and governance checks (e.g., memory controls, clinical records) that general benchmarks don’t provide out of the box.
Find your fit

Which one matches your setup?

Answer any of the questions — the shortlist updates as you go. Recommendations come from the capability passports below, nothing else.

Where should reviews happen?
Can code leave your infrastructure?
What matters most?
Model access?
Pick at least one answer to get a shortlist.
Side by side

Comparison matrix

Axes are extracted from each project's docs by our review pipeline; the maturity score is computed from stars, growth and commit activity — not an opinion. Click a column to sort.

Runs inModelsContextCost to run
2.2k +131/7d
CLIIDECoding-agent pluginFixed providerWhole-repo analysisEvidence-backed findingsSelf-hostableDepends on host and agent (may use paid services)
6.4k +14/7d
Web appFixed providerDiff onlyBasic scoring filtersCloud API (your key)Requires Nonelinear API key (cloud service)
2.8k +4/7d
CLICIFixed providerDiff + related filesGuides & checkpointsCloud via API keyYour API key or model costs may apply
1.6k +85/7d
CLICIBYOKOpenAIAnthropicGeminiWhole-repo analysisBehavioral verificationCloud via API keyYour API key (OpenAI/Anthropic), pay-per-run model costs
546 +9/7d
CLICIBYOKOpenAIAnthropicGeminiFixed providerWhole-repo accessJudge + rulesAPIs with your keyYour API keys + H100 GPU (cluster or rented)
534 +57/7d
CLICoding-agent pluginFixed providerSingle-file analysisSeverity prioritizationFully localFree, local
181 +11/7d
CLICIBYOKOpenAIAnthropicGeminiFixed providerTask-level analysisBasic review toolingSelf-hostableYour API key + upstream inference costs
3.7k +12/7d
CLIOpenAIFixed providerEnv/task onlyNone mentionedAPI key cloudYour OpenAI API key required
536
CLICIFixed providerWhole-repo analysisVerifier & testsSelf-hostableFree, self-hosted (Docker required).
140 +1/7d
CLICIBYOKOpenAIAnthropicGeminiFixed providerEpisode + related filesBasic scoringCloud API keysYour API key for cloud models
53 +1/7d
CLICIOpenAIAnthropicFixed providerTask-onlyPytest verifierCloud APIs (your key)Your API key (OpenAI/Anthropic/OpenRouter); local Docker image required
44
CLIBYOKOpenAIWhole sandboxEvaluation-only checksAPI-key cloudYour API key; large dataset download (~43 GB)
At a glance

Capability profiles

Six axes, 0–5 each. The shape tells you the strategy: a wide hexagon is a generalist, a spike is a specialist. Showing the 8 most established — the rest are in the full catalog.

ContextNoiseCustomPrivacyModelsMaturity

Reviews AI coding workflows end-to-end and produces evidence-bounded, prioritized findings across five Agent Work Loop dimensions rather than only analyzing final diffs.

ContextNoiseCustomPrivacyModelsMaturity

Chinese-first, large-scale benchmark that covers hundreds of commercial and open-source models and ships a >2M-entry defect database for model diagnosis.

ContextNoiseCustomPrivacyModelsMaturity

Provides a TEAQL-focused, auditable evaluation harness that measures software-engineering discipline and token-efficiency for coding agents rather than offering general-purpose agent automation.

ContextNoiseCustomPrivacyModelsMaturity

Provides long-horizon, behaviorally-graded software-engineering tasks with isolated sandbox execution and separate verifier environments (via Pier).

ContextNoiseCustomPrivacyModelsMaturity

Measures autonomous CLI agents' ability to post-train base LLMs within a 10‑hour H100 budget, evaluating agent-driven R&D rather than only inference.

ContextNoiseCustomPrivacyModelsMaturity

Provides modular, reusable agent 'skills' (not just benchmarks), including an error-discovery skill that builds a dependency-free single-file review app for guided failure-mode analysis.

ContextNoiseCustomPrivacyModelsMaturity

Adds explicit support for installed agents in air‑gapped sandboxes and an augmented ATIF v1.7 to produce more faithful, consistent agent trajectories.

ContextNoiseCustomPrivacyModelsMaturity

Provides an end-to-end, multi-environment benchmark specifically for evaluating LLMs as interactive agents with containerized task workers and function-calling support.

All repositories (12)

非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-…

6.4k
+147d
3 yrs

A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)

3.7k
+127d
Python
3 yrs

A model-mediated harness for reliable agentic software development.

2.8k
+47d
Python
7 yrs

An open-source Harness Engineering platform for coding agents—define harnesses as code, run controlled experim…

2.2k
+1317d
JavaScript
1 mo

Measuring frontier coding agents on original, long-horizon engineering tasks

1.6k
+857d
Python
3 mo

Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 1…

546
+97d
Python
9 mo

An open-source game coding agent environment and benchmark for Godot.

536
JavaScript
2 wk

Skills that guide AI coding agents to help you build product-specific AI evals.

534
+577d
2 mo

Pier is a Harbor fork built for DeepSWE, with stronger support for CLI agents in air-gapped (no-internet) task…

181
+117d
Python
4 mo

GateMem: a benchmark and evaluation toolkit for memory governance in multi-principal shared-memory LLM agents.

140
+17d
Python
2 mo

The benchmark tasks and evaluation harness for "PhysicianBench: Evaluating LLM Agents in Real-World EHR Enviro…

53
+17d
Python
4 mo

CoDA-Bench is a benchmark for code agents on data-intensive tasks. 🎈代码智能体能搞定数据密集型任务吗?

44
Python
3 mo
Don't want to self-host?

Hosted alternatives

If running your own reviewer is more ops than you want, these managed services cover the same job.

Braintrust

A hosted evals platform for LLM apps: run scored test suites, track regressions and compare prompts in the cloud rather than maintaining a benchmark harness of your own.

Try Braintrust
LangSmith

LangChain's managed platform for evaluation and tracing, so datasets, graders and run comparisons are hosted instead of assembled and operated by you.

Try LangSmith
Langfuse Cloud

The hosted side of open-source Langfuse: evals, datasets and scoring for agents run for you, an alternative to self-hosting the same stack.

Try Langfuse Cloud

More in Running agents in production