THUDM/AgentBench
A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)
View on GitHub →Where this sits in the catalog
Most-wanted open issues
Ranked by 👍 reactions. 65 open issues total, checked 2026-08-29.
| Issue | 👍 | Comments | Opened |
|---|---|---|---|
Benchmark for mistral models enhancement | 2 | 1 | 2 yrs |
| 2 | 1 | 2 yrs | |
cg和kg都遇到了Worker not responding bug, help wanted | 1 | 1 | 2 yrs |
[Bug/Assistance] A lot of os-std tasks are impossible bug, help wanted | 1 | 1 | 2 yrs |
| 1 | 0 | 1 yr |
What Hacker News made of it
Threads that link to or name this project, ranked by points. The comments are where the objections live — the README never argues back.
- AgentBench: Evaluating LLMs as Agents
2 points · 1 comments · 2 yrs
- Thudm/AgentBench: A Comprehensive Benchmark to Evaluate LLMs as Agents
1 points · 0 comments · 3 yrs
- AgentBench: A Comprehensive Benchmark to Evaluate LLMs as Agents
1 points · 0 comments · 3 yrs
Alternatives in Evals & benchmarks
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-…
A model-mediated harness for reliable agentic software development.
An open-source Harness Engineering platform for coding agents—define harnesses as code, run controlled experim…
Measuring frontier coding agents on original, long-horizon engineering tasks
Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 1…
An open-source game coding agent environment and benchmark for Godot.