Toolbox/Evals & benchmarks/THUDM/AgentBench

THUDM/AgentBench

A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)

View on GitHub →
PythonApache-2.0chatgptgpt-4llmllm-agent
Stars
3.7k
Δ 7d
+12
Δ 30d
+69
Age
3 yrs
Last push
207d ago
3.7k3.7k3.6k3.5k3.4kJun 8Jul 24Sep 4
90-day star history

Where this sits in the catalog

Most-wanted open issues

Ranked by 👍 reactions. 65 open issues total, checked 2026-08-29.

What Hacker News made of it

Threads that link to or name this project, ranked by points. The comments are where the objections live — the README never argues back.

Alternatives in Evals & benchmarks

非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-…

6.4k
+147d
3 yrs

A model-mediated harness for reliable agentic software development.

2.8k
+47d
Python
7 yrs

An open-source Harness Engineering platform for coding agents—define harnesses as code, run controlled experim…

2.2k
+1317d
JavaScript
1 mo

Measuring frontier coding agents on original, long-horizon engineering tasks

1.6k
+857d
Python
3 mo

Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 1…

546
+97d
Python
9 mo

An open-source game coding agent environment and benchmark for Godot.

536
JavaScript
2 wk