ai-evals-course/evals-skills
Skills that guide AI coding agents to help you build product-specific AI evals.
View on GitHub →Stars
534
Δ 7d
+57
Δ 30d
—
Age
2 mo
Last push
3d ago
90-day star history
Where this sits in the catalog
RankingsNot in any top 20 today — see the rankings
Most-wanted open issues
Ranked by 👍 reactions. 1 open issues total, checked 2026-09-01.
| Issue | 👍 | Comments | Opened |
|---|---|---|---|
| 0 | 0 | 1 wk |
Alternatives in Evals & benchmarks
非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-…
⭐6.4k
+147d
3 yrs
A model-mediated harness for reliable agentic software development.
⭐2.8k
+47d
Python
7 yrs
An open-source Harness Engineering platform for coding agents—define harnesses as code, run controlled experim…
⭐2.2k
+1317d
JavaScript
1 mo
Measuring frontier coding agents on original, long-horizon engineering tasks
⭐1.6k
+857d
Python
3 mo
Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 1…
⭐546
+97d
Python
9 mo