Toolbox/Evals & benchmarks/jeinlee1991/chinese-llm-benchmark

jeinlee1991/chinese-llm-benchmark

非线智能 NoneLinear - ReLE评测:中文AI大模型能力评测(持续更新):目前已囊括374个大模型,覆盖chatgpt、gpt-5.4、谷歌gemini-3.1-pro、Claude-4.6、文心ERNIE-X1.1、ERNIE-5.0、qwen3.6-max、qwen3.6-plus、百川、讯飞星火、商汤senseChat等商用模型, 以及step3.5-flash、kimi-k2.6、ernie4.5、MiniMax-M2.7、deepseek-v4、Qwen3.6、llama4、智谱GLM-5.1、MiMo-V2、LongCat、gemma4、mistral等开源大模型。不仅提供排行榜,也提供规模超200万的大模型缺陷库!方便广大社区研究分析、改进大模型。

View on GitHub →
agentic-aiartificial-intelligencellm-agentllm-evaluation
Stars
6.4k
Δ 7d
+11
Δ 30d
+67
Age
3 yrs
Last push
1d ago
6.4k6.4k6.3k6.2k6.1kJun 8Jul 21Sep 4
90-day star history

Where this sits in the catalog

Most-wanted open issues

Ranked by 👍 reactions. 16 open issues total, checked 2026-08-29.

Head to head with its peers

The same six axes and the same star history, two projects at a time.

Alternatives in Evals & benchmarks

A Comprehensive Benchmark to Evaluate LLMs as Agents (ICLR'24)

3.7k
+107d
Python
3 yrs

A model-mediated harness for reliable agentic software development.

2.8k
+37d
Python
7 yrs

An open-source Harness Engineering platform for coding agents—define harnesses as code, run controlled experim…

2.2k
+1247d
JavaScript
1 mo

Measuring frontier coding agents on original, long-horizon engineering tasks

1.6k
+817d
Python
3 mo

Measuring how well CLI agents like Claude Code or Codex CLI can post-train base LLMs on a single H100 GPU in 1…

546
+97d
Python
9 mo

An open-source game coding agent environment and benchmark for Godot.

536
-37d
JavaScript
2 wk