jeinlee1991/chinese-llm-benchmark
Frameworks & Eval · tested 2026-08-27 · by the Hlido desk, not the vendor
In short: A continuously-updated Chinese-language LLM leaderboard (ReLE, formerly CLiB) covering 380+ commercial and open models, plus a large model-defect library — a reference resource, not an agent you run.
2 PASS · 3 FAIL of 5 public-surface claims
Quick answer
jeinlee1991/chinese-llm-benchmark scores 70/100 (STEADY) on Hlido’s independent, hands-on test (reviewed 2026-08-27). STEADY (70). Pricing was not findable on the public surface when tested.
ReLE (Really Reliable Live Evaluation, formerly CLiB) is a substantial, continuously-updated evaluation resource: a Chinese-language capability leaderboard spanning 380+ models — GPT, Gemini, Claude, ERNIE, Qwen, DeepSeek, GLM and more — alongside a large-scale defect library the project says exceeds two million entries for community analysis. Its value is as a reference and research artifact, and the breadth plus the "live/continuously updated" posture is its main strength. The honest limits are inherent to the format: the methodology and scoring behind the rankings are the maintainers' own, non-independent choices, and there is no programmatic interface here for an agent to consume — it is a leaderboard to read, not an API to call. Useful primarily for teams selecting models for Chinese-language workloads who want a single continuously-maintained comparison point.
Why STEADY
STEADY (70). Genuine breadth (380+ models), a continuously-updated posture and a sizeable defect library clear STEADY as a reference resource. VITAL is withheld because it is a maintainer-curated leaderboard with no independent methodology guarantee, no programmatic surface, and no hands-on verification of the underlying scores.
Public-surface checklist
- PASS Homepage loads (required)
- PASS Primary value prop (required)
- FAIL Cta present (required)
- FAIL Pricing or access
- FAIL Evidence or demo
What we saw
1 screenshot captured by the Hlido engine during the reviewed run (run-1c0a5e1e0687488f-github-com). Our own captures — not vendor marketing material.
What it does well
- Continuously-updated leaderboard covering 380+ Chinese and international models
- Includes both commercial and open-source models in one comparison
- Ships a large-scale model-defect library for community research
- Fills a real gap for Chinese-language LLM capability comparison
What it fails at
- Rankings rest on the maintainers' own methodology — not independently audited
- No programmatic/API surface for agents to consume the data
- Underlying scores unverified by any hands-on check here
Best for
- Teams selecting models for Chinese-language workloads
- Researchers wanting a broad, maintained model comparison and defect corpus
Not recommended for
- Anyone needing an independently-audited benchmark methodology
- Agent pipelines wanting to query results programmatically
Pricing & access
- Pricing findable on the public surfaceFAIL
Derived from Hlido-held evidence only (engine checklist + editorial text); quotes are verbatim from the scorecard; not vendor-supplied; re-derived daily. Verify current prices on the vendor's pricing page. Last verified 2026-06-14.
Related agents
Agent relevance
No programmatic surfaces
Agentic-Commerce Readiness 16/100 · CLOSED
Independent readiness for agent delegation & transaction. How it’s scored · check live
A published leaderboard and defect dataset consumed by people reading the rankings; no programmatic interface for an agent to call. No hands-on behavioral test performed.
Agent-friendly score: 2/10
Score over time
The longitudinal record — every point is the score as published on that date. Raw series.
