jeinlee1991/chinese-llm-benchmark

Frameworks & Eval · tested 2026-08-27 · by the Hlido desk, not the vendor

In short: A continuously-updated Chinese-language LLM leaderboard (ReLE, formerly CLiB) covering 380+ commercial and open models, plus a large model-defect library — a reference resource, not an agent you run.

2 PASS · 3 FAIL of 5 public-surface claims

Quick answer

jeinlee1991/chinese-llm-benchmark scores 70/100 (STEADY) on Hlido’s independent, hands-on test (reviewed 2026-08-27). STEADY (70). Pricing was not findable on the public surface when tested.

ReLE (Really Reliable Live Evaluation, formerly CLiB) is a substantial, continuously-updated evaluation resource: a Chinese-language capability leaderboard spanning 380+ models — GPT, Gemini, Claude, ERNIE, Qwen, DeepSeek, GLM and more — alongside a large-scale defect library the project says exceeds two million entries for community analysis. Its value is as a reference and research artifact, and the breadth plus the "live/continuously updated" posture is its main strength. The honest limits are inherent to the format: the methodology and scoring behind the rankings are the maintainers' own, non-independent choices, and there is no programmatic interface here for an agent to consume — it is a leaderboard to read, not an API to call. Useful primarily for teams selecting models for Chinese-language workloads who want a single continuously-maintained comparison point.

Why STEADY

STEADY (70). Genuine breadth (380+ models), a continuously-updated posture and a sizeable defect library clear STEADY as a reference resource. VITAL is withheld because it is a maintainer-curated leaderboard with no independent methodology guarantee, no programmatic surface, and no hands-on verification of the underlying scores.

Public-surface checklist

What we saw

1 screenshot captured by the Hlido engine during the reviewed run (run-1c0a5e1e0687488f-github-com). Our own captures — not vendor marketing material.

jeinlee1991/chinese-llm-benchmark — run screenshot 1 (home.png)
home.png

What it does well

What it fails at

Best for

  • Teams selecting models for Chinese-language workloads
  • Researchers wanting a broad, maintained model comparison and defect corpus

Not recommended for

  • Anyone needing an independently-audited benchmark methodology
  • Agent pipelines wanting to query results programmatically

Pricing & access

Derived from Hlido-held evidence only (engine checklist + editorial text); quotes are verbatim from the scorecard; not vendor-supplied; re-derived daily. Verify current prices on the vendor's pricing page. Last verified 2026-06-14.

Related agents

Agent relevance

No programmatic surfaces

Agentic-Commerce Readiness 16/100 · CLOSED

Independent readiness for agent delegation & transaction. How it’s scored · check live

A published leaderboard and defect dataset consumed by people reading the rankings; no programmatic interface for an agent to call. No hands-on behavioral test performed.

Agent-friendly score: 2/10

scorecard.json · transparency passport · registry · methodology

More: compare agents · best of · developer tools · incident registry

Verdict by Hlido Editor, our automated editorial system · Method: public-surface-tier-2+editorial-narrative-v2 · Methodology version 2026.05 ·

How this page was produced. The scores, claim verdicts and evidence come from automated hands-on testing of the product’s public surface. The written analysis is drafted by an AI system, and pages publish without a person reviewing each one. Hlido publishes this record and answers for it — tell us if anything here is wrong and we will correct it.

Embed this trust badge

Hlido trust score

Live, always-current independent score — free to embed on your site or README. No vendor pays for placement.

Markdown

[![Hlido trust score](https://hlido.eu/badge/jeinlee1991-chinese-llm-benchmark.svg)](https://hlido.eu/check/?agent=jeinlee1991-chinese-llm-benchmark)

HTML

<a href="https://hlido.eu/check/?agent=jeinlee1991-chinese-llm-benchmark"><img src="https://hlido.eu/badge/jeinlee1991-chinese-llm-benchmark.svg" alt="Hlido trust score"></a>