Tool Definition Quality Score (TDQS)
Frameworks & Eval · tested 2026-08-09 · re-test due 2026-11-09 · by the Hlido desk, not the vendor
In short: An open-sourced six-dimension score for MCP tool descriptions, grounded in two cited studies rather than opinion — reviewed here from its announcement post, which is the only surface that resolved.
4 PASS · 3 FAIL of 7 public-surface claims
Quick answer
Tool Definition Quality Score (TDQS) scores 73/100 (STEADY) on Hlido’s independent, hands-on test (reviewed 2026-08-09). STEADY (73) for a measure that starts from cited, quantified research rather than opinion, decomposes into six inspectable dimensions instead of one opaque number, targets a genuine failure mode (tool selection is driven
TDQS scores every tool on an MCP server across six dimensions — purpose clarity, usage guidelines, behavioural transparency, parameter semantics, conciseness and contextual completeness — at 1–5 each, rolled up into a tier. What makes it worth attention is that it is anchored in published research rather than in the author's taste: it cites 'MCP Tool Descriptions Are Smelly' (SAIL Research, arXiv 2602.14878), which found quality defects in 97% of 856 tools across 103 servers, with 56% not clearly stating what the tool does and 89% failing to say when it should or should not be used; and 'From Docs to Descriptions' (arXiv 2602.18914), a study of 10,831 servers reporting that well-described tools are selected 260% more often in competitive settings and that fixing descriptions lifts task success by about 6 percentage points. That is the correct order of operations — establish the problem is real and quantified, then build the measure. The framework has since been open-sourced, which matters for a scoring system: a score nobody can inspect is a ranking, not a measurement, and Hlido has no standing to say otherwise about anyone else's methodology if it will not credit that distinction. Two caveats on this review. First, the surface that resolved is the April 2026 announcement post, not the framework repository or a live scoring endpoint, so what is assessed is the design and its justification, not the implementation. Second, TDQS is operated by Glama, a directory that also ranks the servers it scores; the post does not discuss that conflict, and a scorer with commercial standing in what it scores should address it.
Why STEADY
STEADY (73) for a measure that starts from cited, quantified research rather than opinion, decomposes into six inspectable dimensions instead of one opaque number, targets a genuine failure mode (tool selection is driven entirely by the description an agent reads), and has been open-sourced so the method can be checked. Held below VITAL because the only surface that resolved is the announcement post rather than the framework or a live scorer — the implementation is unverified — and because the operator also ranks the servers it scores without addressing that conflict.
Public-surface checklist
- PASS Homepage loads (required)
- PASS Primary value prop (required) — Six-dimension quality score for MCP tool definitions
- FAIL Cta present (required) — Announcement post — no install, signup or run path on the captured page
- PASS Pricing or access — Framework stated as open-sourced
- PASS Claims cited — Two arXiv studies cited with identifiers and figures
- FAIL Conflict disclosed — Operator also ranks the servers scored; not addressed
- FAIL Third party validation — No calibration or reliability evidence for the score
What we saw
2 screenshots captured by the Hlido engine during the reviewed run (run-af80ddab61879473-glama-ai). Our own captures — not vendor marketing material.
What it does well
- Anchored in two cited studies rather than in the author's judgement
- Quantifies the problem first: 97% of 856 tools across 103 servers carry at least one description defect
- Cites a measured impact — well-described tools selected 260% more often; ~6pp task-success lift from fixing descriptions
- Six named dimensions scored 1–5 each, so a bad score decomposes into what to fix
- Framework has been open-sourced, so the method can be inspected rather than trusted
- Targets the right thing — the description is the only signal an agent reads when choosing a tool
- Publicly renamed 'Description' to 'Definition' to describe what is actually measured
What it fails at
- Reviewed from an announcement post — the framework repository and any live scoring surface were not the page that resolved
- Operated by a directory that also ranks the servers it scores, a conflict the post does not address
- No inter-rater reliability, calibration or validation of the score itself against downstream outcomes
- Scoring appears to be LLM-graded, with no discussion of variance between runs
- No adoption evidence — how many servers are scored, or whether authors act on it
Red flags
- TDQS is operated by Glama, which also lists and ranks the MCP servers being scored. The announcement does not address that conflict of interest.
Best for
- MCP server authors who want specific, actionable defects rather than 'write better docs'
- Teams choosing between servers offering similar functionality
- Anyone building agent tool-selection who needs a vocabulary for description quality
Not recommended for
- Anyone needing a validated, calibrated metric — no reliability evidence is published
- Users who want a score free of operator conflict of interest
- Assessing anything beyond the tool definition, which is all this measures
Pricing & access
- Pricing findable on the public surfacePASS Framework stated as open-sourced (tested 2026-08-06)
Derived from Hlido-held evidence only (engine checklist + editorial text); quotes are verbatim from the scorecard; not vendor-supplied; re-derived daily. Verify current prices on the vendor's pricing page. Last verified 2026-08-06.
Related agents
Agent relevance
No programmatic surfaces
Agentic-Commerce Readiness 32/100 · SURFACE-ONLY
Independent readiness for agent delegation & transaction. How it’s scored · check live
Not an agent-callable tool. TDQS is a scoring framework applied to MCP servers; its output is consumed by server authors improving definitions and by users comparing servers. The open-sourced framework could be run in CI, but no agent-facing endpoint is described on the captured surface.
Agent-friendly score: 3/10
Score over time
The longitudinal record — every point is the score as published on that date. Raw series.
Evidence
- Six dimensions — purpose clarity, usage guidelines, behavioural transparency, parameter semantics, conciseness, contextual completeness — scored 1–5 and rolled into a tier — source (2026-08-09) verified
- Cites SAIL Research arXiv 2602.14878: 97% of 856 tools across 103 servers have at least one defect; 56% unclear purpose; 89% no usage guidance — source (2026-08-09) verified
- Cites arXiv 2602.18914 over 10,831 servers: 260% more selection for well-described tools; ~6pp task-success improvement — source (2026-08-09) verified
- The TDQS framework has been open-sourced — source (2026-08-09) verified
- Renamed from 'Tool Description Quality Score' to 'Tool Definition Quality Score' — source (2026-08-09) verified
- Validation, calibration or inter-rater reliability of the score itself — source (2026-08-09)
- Disclosure of the operator's commercial interest in the servers it scores — source (2026-08-09)

