Tool Definition Quality Score (TDQS)

Frameworks & Eval · tested 2026-08-09 · re-test due 2026-11-09 · by the Hlido desk, not the vendor

In short: An open-sourced six-dimension score for MCP tool descriptions, grounded in two cited studies rather than opinion — reviewed here from its announcement post, which is the only surface that resolved.

4 PASS · 3 FAIL of 7 public-surface claims

Quick answer

Tool Definition Quality Score (TDQS) scores 73/100 (STEADY) on Hlido’s independent, hands-on test (reviewed 2026-08-09). STEADY (73) for a measure that starts from cited, quantified research rather than opinion, decomposes into six inspectable dimensions instead of one opaque number, targets a genuine failure mode (tool selection is driven

TDQS scores every tool on an MCP server across six dimensions — purpose clarity, usage guidelines, behavioural transparency, parameter semantics, conciseness and contextual completeness — at 1–5 each, rolled up into a tier. What makes it worth attention is that it is anchored in published research rather than in the author's taste: it cites 'MCP Tool Descriptions Are Smelly' (SAIL Research, arXiv 2602.14878), which found quality defects in 97% of 856 tools across 103 servers, with 56% not clearly stating what the tool does and 89% failing to say when it should or should not be used; and 'From Docs to Descriptions' (arXiv 2602.18914), a study of 10,831 servers reporting that well-described tools are selected 260% more often in competitive settings and that fixing descriptions lifts task success by about 6 percentage points. That is the correct order of operations — establish the problem is real and quantified, then build the measure. The framework has since been open-sourced, which matters for a scoring system: a score nobody can inspect is a ranking, not a measurement, and Hlido has no standing to say otherwise about anyone else's methodology if it will not credit that distinction. Two caveats on this review. First, the surface that resolved is the April 2026 announcement post, not the framework repository or a live scoring endpoint, so what is assessed is the design and its justification, not the implementation. Second, TDQS is operated by Glama, a directory that also ranks the servers it scores; the post does not discuss that conflict, and a scorer with commercial standing in what it scores should address it.

Why STEADY

STEADY (73) for a measure that starts from cited, quantified research rather than opinion, decomposes into six inspectable dimensions instead of one opaque number, targets a genuine failure mode (tool selection is driven entirely by the description an agent reads), and has been open-sourced so the method can be checked. Held below VITAL because the only surface that resolved is the announcement post rather than the framework or a live scorer — the implementation is unverified — and because the operator also ranks the servers it scores without addressing that conflict.

Public-surface checklist

What we saw

2 screenshots captured by the Hlido engine during the reviewed run (run-af80ddab61879473-glama-ai). Our own captures — not vendor marketing material.

Tool Definition Quality Score (TDQS) — run screenshot 1 (home.png)
home.png
Tool Definition Quality Score (TDQS) — run screenshot 2 (page_main-content.png)
page_main-content.png

What it does well

What it fails at

Red flags

Best for

  • MCP server authors who want specific, actionable defects rather than 'write better docs'
  • Teams choosing between servers offering similar functionality
  • Anyone building agent tool-selection who needs a vocabulary for description quality

Not recommended for

  • Anyone needing a validated, calibrated metric — no reliability evidence is published
  • Users who want a score free of operator conflict of interest
  • Assessing anything beyond the tool definition, which is all this measures

Pricing & access

Derived from Hlido-held evidence only (engine checklist + editorial text); quotes are verbatim from the scorecard; not vendor-supplied; re-derived daily. Verify current prices on the vendor's pricing page. Last verified 2026-08-06.

Related agents

Agent relevance

No programmatic surfaces

Agentic-Commerce Readiness 32/100 · SURFACE-ONLY

Independent readiness for agent delegation & transaction. How it’s scored · check live

Not an agent-callable tool. TDQS is a scoring framework applied to MCP servers; its output is consumed by server authors improving definitions and by users comparing servers. The open-sourced framework could be run in CI, but no agent-facing endpoint is described on the captured surface.

Agent-friendly score: 3/10

Evidence

scorecard.json · transparency passport · registry · methodology

More: compare agents · best of · developer tools · incident registry

Verdict by Hlido Editor, our automated editorial system · Method: public-surface-tier-2+editorial-narrative-v2 · Methodology version 2026.05 · Next review due 2026-11-09

How this page was produced. The scores, claim verdicts and evidence come from automated hands-on testing of the product’s public surface. The written analysis is drafted by an AI system, and pages publish without a person reviewing each one. Hlido publishes this record and answers for it — tell us if anything here is wrong and we will correct it.

Embed this trust badge

Hlido trust score

Live, always-current independent score — free to embed on your site or README. No vendor pays for placement.

Markdown

[![Hlido trust score](https://hlido.eu/badge/glama-ai-tool-definition-quality-score.svg)](https://hlido.eu/check/?agent=glama-ai-tool-definition-quality-score)

HTML

<a href="https://hlido.eu/check/?agent=glama-ai-tool-definition-quality-score"><img src="https://hlido.eu/badge/glama-ai-tool-definition-quality-score.svg" alt="Hlido trust score"></a>