Dingo

Frameworks & Eval · tested 2026-09-05 · re-test due 2026-12-04 · by the Hlido desk, not the vendor

In short: Comprehensive open-source data/model/application quality-evaluation toolkit that spans rule-based checks, LLM-as-a-judge and Agent-as-a-judge, including agent-trace evaluation — genuinely on-thesis for the agent era, reviewed here only at the surface.

5 PASS · 0 FAIL of 5 public-surface claims

Quick answer

Dingo scores 72/100 (STEADY) on Hlido’s independent, hands-on test (reviewed 2026-09-05). STEADY (72): a broad, clearly-documented, open-source evaluation toolkit that explicitly covers agent-trace evaluation, with real traction and a coherent architecture. Pricing: Open source (free entry point documented).

Dingo positions itself as a 'data quality inspection assistant' that scales from single-sample checks up to continuous agent-trace evaluation, and the public surface describes a coherent three-mode design: Quick Try for fast validation, Full Pipeline for batch dataset evaluation, and an Agent Evaluation mode for online monitoring of agent traces (task completion, tool usage, plan adherence, error recovery, latency and token trends). That last mode is what makes it interesting to Hlido's readership — evaluating agents, not just training data, is exactly where the eval space is heading. The breadth is credible: 20+ built-in rules for dataset hygiene, LLM integration (OpenAI, Kimi, local Llama3) for semantic assessment, and a bundled HHEM model for RAG/hallucination consistency. It is open source with meaningful traction (750+ stars) and clear documentation of its architecture. The limits are the honest Tier-1 ones: the accuracy and usefulness of the judges, the real cost of running them, and how well the agent-trace evaluation performs are all things you learn by running it, not by reading the site. There is also a naming/homepage inconsistency worth a second look at re-review (README repo namespace vs. the DataEval/MigoXLab labels), which is cosmetic but worth confirming.

Why STEADY

STEADY (72): a broad, clearly-documented, open-source evaluation toolkit that explicitly covers agent-trace evaluation, with real traction and a coherent architecture. Held mid-band because judge accuracy, run cost and the agent-eval mode's real-world performance are untested at Tier-1, and the toolkit's breadth means depth-per-mode is unproven from the surface alone.

Public-surface checklist

What it does well

What it fails at

Best for

  • ML/LLM teams needing dataset-quality inspection before training or RAG
  • Agent teams that want to evaluate traces (tool use, plan adherence) continuously
  • Privacy-sensitive orgs that want to run judges against local models
  • Anyone standardising a data/model quality gate in a pipeline

Not recommended for

  • Teams wanting a fully hosted, zero-setup eval SaaS with SLAs
  • Buyers who need an independently benchmarked judge-accuracy guarantee up front
  • Non-technical users who can't operate a Python/pipeline toolkit

Pricing & access

Derived from Hlido-held evidence only (engine checklist + editorial text); quotes are verbatim from the scorecard; not vendor-supplied; re-derived daily. Verify current prices on the vendor's pricing page. Last verified 2026-09-05.

Compared to

Agent relevance

API CLI SDK Behavioral-testable

An evaluation toolkit that can score agent traces (tool usage, plan adherence, error recovery). Directly relevant as a quality gate in agent pipelines; behaviourally testable, though not exercised at Tier-1.

Agent-friendly score: 6/10

Evidence

scorecard.json · registry · methodology

More: compare agents · best of · developer tools · incident registry

Verdict by Hlido Editor, our automated editorial system · Method: public-surface-tier-1+editorial-narrative-v2 · Methodology version 2026.05 · Next review due 2026-12-04

How this page was produced. The scores, claim verdicts and evidence come from automated hands-on testing of the product’s public surface. The written analysis is drafted by an AI system, and pages publish without a person reviewing each one. Hlido publishes this record and answers for it — tell us if anything here is wrong and we will correct it.

Embed this trust badge

Hlido trust score

Live, always-current independent score — free to embed on your site or README. No vendor pays for placement.

Markdown

[![Hlido trust score](https://hlido.eu/badge/dataeval-dingo.svg)](https://hlido.eu/check/?agent=dataeval-dingo)

HTML

<a href="https://hlido.eu/check/?agent=dataeval-dingo"><img src="https://hlido.eu/badge/dataeval-dingo.svg" alt="Hlido trust score"></a>