We hands-on tested 26 AI research agents. 62% are fading.
Hlido independently tested 26 AI research agents — literature review, deep-research, and answer-engine tools. Only 2 are VITAL; 62% are FADING under a real test.
By the Hlido Editor · 2026-08-03
Everyone is shipping a "research agent." Most of them don't hold up.
Hlido independently tested 26 AI research agents — literature-review tools, deep-research assistants, and answer engines that claim to find, synthesize, or cite sources for you. Each one gets one evidence-backed verdict: a score from 0 to 100, hands-on, claim by claim. No vendor surveys, no self-reported benchmarks.
Here's what the Research category looks like right now.
The distribution is bottom-heavy
- VITAL (90+): 8% — 2 of 26. Genuinely deliver, evidence backs the claims.
- STEADY (70–89): 31% — 8 of 26. Solid, with caveats.
- FADING (40–69): 62% — 16 of 26. The largest group by far. The product exists, the pitch is confident, but the core promise wobbles under a real test.
Nearly two in three research agents we tested fall into FADING. That's a worse split than the AI-agent category overall (46% FADING across the full corpus) — research is a harder job to fake than it looks, and most tools built for it haven't earned the claim yet.
Where it actually works
Only two agents cleared the VITAL bar in this category: Exa and Elicit, both scoring 90. Both are narrow by design — Exa is a search API built specifically for retrieval, Elicit is scoped to literature review and evidence synthesis. Neither tries to be a general-purpose "AI researcher."
That tracks with a pattern Hlido sees across categories: tools that do one well-scoped research job — retrieval, citation synthesis, literature screening — hold up better than tools promising to "research anything for you."
The takeaway
If a research agent's pitch is broad ("ask me anything, I'll find the answer"), demand evidence before you trust it — the corpus says 62% of the category doesn't deliver on that promise under a real test. If the pitch is narrow and specific, it's more likely to be one of the 39% that does.
Methodology: every agent is tested hands-on and scored on a private rubric; outcomes, claim audits, and signed evidence are public. Explore all Research-category verdicts at hlido.eu/reviews, or query them over MCP at hlido.eu/mcp.