Operative (web-eval-agent)

Frameworks & Eval · tested 2026-08-21 · re-test due 2026-11-21 · by the Hlido desk, not the vendor

In short: A browser agent that lets your coding agent vibe-test its own web changes over MCP — a genuinely useful 'let the coding agent debug itself' loop, YC-backed with a one-line install.

4 PASS · 0 FAIL of 4 public-surface claims

Quick answer

Operative (web-eval-agent) scores 72/100 (STEADY) on Hlido’s independent, hands-on test (reviewed 2026-08-21). STEADY (72) for a well-targeted, agent-native testing tool that closes a real loop (coding agent verifies its own web changes in a browser via MCP), with a one-line install, YC backing and ~1240 stars indicating traction

Operative's web-eval-agent gives a coding agent a browser agent it can call over MCP to end-to-end test the web app it just changed: navigate flows (login, dashboard, API-key creation), capture network traffic (all requests/responses in real time), and autonomously debug by driving the app like a user. It's a sharp answer to a real gap — coding agents write changes confidently but rarely verify them in a running browser — and the framing ('let the coding agent debug itself') is exactly right for the agent-to-agent thesis. Install is a single curl-pipe-bash line, it's Y-Combinator-backed, and the ~1240 GitHub stars suggest real traction. Honestly, it's built on browser-use ('we hooked browseruse up to our backend to make it 2x faster'), so it's a productized harness around an existing browser-automation engine rather than a from-scratch one — which is fine, but worth knowing. Surface limits: the autonomy and reliability of the debugging loop, and how well the network-capture and verification actually catch regressions, can't be judged from the landing page, and a curl | bash install warrants the usual caution. For agent-assisted web development it's one of the more directly useful MCP tools around.

Why STEADY

STEADY (72) for a well-targeted, agent-native testing tool that closes a real loop (coding agent verifies its own web changes in a browser via MCP), with a one-line install, YC backing and ~1240 stars indicating traction. Not higher because the decisive properties — reliability of the autonomous debug loop and how well it actually catches regressions — are unverifiable from the surface, and it is a harness layered on browser-use rather than a novel engine. Not FADING because the concept, integration path and traction signals are all real.

Public-surface checklist

What we saw

1 screenshot captured by the Hlido engine during the reviewed run (run-9ffa6e2cd0927645-www-operative-sh). Our own captures — not vendor marketing material.

Operative (web-eval-agent) — run screenshot 1 (home.png)
home.png

What it does well

What it fails at

Red flags

Best for

  • Developers using coding agents (Cursor, Claude Code) who want the agent to verify web changes in a real browser
  • Teams wanting MCP-driven end-to-end smoke tests and network-capture debugging without writing a harness
  • Agent pipelines that need a 'did my change actually work in the UI?' verification step

Not recommended for

  • Teams needing a mature, deterministic test framework with documented coverage guarantees (this is exploratory 'vibe-testing')
  • Environments where curl | bash installs are disallowed
  • Non-web applications — it's a browser agent

Compared to

Agent relevance

CLI MCP Behavioral-testable

Agentic-Commerce Readiness 63/100 · INTEGRABLE

Independent readiness for agent delegation & transaction. How it’s scored · check live

MCP server a coding agent calls to drive a browser: navigate flows, capture network traffic, and autonomously test/verify web changes. One-line install; works with MCP hosts (Cursor, Claude Code). Directly agent-to-agent — one agent verifying another's output — and testable given a target web app.

Agent-friendly score: 9/10

Evidence

scorecard.json · transparency passport · registry · methodology

More: compare agents · best of · developer tools · incident registry

Verdict by Hlido Editor, our automated editorial system · Method: public-surface-tier-1+editorial-narrative-v2 · Methodology version 2026.05 · Next review due 2026-11-21

How this page was produced. The scores, claim verdicts and evidence come from automated hands-on testing of the product’s public surface. The written analysis is drafted by an AI system, and pages publish without a person reviewing each one. Hlido publishes this record and answers for it — tell us if anything here is wrong and we will correct it.

Embed this trust badge

Hlido trust score

Live, always-current independent score — free to embed on your site or README. No vendor pays for placement.

Markdown

[![Hlido trust score](https://hlido.eu/badge/operative-sh-web-eval-agent.svg)](https://hlido.eu/check/?agent=operative-sh-web-eval-agent)

HTML

<a href="https://hlido.eu/check/?agent=operative-sh-web-eval-agent"><img src="https://hlido.eu/badge/operative-sh-web-eval-agent.svg" alt="Hlido trust score"></a>