Operative (web-eval-agent)
Frameworks & Eval · tested 2026-08-21 · re-test due 2026-11-21 · by the Hlido desk, not the vendor
In short: A browser agent that lets your coding agent vibe-test its own web changes over MCP — a genuinely useful 'let the coding agent debug itself' loop, YC-backed with a one-line install.
4 PASS · 0 FAIL of 4 public-surface claims
Quick answer
Operative (web-eval-agent) scores 72/100 (STEADY) on Hlido’s independent, hands-on test (reviewed 2026-08-21). STEADY (72) for a well-targeted, agent-native testing tool that closes a real loop (coding agent verifies its own web changes in a browser via MCP), with a one-line install, YC backing and ~1240 stars indicating traction
Operative's web-eval-agent gives a coding agent a browser agent it can call over MCP to end-to-end test the web app it just changed: navigate flows (login, dashboard, API-key creation), capture network traffic (all requests/responses in real time), and autonomously debug by driving the app like a user. It's a sharp answer to a real gap — coding agents write changes confidently but rarely verify them in a running browser — and the framing ('let the coding agent debug itself') is exactly right for the agent-to-agent thesis. Install is a single curl-pipe-bash line, it's Y-Combinator-backed, and the ~1240 GitHub stars suggest real traction. Honestly, it's built on browser-use ('we hooked browseruse up to our backend to make it 2x faster'), so it's a productized harness around an existing browser-automation engine rather than a from-scratch one — which is fine, but worth knowing. Surface limits: the autonomy and reliability of the debugging loop, and how well the network-capture and verification actually catch regressions, can't be judged from the landing page, and a curl | bash install warrants the usual caution. For agent-assisted web development it's one of the more directly useful MCP tools around.
Why STEADY
STEADY (72) for a well-targeted, agent-native testing tool that closes a real loop (coding agent verifies its own web changes in a browser via MCP), with a one-line install, YC backing and ~1240 stars indicating traction. Not higher because the decisive properties — reliability of the autonomous debug loop and how well it actually catches regressions — are unverifiable from the surface, and it is a harness layered on browser-use rather than a novel engine. Not FADING because the concept, integration path and traction signals are all real.
Public-surface checklist
- PASS Homepage loads (required)
- PASS Primary value prop (required) — A browser agent that lets your coding agent vibe-test its own web changes over M
- PASS Cta present
- PASS Evidence or demo — 1 screenshot(s) captured
What we saw
1 screenshot captured by the Hlido engine during the reviewed run (run-9ffa6e2cd0927645-www-operative-sh). Our own captures — not vendor marketing material.
What it does well
- Closes a genuinely missing loop: gives a coding agent a browser agent to end-to-end test the web changes it just made, over MCP
- Real-time network traffic capture (all requests/responses) for comprehensive debugging
- Autonomous flow testing (login, dashboard, API-key creation) driving the app like a user
- One-line install (curl | bash), Y-Combinator-backed, ~1240 GitHub stars indicating traction
What it fails at
- Built on browser-use ('we hooked browseruse up to our backend') — a productized harness rather than a novel engine
- Autonomy/reliability of the debug loop and its true regression-catching ability are unverifiable from the landing page
- curl | bash install pattern warrants the usual supply-chain caution
- Depth of assertions/coverage the browser agent actually performs isn't documented on the surface
Red flags
- 'Vibe-test' autonomy is the whole pitch, but the reliability and regression-catching depth of the debug loop are unverifiable from the surface — validate it finds the bugs you care about before trusting it as a gate.
- Installs via curl | bash; review the script before running in a sensitive environment.
Best for
- Developers using coding agents (Cursor, Claude Code) who want the agent to verify web changes in a real browser
- Teams wanting MCP-driven end-to-end smoke tests and network-capture debugging without writing a harness
- Agent pipelines that need a 'did my change actually work in the UI?' verification step
Not recommended for
- Teams needing a mature, deterministic test framework with documented coverage guarantees (this is exploratory 'vibe-testing')
- Environments where curl | bash installs are disallowed
- Non-web applications — it's a browser agent
Compared to
-
Browser Use
ready-made-agent-self-test-harness-vs-general-browser-automation-engine
browser-use is the general-purpose browser-automation engine agents use to drive web pages; Operative is a productized, MCP-native harness built on top of it, aimed specifically at letting a coding agent test and debug its own web changes (with network capture and one-line install). Choose browser-use to build your own automation, Operative for a ready-made agent self-testing loop.
Agent relevance
CLI MCP Behavioral-testable
Agentic-Commerce Readiness 63/100 · INTEGRABLE
Independent readiness for agent delegation & transaction. How it’s scored · check live
MCP server a coding agent calls to drive a browser: navigate flows, capture network traffic, and autonomously test/verify web changes. One-line install; works with MCP hosts (Cursor, Claude Code). Directly agent-to-agent — one agent verifying another's output — and testable given a target web app.
Agent-friendly score: 9/10
Score over time
The longitudinal record — every point is the score as published on that date. Raw series.
Evidence
- Browser agent that lets your coding agent vibe-test web applications via MCP ('let the coding agent debug itself') — source (2026-08-21) verified
- One-line install: curl -LSf https://operative.sh/install.sh -o install.sh && bash install.sh — source (2026-08-21) verified
- Network traffic capture: monitor all requests/responses in real time for debugging — source (2026-08-21) verified
- Autonomous debugging: browser-use agent tests and verifies the app end-to-end — source (2026-08-21) verified
- Built on browser-use ('hooked browseruse up to our backend to make it 2x faster'); Y-Combinator-backed; ~1240 GitHub stars — source (2026-08-21) verified
