agent-device
Coding · tested 2026-08-25 · re-test due 2026-11-23 · by the Hlido desk, not the vendor
In short: Playwright's snapshot-and-act model, ported to native mobile and TV — a token-efficient device-automation CLI built specifically for agents, from a team that knows the mobile tooling space.
4 PASS · 1 FAIL of 5 public-surface claims
Quick answer
agent-device scores 80/100 (STEADY) on Hlido’s independent, hands-on test (reviewed 2026-08-25). STEADY (80) because agent-device applies a proven, token-conscious automation model to an underserved surface (native mobile/TV), ships as a clean single-install CLI designed for agents, and comes from a team with real m Pricing: Open source (free entry point documented).
agent-device takes the pattern that made web agents practical — expose the page as a structured accessibility tree with stable references, let the agent read it, then act on those references — and brings it to native iOS, Android, TV and desktop apps. The insight is explicitly about tokens: instead of feeding an agent raw screenshots or verbose UI dumps, `snapshot -i` returns only the interactive elements with stable refs like @e2, keeping the context small enough that a real exploration loop stays affordable. From there the agent acts with either those refs or semantic selectors ('find Sign In click', 'find role button click'), which keeps generated flows readable and resilient to layout churn. It is a single global npm install, one mental model across simulators, physical QA devices and TV targets, and evidence (screenshots) is captured only when needed rather than on every step. Two things earn it real credibility: it is built by callstack, a well-known React Native consultancy, so the mobile-tooling competence is not in doubt; and the 4.2k GitHub stars signal genuine early traction for what is a narrow, technical tool. What keeps it in the middle of the STEADY band rather than the top is maturity and proof — the value proposition is clean but the surface leans on the concept demo rather than published reliability data across real apps, device farms, or CI, and there is no pricing or hosted-service story visible (it reads as an open-source CLI). For an agent that needs to drive a real mobile app, though, this is the right shape of tool and a rare one.
Why STEADY
STEADY (80) because agent-device applies a proven, token-conscious automation model to an underserved surface (native mobile/TV), ships as a clean single-install CLI designed for agents, and comes from a team with real mobile-tooling pedigree plus early traction (4.2k stars). It sits mid-band rather than higher because the public surface is concept-and-command led — it does not yet publish reliability evidence across real apps or a CI/device-farm story — and the commercial/support model is unstated.
Public-surface checklist
- PASS Homepage loads (required)
- PASS Primary value prop (required) — The mobile verification for AI Agents — device automation CLI for AI agents
- PASS Cta present (required) — Read the Docs / npm install -g agent-device
- FAIL Pricing or access — No pricing surfaced; reads as an open-source CLI installed via npm
- PASS Evidence or demo — Worked terminal examples of open/snapshot/press/fill/find; cross-platform command samples
What we saw
4 screenshots captured by the Hlido engine during the reviewed run (run-7469c725a338cb3a-agent-device-dev). Our own captures — not vendor marketing material.
What it does well
- Token-efficient by design: `snapshot -i` exposes only interactive elements with stable refs instead of raw screenshots or verbose dumps
- Semantic selectors ('find Sign In click') make agent-authored flows readable and resilient to layout changes
- One mental model across iOS, Android, TV and desktop — simulators, physical QA devices and TV targets alike
- Purpose-built for agents rather than a human test framework with an agent wrapper
- Single global npm install; evidence captured only when needed
- Built by callstack, a recognised React Native tooling team — real domain competence
- 4.2k GitHub stars indicate genuine early traction for a narrow technical tool
What it fails at
- Surface leans on concept demos rather than published reliability data across real apps or device farms
- No CI-integration or device-farm story surfaced on the reviewed page
- No pricing, support or hosted-service model stated — reads as an open-source CLI you self-operate
- Native device automation is inherently brittle across OS versions; the surface does not address that head-on
- No named production users or case studies as evidence
Best for
- Agent builders who need to drive real native iOS/Android/TV apps, not just web pages
- Mobile QA teams wiring an AI agent into device testing with token cost in mind
- Developers who want a single automation model across simulators and physical devices
Not recommended for
- Web-only agent workflows already served by browser automation
- Teams needing a supported, SLA-backed commercial product out of the box
- Buyers who require published cross-OS reliability evidence before adopting
Pricing & access
- ModelOpen source
- Free entry pointYes — a free tier or open-source edition is documented
- Pricing findable on the public surfaceFAIL No pricing surfaced; reads as an open-source CLI installed via npm (tested 2026-08-25)
Derived from Hlido-held evidence only (engine checklist + editorial text); quotes are verbatim from the scorecard; not vendor-supplied; re-derived daily. Verify current prices on the vendor's pricing page. Last verified 2026-08-25.
Compared to
-
Playwright MCP Server (ExecuteAutomation)
native-mobile-vs-web
The Playwright MCP server automates web browsers; agent-device brings the same snapshot-and-act model to native mobile and TV apps. Playwright MCP for the web surface, agent-device for native app targets a browser tool cannot reach.
-
Browser Use
device-vs-browser-target
browser-use gives agents structured control of a web browser; agent-device is its native-app analogue with an explicit token-efficiency focus. Same philosophy, different surface — pick by whether the target is a website or a real device app.
Agent relevance
CLI Behavioral-testable
Agentic-Commerce Readiness 48/100 · SURFACE-ONLY
Independent readiness for agent delegation & transaction. How it’s scored · check live
A CLI built for agents: an agent shells out to `agent-device open`, `snapshot`, `press`, `fill`, and `find` to explore and drive a native app. The snapshot output is formatted for LLM consumption (interactive-only, stable refs), and semantic selectors let the agent target elements by text or role. No MCP server is documented on the surface, but the CLI itself is the agent interface.
Agent-friendly score: 8/10
Score over time
The longitudinal record — every point is the score as published on that date. Raw series.
Evidence
- Device automation CLI for AI agents across real iOS, Android, TV and desktop apps — source (2026-08-25) verified
- snapshot exposes the accessibility tree with stable refs, keeping context smaller than screenshots — source (2026-08-25) verified
- Acts via stable refs or semantic selectors and installs via a single global npm command — source (2026-08-25) verified
- Approximately 4.2k GitHub stars indicating early traction — source (2026-08-25) verified


