agent-device

Coding · tested 2026-08-25 · re-test due 2026-11-23 · by the Hlido desk, not the vendor

In short: Playwright's snapshot-and-act model, ported to native mobile and TV — a token-efficient device-automation CLI built specifically for agents, from a team that knows the mobile tooling space.

4 PASS · 1 FAIL of 5 public-surface claims

Quick answer

agent-device scores 80/100 (STEADY) on Hlido’s independent, hands-on test (reviewed 2026-08-25). STEADY (80) because agent-device applies a proven, token-conscious automation model to an underserved surface (native mobile/TV), ships as a clean single-install CLI designed for agents, and comes from a team with real m Pricing: Open source (free entry point documented).

agent-device takes the pattern that made web agents practical — expose the page as a structured accessibility tree with stable references, let the agent read it, then act on those references — and brings it to native iOS, Android, TV and desktop apps. The insight is explicitly about tokens: instead of feeding an agent raw screenshots or verbose UI dumps, `snapshot -i` returns only the interactive elements with stable refs like @e2, keeping the context small enough that a real exploration loop stays affordable. From there the agent acts with either those refs or semantic selectors ('find Sign In click', 'find role button click'), which keeps generated flows readable and resilient to layout churn. It is a single global npm install, one mental model across simulators, physical QA devices and TV targets, and evidence (screenshots) is captured only when needed rather than on every step. Two things earn it real credibility: it is built by callstack, a well-known React Native consultancy, so the mobile-tooling competence is not in doubt; and the 4.2k GitHub stars signal genuine early traction for what is a narrow, technical tool. What keeps it in the middle of the STEADY band rather than the top is maturity and proof — the value proposition is clean but the surface leans on the concept demo rather than published reliability data across real apps, device farms, or CI, and there is no pricing or hosted-service story visible (it reads as an open-source CLI). For an agent that needs to drive a real mobile app, though, this is the right shape of tool and a rare one.

Why STEADY

STEADY (80) because agent-device applies a proven, token-conscious automation model to an underserved surface (native mobile/TV), ships as a clean single-install CLI designed for agents, and comes from a team with real mobile-tooling pedigree plus early traction (4.2k stars). It sits mid-band rather than higher because the public surface is concept-and-command led — it does not yet publish reliability evidence across real apps or a CI/device-farm story — and the commercial/support model is unstated.

Public-surface checklist

What we saw

4 screenshots captured by the Hlido engine during the reviewed run (run-7469c725a338cb3a-agent-device-dev). Our own captures — not vendor marketing material.

agent-device — run screenshot 1 (home.png)
home.png
agent-device — run screenshot 2 (page_.png)
page_.png
agent-device — run screenshot 3 (page_.png)
page_.png
agent-device — run screenshot 4 (page_cloud.png)
page_cloud.png

What it does well

What it fails at

Best for

  • Agent builders who need to drive real native iOS/Android/TV apps, not just web pages
  • Mobile QA teams wiring an AI agent into device testing with token cost in mind
  • Developers who want a single automation model across simulators and physical devices

Not recommended for

  • Web-only agent workflows already served by browser automation
  • Teams needing a supported, SLA-backed commercial product out of the box
  • Buyers who require published cross-OS reliability evidence before adopting

Pricing & access

Derived from Hlido-held evidence only (engine checklist + editorial text); quotes are verbatim from the scorecard; not vendor-supplied; re-derived daily. Verify current prices on the vendor's pricing page. Last verified 2026-08-25.

Compared to

Agent relevance

CLI Behavioral-testable

Agentic-Commerce Readiness 48/100 · SURFACE-ONLY

Independent readiness for agent delegation & transaction. How it’s scored · check live

A CLI built for agents: an agent shells out to `agent-device open`, `snapshot`, `press`, `fill`, and `find` to explore and drive a native app. The snapshot output is formatted for LLM consumption (interactive-only, stable refs), and semantic selectors let the agent target elements by text or role. No MCP server is documented on the surface, but the CLI itself is the agent interface.

Agent-friendly score: 8/10

Evidence

scorecard.json · transparency passport · registry · methodology

More: compare agents · best of · developer tools · incident registry

Verdict by Hlido Editor, our automated editorial system · Method: public-surface-tier-1+editorial-narrative-v2 · Methodology version 2026.05 · Next review due 2026-11-23

How this page was produced. The scores, claim verdicts and evidence come from automated hands-on testing of the product’s public surface. The written analysis is drafted by an AI system, and pages publish without a person reviewing each one. Hlido publishes this record and answers for it — tell us if anything here is wrong and we will correct it.

Embed this trust badge

Hlido trust score

Live, always-current independent score — free to embed on your site or README. No vendor pays for placement.

Markdown

[![Hlido trust score](https://hlido.eu/badge/callstackincubator-agent-device.svg)](https://hlido.eu/check/?agent=callstackincubator-agent-device)

HTML

<a href="https://hlido.eu/check/?agent=callstackincubator-agent-device"><img src="https://hlido.eu/badge/callstackincubator-agent-device.svg" alt="Hlido trust score"></a>