Statewright
Workflow & Automation · tested 2026-08-25 · re-test due 2026-11-23 · by the Hlido desk, not the vendor
In short: Protocol-level guardrails for coding agents — turns a large task into bounded workflow phases with per-phase tool policy and model routing, so a destructive tool literally does not exist in a read-only state.
4 PASS · 1 FAIL of 5 public-surface claims
Quick answer
Statewright scores 73/100 (STEADY) on Hlido’s independent, hands-on test (reviewed 2026-08-25). STEADY (73), lower-middle, because Statewright targets a real agent failure mode (looping, scope creep, unsafe tool use) with a materially stronger answer than prompting — protocol-level per-phase tool enforcement — plus Pricing was not findable on the public surface when tested.
Statewright's thesis is that structure beats reasoning: rather than trusting a prompt to keep an agent in line, it decomposes a large agent task into bounded workflow phases, each carrying its own model, reasoning level, tool policy and budget, and enforces the boundaries at the protocol level. The sharpest expression of this is tool enforcement per phase — 'destructive tools don't exist in read-only states; the agent can't call what it can't see' — which is a materially stronger safety posture than the usual prompt-level 'please don't delete anything'. Around that sit decision checkpoints that force progress or fail (no idle looping), read-deduplication, edit guards against scope explosion, and — in the 0.3.0 plugin — native autonomous model routing that switches Claude or Codex between tiers at phase boundaries while the developer stays in the TUI they already use. The workflow itself is built visually (drag states, draw transitions, assign tools per phase), and it integrates with Codex, Claude Code, opencode and Cursor. This is a genuinely good idea addressing a real failure mode — agents that loop, over-read, or reach for tools they should not — and the enforcement-not-suggestion framing is the right one. It sits in the lower-middle of STEADY because it is early (a 0.3.0 plugin), the reliability of the enforcement across real long-running tasks is asserted rather than evidenced on the surface, and pricing sits behind a sign-up ('Start Free' with no published tiers). The concept is strong; the proof and the commercial detail are still thin.
Why STEADY
STEADY (73), lower-middle, because Statewright targets a real agent failure mode (looping, scope creep, unsafe tool use) with a materially stronger answer than prompting — protocol-level per-phase tool enforcement — plus model routing and a visual builder across major agent clients. It is held down because it is early (0.3.0), the enforcement's reliability on real long tasks is asserted rather than demonstrated, and pricing is behind sign-up with no published tiers.
Public-surface checklist
- PASS Homepage loads (required)
- PASS Primary value prop (required) — Autonomous runs. Deliberate boundaries. — bounded, enforced workflow phases for AI coding agents
- PASS Cta present (required) — Start Free / View on GitHub
- FAIL Pricing or access — 'Start Free' with no published pricing tiers on the surface
- PASS Evidence or demo — Three-step workflow walkthrough, visual editor, per-client init commands (Codex/Claude Code/opencode/Cursor)
What we saw
2 screenshots captured by the Hlido engine during the reviewed run (run-3f6c85effd018606-statewright-ai). Our own captures — not vendor marketing material.
What it does well
- Protocol-level tool enforcement: destructive tools are absent from read-only phases, not merely discouraged by a prompt
- Bounds a large task into phases each with its own model, reasoning level, tool policy and budget
- Decision checkpoints force progress or fail, curbing idle looping
- Read-deduplication and edit guards curb repeated reads and scope explosion
- Native autonomous model routing (0.3.0) switches Claude/Codex tiers at phase boundaries inside the existing TUI
- Visual workflow editor to design states, transitions and per-phase tool policy
- Integrates with Codex, Claude Code, opencode and Cursor
What it fails at
- Early-stage (0.3.0 plugin); the model and its guarantees are still moving
- Enforcement reliability across real long-running tasks is asserted, not evidenced on the surface
- Pricing is behind 'Start Free' sign-up with no published tiers
- Designing good workflows is itself work — the safety depends on the human modelling phases well
- No named production adopters or case studies as evidence
Best for
- Teams running autonomous coding agents who need hard, protocol-level limits on destructive actions
- Developers who want per-phase model routing to spend frontier reasoning only where it earns its keep
- Anyone burned by agents that loop, over-read, or exceed their intended scope
Not recommended for
- Users wanting a zero-configuration agent — Statewright requires modelling the workflow
- Buyers who need published pricing and proven long-task reliability before adopting
- Simple single-step tasks where phase decomposition is overhead
Pricing & access
- Pricing findable on the public surfaceFAIL 'Start Free' with no published pricing tiers on the surface (tested 2026-08-25)
Derived from Hlido-held evidence only (engine checklist + editorial text); quotes are verbatim from the scorecard; not vendor-supplied; re-derived daily. Verify current prices on the vendor's pricing page. Last verified 2026-08-25.
Compared to
-
goose
enforced-workflow-control
Goose is an open, extensible agent that executes tasks; Statewright is a control layer that constrains an agent (including ones like these) into enforced phases. Goose to do the work, Statewright to bound how an agent is allowed to do it.
-
@langchain/langgraph-supervisor
guardrails-vs-orchestration
LangGraph Supervisor orchestrates multi-agent flows in code; Statewright enforces phase boundaries and tool policy for a single coding agent with a visual builder and a TUI. LangGraph for programmatic multi-agent orchestration, Statewright for protocol-level guardrails on an interactive coding agent.
Agent relevance
CLI Behavioral-testable
Agentic-Commerce Readiness 48/100 · SURFACE-ONLY
Independent readiness for agent delegation & transaction. How it’s scored · check live
Statewright is a control layer for coding agents rather than an agent itself. It installs as a plugin/CLI (e.g. `npx statewright-codex@latest init`) into Codex, Claude Code, opencode or Cursor, then enforces the phase workflow — tool policy, model tier, budget — as the host agent runs. The integration path is agent-facing by design; it constrains an agent rather than exposing tools to one.
Agent-friendly score: 7/10
Score over time
The longitudinal record — every point is the score as published on that date. Raw series.
Evidence
- Turns a large agent task into bounded workflow phases with per-phase model, reasoning level, tool policy and budget — source (2026-08-25) verified
- Tool enforcement is protocol-level: destructive tools do not exist in read-only states — source (2026-08-25) verified
- 0.3.0 adds native autonomous model routing for Claude and Codex at workflow boundaries — source (2026-08-25) verified
- Pricing is behind sign-up with no published tiers on the surface — source (2026-08-25)

