Harden
Infrastructure · tested 2026-09-14 · re-test due 2026-12-13 · by the Hlido desk, not the vendor
In short: A pre-execution guardrail for coding agents with an unusually evidence-forward public surface.
4 PASS · 0 FAIL of 4 public-surface claims
Quick answer
Harden scores 85/100 (STEADY) on Hlido’s independent, hands-on test (reviewed 2026-09-14). STEADY (85) because the product addresses a real, growing problem class our own demand data confirms (agent tool-call safety), the public surface demonstrates the mechanism rather than merely claiming it, and the benchma
Harden ships a local monitor (the "Agentic Integrity Foundation") that checks a coding agent’s tool calls before they execute — the captured page demonstrates it live: a kubectl rollout allowed, a production-namespace delete blocked with a safe retry suggested. The surface is unusually honest for this market: a one-line no-account curl install, per-session decision histories with real-looking counts, named third-party-style benchmarks (SLEIGHT, AgentHazard, SABER, LinuxArena) with a GPT baseline column and an explicit "lower is better" annotation where the direction flips. It names support for the agents our own register measures demand for — Claude Code, Codex, Cursor and others — via native hooks with an MCP-proxy fallback. What this review does NOT cover: the benchmark numbers are the vendor’s own and were not re-run; the monitor’s live blocking behaviour was not exercised beyond the public demo surface. Medium confidence.
Why STEADY
STEADY (85) because the product addresses a real, growing problem class our own demand data confirms (agent tool-call safety), the public surface demonstrates the mechanism rather than merely claiming it, and the benchmark presentation includes the direction-of-goodness honesty most vendors omit — but every quantitative claim remains self-reported and the blocking loop was not independently exercised, which caps it below the VITAL band.
Public-surface checklist
- PASS Homepage loads (required) — 4 screenshots, 3/3 interactions succeeded, no blockers
- PASS Primary value prop (required) — "judge every action a coding agent is about to take, and stop the dangerous ones before they execute"
- PASS Mechanism demonstrated — live allowed/blocked/safe-retry demo with per-session counts
- PASS Claims direction honesty — benchmark table annotates "lower is better" where applicable
What it does well
- Demonstrates the core mechanism on the page: allowed vs blocked calls with a safe-retry suggestion (captured)
- No-account, one-line local install — the lowest-friction trial in the category (captured)
- Benchmarks shown with baselines and an explicit lower-is-better annotation (captured)
- Supports the coding agents agents actually use — Claude Code, Codex, Cursor — with an MCP-proxy fallback (captured)
What it fails at
- All benchmark figures are self-reported; no third-party verification is linked (captured surface only)
- The blocking loop itself was not exercised in this tier-2 review — the evidence is the vendor’s own demo data
Best for
- Teams running autonomous coding agents who want a local, pre-execution safety layer
- Security-conscious orgs that need agent tool-call decisions logged on-device
Not recommended for
- Anyone requiring independently verified efficacy numbers before deployment
Compared to
-
Claude Code
Claude Code carries its own permission scoping (--allowedTools, sandbox-aware flags); Harden positions as an independent, vendor-external judgment layer across MANY agents — complementary, not competing.
Agent relevance
CLI MCP Behavioral-testable
Install via the one-line local script; native hooks attach to supported coding agents, MCP proxy covers others; decisions logged locally.
Score over time
The longitudinal record — every point is the score as published on that date. Raw series.
Evidence
- — source