hotcell

Infrastructure · tested 2026-08-18 · re-test due 2026-11-16 · by the Hlido desk, not the vendor

In short: The most technically honest agent-sandboxing surface we have reviewed this cycle — Apache-2.0, self-hosted on hardware you already own, with a containment model spelled out down to which guarantees are kernel-enforced and which are only advisory.

Quick answer

hotcell scores 80/100 (STEADY) on Hlido’s independent, hands-on test (reviewed 2026-08-18). STEADY (80) because the captured surface is exceptionally complete and specific — a coherent containment model (per-sandbox token swap at an egress gateway, keys never in the cell, hard spend caps, three isolation driver Pricing: Open source (free entry point documented).

hotcell is what serious infrastructure documentation looks like. It solves a real and sharpening problem — running untrusted AI-agent code at scale without leaking credentials, hemorrhaging spend, or letting a compromised agent phone your data out — and it does so on hardware you already own (a Mac Mini, a Linux VM, bare metal), self-hosted under Apache-2.0, with one daemon. The containment model is specific rather than aspirational: every model call leaves a sandbox through a single egress gateway that swaps a per-sandbox scoped token for the real provider key, so the real key never enters the cell and dies with it on teardown; hard USD spend caps and provider allowlists bound blast radius; three isolation drivers (Docker, Firecracker, Apple VZ) sit behind one interface. What lifts hotcell above the category norm is its honesty. The comparison table against E2B, Daytona, NVIDIA OpenShell and Tencent CubeSandbox marks its own weaknesses plainly — default-deny egress is 'kernel-enforced on Linux and no-NIC microVMs; advisory on the microVM-NIC and macOS-Docker paths' — and it invites corrections via GitHub issue. A vendor that tells you where its guarantee is only advisory is a vendor worth more trust, not less. The honest caveats are the rating's ceiling too: it is built by essentially one person plus contained agents, the managed Cloud is an early-access waitlist, and there is no third-party audit or adoption signal on the surface — so the isolation strength is well-described but not independently verified. As a self-hosted, agent-drivable containment layer with genuine engineering depth and unusual candour, it is a standout.

Why STEADY

STEADY (80) because the captured surface is exceptionally complete and specific — a coherent containment model (per-sandbox token swap at an egress gateway, keys never in the cell, hard spend caps, three isolation drivers), an Apache-2.0 self-hosted architecture, multiple SDKs and a REST surface — and it is candid about exactly which guarantees are kernel-enforced versus advisory. Not VITAL because it is single-maintainer with the managed Cloud still a waitlist, and no independent security audit, isolation-escape testing or adoption evidence appears on the surface, so the strong containment claims remain well-argued rather than externally verified.

What it does well

What it fails at

Best for

  • Teams running untrusted or autonomous agent code who want provider keys to never enter the execution environment
  • Developers who need parallel isolated sandboxes on their own hardware without per-second managed-sandbox billing
  • Data-residency-constrained workloads that must keep egress locked to an allowlist and audited
  • Agent builders wanting a self-hosted, SDK- and REST-drivable containment layer with real spend caps

Not recommended for

  • Buyers who require a completed third-party isolation audit before trusting a containment boundary
  • Teams relying on default-deny egress on macOS-Docker, where it is advisory rather than kernel-enforced
  • Organisations that need a vendor-managed, multi-tenant SLA today rather than self-hosting
  • Users unwilling to operate their own daemon and hardware

Pricing & access

Derived from Hlido-held evidence only (engine checklist + editorial text); quotes are verbatim from the scorecard; not vendor-supplied; re-derived daily. Verify current prices on the vendor's pricing page. Last verified 2026-08-18.

Compared to

Agent relevance

API CLI SDK Behavioral-testable

Agentic-Commerce Readiness 57/100 · INTEGRABLE

Independent readiness for agent delegation & transaction. How it’s scored · check live

hotcell is containment infrastructure agents run inside and can drive: every CLI command is also a REST call, and TypeScript and Python SDKs ship with it, so an orchestrator can create sandboxes, wire keyless egress, exec streamed commands, and tear everything down programmatically. It runs Claude Code, Codex, OpenCode and Mastra inside its cells. No MCP server is advertised, but the REST + SDK surface is fully agent-drivable.

Agent-friendly score: 9/10

Evidence

scorecard.json · registry · methodology

More: compare agents · best of · developer tools · incident registry

Verdict by Hlido Editor, our automated editorial system · Method: public-surface-tier-1+editorial-narrative-v2 · Methodology version 2026.05 · Next review due 2026-11-16

How this page was produced. The scores, claim verdicts and evidence come from automated hands-on testing of the product’s public surface. The written analysis is drafted by an AI system, and pages publish without a person reviewing each one. Hlido publishes this record and answers for it — tell us if anything here is wrong and we will correct it.

Embed this trust badge

Hlido trust score

Live, always-current independent score — free to embed on your site or README. No vendor pays for placement.

Markdown

[![Hlido trust score](https://hlido.eu/badge/sinameraji-hotcell.svg)](https://hlido.eu/check/?agent=sinameraji-hotcell)

HTML

<a href="https://hlido.eu/check/?agent=sinameraji-hotcell"><img src="https://hlido.eu/badge/sinameraji-hotcell.svg" alt="Hlido trust score"></a>