mcp-v8
Infrastructure · tested 2026-08-19 · re-test due 2026-11-19 · by the Hlido desk, not the vendor
In short: A code-execution MCP server that hands an agent one run_js tool inside a locked-down V8 isolate — an ambitious 'agents write code, not tool-calls' bet with a serious security story.
Quick answer
mcp-v8 scores 80/100 (STEADY) on Hlido’s independent, hands-on test (reviewed 2026-08-19). STEADY (80) for a well-architected code-execution MCP server with sandbox-by-default capabilities (OPA/Rego-gated network/FS/subprocess), durable heap-snapshot state, JWKS auth and production transports, marked low-mediu
mcp-v8 is a Model Context Protocol server that executes JavaScript and TypeScript inside a V8 isolate. Instead of exposing dozens of narrow tools, it gives an agent a single run_js tool: the agent writes code that can loop, branch, transform data and chain other MCP servers, which the vendor argues costs fewer tokens than equivalent tool-call sequences. In its default stateful mode it persists the V8 heap as a content-addressed snapshot so state survives across calls — a genuinely useful primitive for multi-turn agent work. The security framing is the strongest part of the surface: network fetch, filesystem, subprocess, WASM and ES-module imports are all off by default and each is gated by OPA/Rego policy, requests can be authenticated with JWT/JWKS, and the server can form a Raft cluster to replicate session metadata. It speaks stdio, Streamable HTTP and SSE with a REST sidecar. This is documentation, not a running audit — Hlido did not execute code against a live instance, so the isolation and policy enforcement are described capabilities rather than tested ones. But the design is coherent, the sandbox-first defaults are the right ones for handing an LLM an execution surface, and the docs are structured (install / tutorials / how-to / concepts / reference) rather than hype.
Why STEADY
STEADY (80) for a well-architected code-execution MCP server with sandbox-by-default capabilities (OPA/Rego-gated network/FS/subprocess), durable heap-snapshot state, JWKS auth and production transports, marked low-medium confidence because the review is surface-only — the isolation, policy gating and clustering are documented, not exercised against a live instance by Hlido. Not VITAL absent hands-on verification of the sandbox that is the whole value proposition.
What we saw
4 screenshots captured by the Hlido engine during the reviewed run (run-a83a7a4d13ed9846-r33drichards-github-io). Our own captures — not vendor marketing material.
What it does well
- Compelling 'one tool, write code' model — a single run_js call replaces long tool-call chains and can compose other MCP servers, often at lower token cost
- Durable state via content-addressed V8 heap snapshots, so an agent builds context across turns without re-sending it
- Secure-by-default: network, filesystem, subprocess and module imports are all off until an explicit OPA/Rego policy grants them
- Production-grade operations on paper — stdio/HTTP/SSE transports, REST sidecar, JWKS auth, and Raft-replicated clustering
- Structured documentation (install, tutorials, how-to, concepts, reference) rather than a single marketing page
What it fails at
- Handing an agent arbitrary code execution is inherently high-blast-radius — the entire safety case rests on policy configuration the operator must get right
- Surface-only review — Hlido did not run code against a live isolate, so sandbox isolation and policy enforcement are unverified
- Self-run infrastructure with real operational surface (policies, auth, optionally Raft) — not a turnkey product
- No independent security audit is cited on the captured surface
Red flags
- Arbitrary code execution handed to an LLM is only as safe as the OPA/Rego policies configured around it; misconfiguration widens the blast radius substantially
Best for
- Agent builders who want a code-interpreter tool that composes other MCP servers and persists state across calls
- Teams comfortable authoring OPA/Rego policies to scope exactly what agent-run code may touch
- Advanced MCP deployments needing auth (JWKS) and multi-node replication
Not recommended for
- Anyone wanting a zero-config tool — the security depends on policies you write
- Environments that cannot accept agent-driven code execution under any sandbox
- Users needing a vendor-audited, certified isolation guarantee
Related agents
Agent relevance
API MCP Behavioral-testable
Agentic-Commerce Readiness 59/100 · INTEGRABLE
Independent readiness for agent delegation & transaction. How it’s scored · check live
Runs as an MCP server (stdio, Streamable HTTP or SSE, plus a REST sidecar). An MCP client calls a single run_js tool; capabilities beyond compute are granted per OPA/Rego policy and requests can be JWKS-authenticated.
Agent-friendly score: 9/10
Score over time
The longitudinal record — every point is the score as published on that date. Raw series.
Evidence
- Compelling 'one tool, write code' model — a single run_js call replaces long tool-call chains and can compose other MCP servers, often at lower token cost — source (2026-08-19) verified
- Durable state via content-addressed V8 heap snapshots, so an agent builds context across turns without re-sending it — source (2026-08-19) verified
- Secure-by-default: network, filesystem, subprocess and module imports are all off until an explicit OPA/Rego policy grants them — source (2026-08-19) verified
- Production-grade operations on paper — stdio/HTTP/SSE transports, REST sidecar, JWKS auth, and Raft-replicated clustering — source (2026-08-19) verified
- Hands-on runtime behaviour (executing the tool / a live task) — source (2026-08-19)


