Codex CLI

Coding · tested 2026-08-24 · re-test due 2026-11-22 · by the Hlido desk, not the vendor

In short: A fast, capable vendor CLI with real safety ideas — held back by a failure mode its own doctor would diagnose.

4 PASS · 1 FAIL of 5 public-surface claims

Quick answer

Codex CLI scores 85/100 (STEADY) on Hlido’s independent, hands-on test (reviewed 2026-08-24). STEADY (85) because the public surface is broad and mostly excellent — sandbox subcommand, trusted-directory guard, self-hosting as an MCP server, an actionable doctor — but the one behavior that matters most for unatten

We installed Codex CLI cold from npm (codex-cli 0.149.1) and probed its public surface. Much of it is genuinely strong: a rich command set (non-interactive exec, code review, session resume/fork), MCP management plus the ability to run Codex itself as an MCP server, a built-in sandbox subcommand, and a trusted-directory guard that refuses to run outside a trusted git repository unless explicitly overridden — a safe default we verified. Its doctor is excellent: run in our restricted sandbox it named every real problem (missing credentials, blocked egress) with a concrete fix each. But the flagship check failed: `codex exec` without credentials printed its session banner and then hung, silently, until we killed it at 40 seconds. No auth error, no timeout, no hint — while its own doctor knows exactly what is wrong. A CLI that other agents will drive non-interactively must fail fast; this one does not. What this review does NOT cover: the agentic loop requires paid credentials and was not exercised; medium confidence. Disclosure: Hlido’s review pipeline runs on Claude models from Anthropic, a direct competitor of OpenAI. The mechanical checks are reproducible by anyone; weigh the editorial layer with that knowledge.

Why STEADY

STEADY (85) because the public surface is broad and mostly excellent — sandbox subcommand, trusted-directory guard, self-hosting as an MCP server, an actionable doctor — but the one behavior that matters most for unattended and agent-driven use, failing fast on missing credentials, failed our test outright (40s silent hang, killed by timeout). That is a production-path defect on a CLI explicitly designed for non-interactive execution, and it caps the craft dimension below the VITAL band until fixed.

Public-surface checklist

What it does well

What it fails at

Red flags

Best for

  • Developers in the OpenAI stack who want a scriptable terminal agent with session management
  • Teams that value an explicit sandbox boundary for agent-executed commands
  • Agent builders who want a coding agent addressable AS an MCP server

Not recommended for

  • Unattended pipelines that need deterministic fail-fast behavior on auth/config errors (verified defect)
  • Anyone needing a credential-free trial of the actual editing loop

Compared to

Agent relevance

API CLI MCP SDK Behavioral-testable

Agentic-Commerce Readiness 65/100 · INTEGRABLE

Independent readiness for agent delegation & transaction. How it’s scored · check live

Install via npm (@openai/codex); drive non-interactively with `codex exec`; run as an MCP server with `codex mcp-server`; requires OpenAI credentials. Guard automation with explicit timeouts — verified: missing credentials hang rather than error.

Evidence

scorecard.json · registry · methodology

More: compare agents · best of · developer tools · incident registry

Verdict by Hlido Editor, our automated editorial system · Method: cli-tier-1+editorial-narrative-v2+claude-native · Methodology version 2026.05 · Next review due 2026-11-22

How this page was produced. The scores, claim verdicts and evidence come from automated hands-on testing of the product’s public surface. The written analysis is drafted by an AI system, and pages publish without a person reviewing each one. Hlido publishes this record and answers for it — tell us if anything here is wrong and we will correct it.

Embed this trust badge

Hlido trust score

Live, always-current independent score — free to embed on your site or README. No vendor pays for placement.

Markdown

[![Hlido trust score](https://hlido.eu/badge/openai-codex.svg)](https://hlido.eu/check/?agent=openai-codex)

HTML

<a href="https://hlido.eu/check/?agent=openai-codex"><img src="https://hlido.eu/badge/openai-codex.svg" alt="Hlido trust score"></a>