OmYarewar/PHANTOM
Specialized verticals · tested 2026-09-30 · by the Hlido desk, not the vendor
In short: An AI pentesting command centre with a clean repo and agent-consumable design — promising for authorised security testing, unverified and unversioned.
9 PASS · 1 FAIL of 10 public-surface claims
Quick answer
OmYarewar/PHANTOM scores 82/100 (STEADY) on Hlido’s independent, hands-on test (reviewed 2026-09-30). STEADY (82) on solid repo hygiene (licence, maintenance, CI, documented install) plus the agent-consumable pass. Pricing: Open source (free entry point documented).
PHANTOM presents as an AI-powered pentesting command centre — autonomous security testing with real-time streaming, self-improving iterations and a polished dark UI. On repo signals it is a tidy, real project: MIT-licensed, actively pushed, CI present, install documented, and it passes agent-consumable. Traction is modest (59 stars), which for a security tool is unremarkable and not a strike against it. The category is what a reader should hold onto. Autonomous, self-improving security testing is genuinely useful for defenders and authorised red teams, but it is dual-use by nature, and the surface track cannot tell us anything about the guardrails: whether it scopes targets, requires authorisation, or constrains 'unlimited tool iterations' against systems the operator does not own. We have not hands-on tested it, so the autonomy claims and — more importantly — the safety posture are entirely unverified. There are also no tagged releases, so anything wired into a workflow tracks a moving branch. Treat PHANTOM as a credible-looking project for authorised, scoped security work by people who understand the legal and safety boundaries; do not read the clean checklist as any statement about whether its autonomous behaviour is safe.
Why STEADY
STEADY (82) on solid repo hygiene (licence, maintenance, CI, documented install) plus the agent-consumable pass. Below VITAL by the repo-surface ceiling and the missing releases; the dual-use safety posture is unverified and flagged as the key caveat rather than scored, since the surface cannot measure it.
Public-surface checklist
- PASS Repo reachable (required) — GH API 200 for OmYarewar/PHANTOM
- PASS Readme present (required) — README length 10117
- PASS License present (required) — MIT
- PASS Install documented (required) — install/usage section found in README
- PASS Active 12mo (required) — last push 0d ago
- FAIL Releases present — no releases
- PASS Community traction — 59 stars
- PASS Ci or tests — 2 workflow file(s)
- PASS Recent commit 90d — last push 0d ago
- PASS Agent consumable — MCP server markers in README
What we saw
1 screenshot captured by the Hlido engine during the reviewed run (run-c1b452d18d9dc5da-github-com). Our own captures — not vendor marketing material.
What it does well
- MIT-licensed, actively maintained, CI present, install documented
- Passes agent-consumable — designed to be driven programmatically
- Real-time streaming and an iterative testing loop suit authorised red-team work
- Coherent, purpose-built command-centre design
What it fails at
- Dual-use with no visible guardrail evidence — scoping/authorisation controls unverified on the surface
- Not hands-on tested: autonomy and, critically, safety behaviour are unverified
- No tagged releases — workflow consumers track a moving branch
- Autonomous "unlimited tool iterations" framing warrants caution against un-owned targets
Best for
- Authorised red teams and defenders doing scoped, legal security testing
- Security engineers who can vet the tool before pointing it at anything
- Labs and CTF/research contexts with clear authorisation
Not recommended for
- Anyone without explicit authorisation for the targets tested
- Buyers who would mistake a clean repo checklist for a safety guarantee
- Production automation needing versioned, release-pinned dependencies
Pricing & access
- ModelOpen source
- Free entry pointYes — a free tier or open-source edition is documented
Derived from Hlido-held evidence only (engine checklist + editorial text); quotes are verbatim from the scorecard; not vendor-supplied; re-derived daily. Verify current prices on the vendor's pricing page. Last verified 2026-07-16.
Related agents
Agent relevance
No programmatic surfaces
Agentic-Commerce Readiness 40/100 · SURFACE-ONLY
Independent readiness for agent delegation & transaction. How it’s scored · check live
Documented as agent-consumable and driven programmatically as a security-testing command centre. Use only against authorised, scoped targets; the surface track verifies none of its safety controls.
Agent-friendly score: 5/10
Score over time
The longitudinal record — every point is the score as published on that date. Raw series.
