Serena
Coding · tested 2026-08-19 · re-test due 2026-11-19 · by the Hlido desk, not the vendor
In short: An agent-first, IDE-grade semantic toolkit that plugs into any MCP client — one of the cleaner examples of tooling designed for the model, not the human.
Quick answer
Serena scores 84/100 (STEADY) on Hlido’s independent, hands-on test (reviewed 2026-08-19). STEADY (84) for a mature, MCP-native coding toolkit with a genuinely agent-first design, broad documented client support, and an unusually honest evaluation posture, held at low-medium confidence because the review is su
Serena provides semantic code retrieval, editing, refactoring and debugging tools that operate at the symbol level rather than on line numbers or raw text, and exposes them to any LLM client over MCP — Claude Code, Codex, Gemini-CLI, Cursor, VS Code, JetBrains and the desktop clients are all listed as supported. The pitch is coherent and, unusually, self-aware: the project states outright that its "end users" are the AI agents that actually call the tools, and it publishes an evaluation where agents run ~20 routine coding tasks with and without Serena and report back. Those testimonials (Opus in Claude Code, GPT in Codex CLI, Copilot CLI on a monorepo) are vendor-run and self-selected, so Hlido treats them as illustrative rather than independent proof. What is verifiable from the surface is the design posture — high-level, symbol-aware abstractions that collapse multi-step text surgery into atomic operations, and honest framing that an LLM still does the actual orchestration. Hlido assessed the documentation site only, not a running session, so the efficiency and reliability claims are unconfirmed hands-on. But the architecture is exactly what the agent-to-agent thesis rewards: structured, machine-first, and standards-based via MCP.
Why STEADY
STEADY (84) for a mature, MCP-native coding toolkit with a genuinely agent-first design, broad documented client support, and an unusually honest evaluation posture, held at low-medium confidence because the review is surface-only (docs site, not a running session) and the headline efficiency/reliability gains rest on vendor-run agent testimonials Hlido did not reproduce. Not VITAL because none of the performance claims were independently verified.
What we saw
4 screenshots captured by the Hlido engine during the reviewed run (run-69fc58ba7b3ed1fd-oraios-github-io). Our own captures — not vendor marketing material.
What it does well
- Genuinely agent-first tool design — symbol-level abstractions (renames, references, refactors) built for the model to call, not line-number text surgery
- Standards-based integration via MCP; documented support across terminal clients (Claude Code, Codex, Gemini-CLI), IDEs (Cursor, VS Code, JetBrains) and desktop/web clients
- Unusually honest framing — it names AI agents as the real end users and publishes an evaluation methodology instead of bare marketing
- Clear explanation that an LLM is still required to orchestrate; no overreach about autonomous capability
What it fails at
- The headline efficiency and reliability gains rest on vendor-run, self-selected agent testimonials — not independent benchmarks
- Surface-only review — Hlido read the documentation site, not a live session, so tool behaviour was not hands-on verified
- No hosted option; it is a self-run server you wire into your own client
- Value is bounded by the host LLM and MCP client — Serena adds tools but does no work on its own
Best for
- Developers running an MCP-capable coding agent (Claude Code, Codex, Cursor) who want IDE-grade semantic navigation and refactoring
- Teams working in large or complex codebases where cross-file refactors and reference lookups are error-prone for text-based agents
- Anyone wanting a client-agnostic, open toolkit rather than a single-vendor coding assistant
Not recommended for
- Users wanting a turnkey hosted coding assistant with no setup
- Non-technical users — this is developer tooling wired into a coding agent
- Anyone needing independently benchmarked performance guarantees before adopting
Related agents
Agent relevance
CLI MCP Behavioral-testable
Agentic-Commerce Readiness 57/100 · INTEGRABLE
Independent readiness for agent delegation & transaction. How it’s scored · check live
Serena runs as an MCP server that any MCP-capable client launches (via a launch command) or connects to over HTTP. Coding agents then call its symbol-level retrieval/edit/refactor tools directly in natural-language workflows.
Agent-friendly score: 9/10
Score over time
The longitudinal record — every point is the score as published on that date. Raw series.
Evidence
- Genuinely agent-first tool design — symbol-level abstractions (renames, references, refactors) built for the model to call, not line-number text surgery — source (2026-08-19) verified
- Standards-based integration via MCP; documented support across terminal clients (Claude Code, Codex, Gemini-CLI), IDEs (Cursor, VS Code, JetBrains) and desktop/web clients — source (2026-08-19) verified
- Unusually honest framing — it names AI agents as the real end users and publishes an evaluation methodology instead of bare marketing — source (2026-08-19) verified
- Clear explanation that an LLM is still required to orchestrate; no overreach about autonomous capability — source (2026-08-19) verified
- Hands-on runtime behaviour (executing the tool / a live task) — source (2026-08-19)



