SkillClaw

Frameworks & Eval · tested 2026-06-16 · re-test due 2026-09-16 · by the Hlido desk, not the vendor

In short: Research-grade collective skill evolution for AI agents — 1,900 stars and an arXiv paper make this more than a weekend project, but it's a research tool, not a production one.

5 PASS · 0 FAIL of 5 public-surface claims

Quick answer

SkillClaw scores 70/100 (STEADY) on Hlido’s independent, hands-on test (reviewed 2026-06-16). STEADY (70) because the arXiv publication gives it more credibility than typical OSS projects, 1,901 stars signal real ML community interest, and the collective learning across agents/devices is a genuinely differentiate Pricing: Open source (free entry point documented).

SkillClaw addresses a real problem that most agent frameworks ignore: agents learn nothing from their interactions. Every session starts fresh, making the same mistakes and rediscovering the same solutions. SkillClaw's approach — collective skill evolution where agent skills improve from every interaction and share that learning across agents, sessions, and devices — is genuinely novel and backed by a published arXiv paper (2604.08377). The 1,901 GitHub stars for a research-lab project suggest the ML community takes the approach seriously. Compatibility with Hermes, OpenClaw, QwenPaw, IronClaw, PicoClaw, and ZeroClaw shows ecosystem investment beyond a single-paper demo. What's cautious here: the gap between research results (which are often measured on controlled benchmarks) and production reliability (which depends on real user interactions, edge cases, and adversarial inputs) is large. SkillClaw's value proposition is specifically that skills evolve from 'real interactions' — which means early users are essentially contributing training signal, with all the quality variance that implies. For teams that want continual skill improvement and are willing to operate in a research-grade framework, SkillClaw is the most thoughtful solution in this space.

Why STEADY

STEADY (70) because the arXiv publication gives it more credibility than typical OSS projects, 1,901 stars signal real ML community interest, and the collective learning across agents/devices is a genuinely differentiated capability. Not VITAL because it's a research tool with production reliability questions, and the skill evolution effectiveness in uncontrolled environments is unverified from the public surface.

Public-surface checklist

What we saw

1 screenshot captured by the Hlido engine during the reviewed run (run-3ffa6d56e0067a6e-github-com). Our own captures — not vendor marketing material.

SkillClaw — run screenshot 1 (home.png)
home.png

What it does well

What it fails at

Best for

  • ML researchers building on agent skill learning foundations
  • Developers already using Hermes or OpenClaw agents who want continual improvement
  • Projects where the same task types recur at scale and improvement from iteration is valuable
  • Research teams that want a published-method foundation rather than proprietary black-box learning

Not recommended for

  • Production systems where skill quality variance is unacceptable
  • One-off or low-volume agent tasks (collective learning needs volume to show value)
  • Teams wanting a standalone agent framework — SkillClaw is a plugin layer, not a full framework
  • Security-sensitive deployments without vetted skill provenance controls

Pricing & access

Derived from Hlido-held evidence only (engine checklist + editorial text); quotes are verbatim from the scorecard; not vendor-supplied; re-derived daily. Verify current prices on the vendor's pricing page. Last verified 2026-06-16.

Compared to

Agent relevance

CLI SDK

Agentic-Commerce Readiness 49/100 · SURFACE-ONLY

Independent readiness for agent delegation & transaction. How it’s scored · check live

Skill plugin for Hermes/OpenClaw and compatible agents. Install via npx skills add. Agents call SkillClaw's skill-store endpoints to retrieve learned skills. Collective evolution happens server-side. No standalone API for external agent consumption.

Agent-friendly score: 6/10

Evidence

scorecard.json · transparency passport · registry · methodology

More: compare agents · best of · developer tools · incident registry

Verdict by Hlido Editor, our automated editorial system · Method: public-surface-tier-2+editorial-narrative-v2 · Methodology version 2026.06 · Next review due 2026-09-16

How this page was produced. The scores, claim verdicts and evidence come from automated hands-on testing of the product’s public surface. The written analysis is drafted by an AI system, and pages publish without a person reviewing each one. Hlido publishes this record and answers for it — tell us if anything here is wrong and we will correct it.

Embed this trust badge

Hlido trust score

Live, always-current independent score — free to embed on your site or README. No vendor pays for placement.

Markdown

[![Hlido trust score](https://hlido.eu/badge/amap-ml-skillclaw.svg)](https://hlido.eu/check/?agent=amap-ml-skillclaw)

HTML

<a href="https://hlido.eu/check/?agent=amap-ml-skillclaw"><img src="https://hlido.eu/badge/amap-ml-skillclaw.svg" alt="Hlido trust score"></a>