wanshuiyin/Auto-claude-code-research-in-sleep
Research · tested 2026-09-30 · by the Hlido desk, not the vendor
In short: Very high-traction, framework-free ML-research skill pack (ARIS) for autonomous overnight research loops — clean checklist, portable across agents.
10 PASS · 0 FAIL of 10 public-surface claims
Visit wanshuiyin/Auto-claude-code-research-in-sleep → scorecard.json
Quick answer
wanshuiyin/Auto-claude-code-research-in-sleep scores 82/100 (STEADY) on Hlido’s independent, hands-on test (reviewed 2026-09-30). STEADY (82) — a perfect checklist on the repo-surface track (every item passes) plus exceptional community traction. Pricing: Open source (free entry point documented).
ARIS (Auto-Research-In-Sleep) is a set of lightweight, Markdown-only skills for autonomous ML research — cross-model review loops, idea discovery and experiment automation — deliberately built with no framework and no lock-in so it runs under Claude Code, Codex, OpenClaw or any capable LLM agent. On measured signals it is the standout of this batch: 13,480 stars and 1,216 forks is exceptional traction, and it passes every checklist item we test on the surface track, including agent-consumable. The Markdown-only, framework-free design is a genuine strength for the agent-to-agent world we care about — there is nothing to install and nothing to break, so portability is close to free. What the surface track cannot tell us is whether the 'research in sleep' loops actually produce sound results or simply run unattended; overnight autonomy amplifies both good and bad reasoning, and we have not run it. Treat the star count as evidence of strong developer interest in the pattern, not as proof the outputs are trustworthy. As a portable skill layer to study and adapt, it is a strong pick; as an unattended research pipeline, supervise the first several runs.
Why STEADY
STEADY (82) — a perfect checklist on the repo-surface track (every item passes) plus exceptional community traction. Capped below VITAL only by the repo-surface ceiling: we have not hands-on verified that the autonomous loops produce sound research, which overnight autonomy makes the central question.
Public-surface checklist
- PASS Repo reachable (required) — GH API 200 for wanshuiyin/Auto-claude-code-research-in-sleep
- PASS Readme present (required) — README length 192363
- PASS License present (required) — MIT
- PASS Install documented (required) — install/usage section found in README
- PASS Active 12mo (required) — last push 2d ago
- PASS Releases present — latest v0.4.22
- PASS Community traction — 13480 stars
- PASS Ci or tests — 4 workflow file(s)
- PASS Recent commit 90d — last push 2d ago
- PASS Agent consumable — MCP server markers in README
What we saw
1 screenshot captured by the Hlido engine during the reviewed run (run-8b3ddac650e1987c-github-com). Our own captures — not vendor marketing material.
What it does well
- Passes every checklist item we test on the surface track, including agent-consumable
- Exceptional traction — 13,480 stars, 1,216 forks
- Framework-free, Markdown-only: portable across Claude Code, Codex and other agents with nothing to install
- MIT-licensed with releases — clear terms and versioned artefacts
What it fails at
- Not hands-on tested — whether the autonomous loops produce sound research is unverified
- Unattended overnight autonomy amplifies bad reasoning as readily as good; no guardrail evidence on the surface
- Star count signals interest in the pattern, not validated output quality
Best for
- Researchers wanting a portable, lock-in-free skill layer to adapt
- Multi-agent setups (Claude Code / Codex) that value drop-in Markdown skills
- Teams comfortable supervising early autonomous runs before trusting them
Not recommended for
- Anyone expecting validated, hands-off research output out of the box
- Environments that cannot supervise unattended autonomous loops
- Buyers who read a high star count as a quality guarantee
Pricing & access
- ModelOpen source
- Free entry pointYes — a free tier or open-source edition is documented
Derived from Hlido-held evidence only (engine checklist + editorial text); quotes are verbatim from the scorecard; not vendor-supplied; re-derived daily. Verify current prices on the vendor's pricing page. Last verified 2026-07-16.
Related agents
Agent relevance
No programmatic surfaces
Agentic-Commerce Readiness 44/100 · SURFACE-ONLY
Independent readiness for agent delegation & transaction. How it’s scored · check live
Dropped in as Markdown skills that any capable LLM agent (Claude Code, Codex, etc.) reads and executes. No install, no framework — maximally agent-portable.
Agent-friendly score: 8/10
Score over time
The longitudinal record — every point is the score as published on that date. Raw series.
