Twill

Coding · tested 2026-08-19 · re-test due 2026-11-19 · by the Hlido desk, not the vendor

In short: A 'software factory' that turns GitHub, Slack and Linear tasks into tested pull requests inside a warm, full-stack dev environment — a strong, concrete pitch that the public surface can't yet prove.

Quick answer

Twill scores 74/100 (STEADY) on Hlido’s independent, hands-on test (reviewed 2026-08-19). STEADY (74) for a coherent, concretely-specified autonomous-coding platform (isolated full-stack task environments, multi-repo reasoning, real integrations, a believable automation catalogue, own-key model routing) with Pricing: Paid.

Twill (twill.ai) markets a hosted 'software factory': a task from GitHub, Slack or Linear spins up Claude Code, Codex or OpenCode in its own isolated copy of your company's environment — repos cloned, dependencies installed, services warm, the app runnable — and returns a pull request with proof attached. The differentiators it names are concrete and credible as a design: multi-repo reasoning across frontend/backend/workers/infra, agents that can install packages, run Docker, seed databases, start dev servers and run tests inside an isolated task fork, and model routing that puts frontier models on hard tasks and cheaper open-source models (Qwen, Kimi, GLM) on routine work using your own keys at provider rates. The integration list (GitHub, Slack, Linear, Notion, Sentry, GCP, AWS, Asana, Datadog) and a specific, believable automation catalogue — Sentry triage-and-fix, daily GitHub issue triage, dependency-update PRs, flaky-test remediation, stale-PR cleanup — make the offering legible rather than hand-wavy, and MCP-server/skill extensibility is advertised. It is 'backed by' an investor and offers a free Pro tier for open source. The gap is the usual one for a demo-gated commercial product: everything here is the vendor's own description, there is no independent evidence or public case study, pricing sits behind the nav, and Hlido reviewed the marketing surface, not a running task. The 'proof attached to every PR' claim is the most interesting and the least verifiable from outside. A coherent, well-scoped pitch; treat the capabilities as claimed until demonstrated.

Why STEADY

STEADY (74) for a coherent, concretely-specified autonomous-coding platform (isolated full-stack task environments, multi-repo reasoning, real integrations, a believable automation catalogue, own-key model routing) with a legible value proposition — discounted to low-medium confidence because it is a demo-gated commercial surface with no public evidence, case studies or transparent pricing, and Hlido reviewed the marketing pages rather than a running task. The 'proof attached to every PR' claim is unverified.

What we saw

4 screenshots captured by the Hlido engine during the reviewed run (run-c82e612be4e3bb4b-twill-ai). Our own captures — not vendor marketing material.

Twill — run screenshot 1 (home.png)
home.png
Twill — run screenshot 2 (page_.png)
page_.png
Twill — run screenshot 3 (page_pricing.png)
page_pricing.png
Twill — run screenshot 4 (page_download.png)
page_download.png

What it does well

What it fails at

Best for

  • Engineering teams wanting to route routine work (triage, dependency updates, flaky-test fixes) to agents that open tested PRs
  • Teams whose tasks span multiple repos and need a full stack running to be done well
  • Cost-conscious adopters who want to run cheaper open-source models on their own keys for routine work
  • Open-source maintainers eligible for the free Pro tier

Not recommended for

  • Buyers who need independent evidence, case studies or transparent pricing before adopting
  • Teams that cannot grant a hosted agent access to run their full stack and open PRs
  • Anyone wanting a self-hosted, on-prem-only solution (this is a hosted factory)

Pricing & access

Derived from Hlido-held evidence only (engine checklist + editorial text); quotes are verbatim from the scorecard; not vendor-supplied; re-derived daily. Verify current prices on the vendor's pricing page. Last verified 2026-08-19.

Related agents

Agent relevance

CLI MCP Behavioral-testable

Agentic-Commerce Readiness 57/100 · INTEGRABLE

Independent readiness for agent delegation & transaction. How it’s scored · check live

Tasks are created from the web app, a desktop app, a CLI, or triggers in GitHub/Slack/Linear; each spins up Claude Code / Codex / OpenCode in an isolated full-stack environment and returns a PR. Extensible via MCP servers and skills; runs on your own model keys.

Agent-friendly score: 7/10

Evidence

scorecard.json · transparency passport · registry · methodology

More: compare agents · best of · developer tools · incident registry

Verdict by Hlido Editor, our automated editorial system · Method: public-surface-tier-2+editorial-narrative-v2 · Methodology version 2026.05 · Next review due 2026-11-19

How this page was produced. The scores, claim verdicts and evidence come from automated hands-on testing of the product’s public surface. The written analysis is drafted by an AI system, and pages publish without a person reviewing each one. Hlido publishes this record and answers for it — tell us if anything here is wrong and we will correct it.

Embed this trust badge

Hlido trust score

Live, always-current independent score — free to embed on your site or README. No vendor pays for placement.

Markdown

[![Hlido trust score](https://hlido.eu/badge/twill.svg)](https://hlido.eu/check/?agent=twill)

HTML

<a href="https://hlido.eu/check/?agent=twill"><img src="https://hlido.eu/badge/twill.svg" alt="Hlido trust score"></a>