Twill
Coding · tested 2026-08-19 · re-test due 2026-11-19 · by the Hlido desk, not the vendor
In short: A 'software factory' that turns GitHub, Slack and Linear tasks into tested pull requests inside a warm, full-stack dev environment — a strong, concrete pitch that the public surface can't yet prove.
Quick answer
Twill scores 74/100 (STEADY) on Hlido’s independent, hands-on test (reviewed 2026-08-19). STEADY (74) for a coherent, concretely-specified autonomous-coding platform (isolated full-stack task environments, multi-repo reasoning, real integrations, a believable automation catalogue, own-key model routing) with Pricing: Paid.
Twill (twill.ai) markets a hosted 'software factory': a task from GitHub, Slack or Linear spins up Claude Code, Codex or OpenCode in its own isolated copy of your company's environment — repos cloned, dependencies installed, services warm, the app runnable — and returns a pull request with proof attached. The differentiators it names are concrete and credible as a design: multi-repo reasoning across frontend/backend/workers/infra, agents that can install packages, run Docker, seed databases, start dev servers and run tests inside an isolated task fork, and model routing that puts frontier models on hard tasks and cheaper open-source models (Qwen, Kimi, GLM) on routine work using your own keys at provider rates. The integration list (GitHub, Slack, Linear, Notion, Sentry, GCP, AWS, Asana, Datadog) and a specific, believable automation catalogue — Sentry triage-and-fix, daily GitHub issue triage, dependency-update PRs, flaky-test remediation, stale-PR cleanup — make the offering legible rather than hand-wavy, and MCP-server/skill extensibility is advertised. It is 'backed by' an investor and offers a free Pro tier for open source. The gap is the usual one for a demo-gated commercial product: everything here is the vendor's own description, there is no independent evidence or public case study, pricing sits behind the nav, and Hlido reviewed the marketing surface, not a running task. The 'proof attached to every PR' claim is the most interesting and the least verifiable from outside. A coherent, well-scoped pitch; treat the capabilities as claimed until demonstrated.
Why STEADY
STEADY (74) for a coherent, concretely-specified autonomous-coding platform (isolated full-stack task environments, multi-repo reasoning, real integrations, a believable automation catalogue, own-key model routing) with a legible value proposition — discounted to low-medium confidence because it is a demo-gated commercial surface with no public evidence, case studies or transparent pricing, and Hlido reviewed the marketing pages rather than a running task. The 'proof attached to every PR' claim is unverified.
What we saw
4 screenshots captured by the Hlido engine during the reviewed run (run-c82e612be4e3bb4b-twill-ai). Our own captures — not vendor marketing material.
What it does well
- Concrete, well-scoped pitch — triggers from GitHub/Slack/Linear produce tested PRs inside a warm, full-stack isolated environment (repos cloned, deps installed, services running)
- Multi-repo reasoning across frontend, backend, workers, infra and shared packages in one task
- Own-key model routing across providers (frontier models for hard tasks, cheaper open-source models for routine work at provider rates) — a real cost lever
- Broad, named integrations (GitHub, Slack, Linear, Notion, Sentry, GCP, AWS, Asana, Datadog) and MCP-server/skill extensibility
- A specific, believable automation catalogue (Sentry triage-and-fix, issue triage, dependency updates, flaky-test remediation, stale-PR cleanup) plus a free Pro tier for open source
What it fails at
- No independent evidence or public case studies — every capability is vendor-described
- The signature 'proof attached to every PR' claim is exactly what a surface review cannot verify
- Pricing is behind the nav, not on the captured surface — a transparency gap for buyers
- Surface-only review — Hlido did not run a task, so environment isolation, test execution and PR quality are unverified
Best for
- Engineering teams wanting to route routine work (triage, dependency updates, flaky-test fixes) to agents that open tested PRs
- Teams whose tasks span multiple repos and need a full stack running to be done well
- Cost-conscious adopters who want to run cheaper open-source models on their own keys for routine work
- Open-source maintainers eligible for the free Pro tier
Not recommended for
- Buyers who need independent evidence, case studies or transparent pricing before adopting
- Teams that cannot grant a hosted agent access to run their full stack and open PRs
- Anyone wanting a self-hosted, on-prem-only solution (this is a hosted factory)
Pricing & access
- ModelPaid
Derived from Hlido-held evidence only (engine checklist + editorial text); quotes are verbatim from the scorecard; not vendor-supplied; re-derived daily. Verify current prices on the vendor's pricing page. Last verified 2026-08-19.
Related agents
Agent relevance
CLI MCP Behavioral-testable
Agentic-Commerce Readiness 57/100 · INTEGRABLE
Independent readiness for agent delegation & transaction. How it’s scored · check live
Tasks are created from the web app, a desktop app, a CLI, or triggers in GitHub/Slack/Linear; each spins up Claude Code / Codex / OpenCode in an isolated full-stack environment and returns a PR. Extensible via MCP servers and skills; runs on your own model keys.
Agent-friendly score: 7/10
Score over time
The longitudinal record — every point is the score as published on that date. Raw series.
Evidence
- Concrete, well-scoped pitch — triggers from GitHub/Slack/Linear produce tested PRs inside a warm, full-stack isolated environment (repos cloned, deps installed, services running) — source (2026-08-19) verified
- Multi-repo reasoning across frontend, backend, workers, infra and shared packages in one task — source (2026-08-19) verified
- Own-key model routing across providers (frontier models for hard tasks, cheaper open-source models for routine work at provider rates) — a real cost lever — source (2026-08-19) verified
- Broad, named integrations (GitHub, Slack, Linear, Notion, Sentry, GCP, AWS, Asana, Datadog) and MCP-server/skill extensibility — source (2026-08-19) verified
- Hands-on runtime behaviour (executing the tool / a live task) — source (2026-08-19)



