Cactus
Infrastructure · tested 2026-08-21 · re-test due 2026-11-21 · by the Hlido desk, not the vendor
In short: On-device AI for phones, wearables and microcontrollers, with a cloud fallback — a focused, credible edge-inference stack (runtime + a 14MB agentic model) with 5.9k+ stars.
4 PASS · 0 FAIL of 4 public-surface claims
Quick answer
Cactus scores 73/100 (STEADY) on Hlido’s independent, hands-on test (reviewed 2026-08-21). STEADY (73) for a focused, credible edge-inference stack with a clear three-part architecture (Hybrid / Needle / Engine), a differentiated 14MB agentic model, and real signals of active development (5.9k+ stars, docs, ch
Cactus is an edge/on-device inference stack with three clearly separated pieces: Cactus Hybrid (post-trained models that know when they're wrong and escalate to the cloud), Cactus Needle (a 14MB agentic LLM doing tool-calling, device use and structured extraction on tiny devices), and Cactus Engine (a resource-constrained runtime with quantization tuned for battery, speed and memory). The positioning is sharp and technically coherent — on-device AI for phones, wearables, robots, home assistants and microcontrollers, with cloud fallback rather than pure-local dogma — and 5.9k+ GitHub stars plus visible docs, blog, a changelog and a compare page signal a real, actively developed project. For agent builders, a 14MB model that does tool-calling and structured extraction locally is a genuinely interesting primitive. The limits of a homepage review apply: the actual quantization quality, inference speed, battery claims and the 'knows when it's wrong' hybrid-escalation behaviour are exactly the things that decide whether an edge stack is usable, and none can be verified without hands-on benchmarking. Promising and well-scoped; verify the performance claims on your target hardware.
Why STEADY
STEADY (73) for a focused, credible edge-inference stack with a clear three-part architecture (Hybrid / Needle / Engine), a differentiated 14MB agentic model, and real signals of active development (5.9k+ stars, docs, changelog, compare page). Not higher because the load-bearing claims — quantization quality, battery/speed, and the hybrid 'knows when it's wrong' escalation — are unverifiable from the public surface and are precisely what determine on-device usability. Not FADING because the product is concrete, technically coherent and clearly maintained.
Public-surface checklist
- PASS Homepage loads (required)
- PASS Primary value prop (required) — On-device AI for phones, wearables and microcontrollers, with a cloud fallback —
- PASS Cta present
- PASS Evidence or demo — 4 screenshot(s) captured
What we saw
4 screenshots captured by the Hlido engine during the reviewed run (run-9f5b6c1990ff7b1d-cactuscompute-com). Our own captures — not vendor marketing material.
What it does well
- Sharp, coherent positioning: on-device AI for phones/wearables/robots/microcontrollers with cloud fallback (not pure-local dogma)
- Cactus Needle — a 14MB agentic LLM doing tool-calling, device use and structured extraction on tiny devices — is a genuinely interesting edge primitive
- Clear three-part separation (Hybrid models, Needle model, Engine runtime) that maps to real deployment decisions
- Signals of an active, real project: 5.9k+ GitHub stars, docs, blog, changelog and a compare page
What it fails at
- The decisive claims — quantization quality, inference speed, battery consumption — are asserted ('SOTA') but unverifiable from the homepage
- The Hybrid 'models that know when they're wrong and request cloud help' behaviour is unproven on the surface and hard to guarantee
- Edge deployment success is highly hardware-dependent; homepage says nothing about supported chips/OS matrix
- No pricing/licensing detail captured on the surface for the commercial pieces
Red flags
- Edge-inference value lives entirely in real-world quantization quality, speed and battery on YOUR hardware — these 'SOTA' claims can't be verified from the site, so benchmark before committing.
- The Hybrid self-doubt-and-escalate mechanism is a strong claim with no surface evidence of how reliable it is.
Best for
- Mobile/embedded developers who need local inference with an optional cloud fallback
- Agent builders wanting a tiny on-device model that can tool-call and extract structured output offline
- Products with privacy, latency or connectivity constraints that rule out cloud-only inference
Not recommended for
- Teams that need certified performance/battery numbers before adoption (benchmark on target hardware first)
- Server-side/cloud-only workloads where on-device constraints add no value
- Anyone needing broad, documented hardware-compatibility guarantees up front
Compared to
-
Ollama
embedded-edge-inference-vs-desktop-local-inference
Ollama makes it trivial to run open models locally on desktops/servers; Cactus targets the harder edge tier — phones, wearables, microcontrollers — with a purpose-built runtime, aggressive quantization and a tiny agentic model plus cloud fallback. Choose Ollama for easy local desktop inference, Cactus for genuinely on-device/embedded deployment.
Agent relevance
API SDK Behavioral-testable
Agentic-Commerce Readiness 58/100 · INTEGRABLE
Independent readiness for agent delegation & transaction. How it’s scored · check live
Provides an on-device inference runtime (Cactus Engine) and models (Needle/Hybrid) with SDK/docs for integrating local tool-calling and structured extraction into apps and agents. Behaviour is testable via the open-source runtime, though real performance is hardware-dependent. Relevant to agents as an edge-inference primitive rather than a hosted agent surface.
Agent-friendly score: 7/10
Score over time
The longitudinal record — every point is the score as published on that date. Raw series.
Evidence
- On-device AI with cloud fallback for phones, wearables, robots, home assistants and microcontrollers — source (2026-08-21) verified
- Cactus Needle: a 14MB agentic LLM with tool calling, device use and structured extraction — source (2026-08-21) verified
- Cactus Hybrid: post-trained models that know when they're wrong and request cloud help — source (2026-08-21) verified
- Cactus Engine: resource-constrained inference runtime with quantization for battery/speed — source (2026-08-21) verified
- 5.9k+ GitHub stars; docs, blog, changelog and compare pages present — source (2026-08-21) verified



