Cactus | Y Combinator
Summer 2025 batch; low-latency on-device AI engine for mobile and wearables.
Weak funding evidence: mention does not include a parsed raise amount/round and may refer to non-funding context.
Loading startup
Market data is refreshed once per day from public sources. Information may be incomplete or outdated — verify independently before making decisions. This is not investment advice.
Evidence-bound summary — expand sections for movement, risks, and signals.
Memo snapshot · May 20, 2026, 6:21 PM
What would change this read
DealFlow OS uses public web data and automated enrichment. Research may be incomplete, outdated, or incorrect. Verify important information before making investment or outreach decisions.
TL;DR
Seed (YC)Cactus - On-device AI for Smartphones, Laptops & Edge One inference engine for on-device AI across smartphones, laptops, and edge hardware
Raised $500K across 1 funding round. Latest: $500K Pre-seed (Jun 2025). Investors: Y Combinator. (High).
Funding
Raised $500K across 1 funding round. Latest: $500K Pre-seed (Jun 2025). Investors: Y Combinator. (High).
Product / news
7 product/news‑styled row(s); headline risk without filings (High).
Verified facts
+5 more in Recent movement below
Summer 2025 batch; low-latency on-device AI engine for mobile and wearables.
Weak funding evidence: mention does not include a parsed raise amount/round and may refer to non-funding context.
No open roles indexed yet.
The index price and activity score are algorithmic estimates based on observed public company-level signals. They may be incomplete, stale, or inaccurate and are not investment, legal, tax, or business advice.
Cactus:reddit_credentials_not_configured
Source types found
Strongest / recent news-style rows
Wed, May 20, 06:21 PM · confidence 88%high quality
wikipedia · Sun, Jun 28, 08:24 PM · confidence 50%medium quality
wikipedia · Sun, Jun 28, 08:24 PM · confidence 50%medium quality
Newest first · 23 event(s)
Source: Blog
Deep dives into on-device AI, inference optimization, and running models on smartphones, laptops, and edge hardware.
Source: Homepage
One inference engine for on-device AI across smartphones, laptops, and edge hardware. Run LLMs, transcription, and embeddings locally with automatic cloud fallback.
Source: official_site
Cactus - On-device AI for Smartphones, Laptops & Edge Cactus Hybrid Cloud On-device AI Docs Blog Talk to us ⌘K Sign in Get Started [NEW] | Needle: a 26M tool-calling model distilled from Gemini Read more Backed by On-device AI with cloud fallback Deploy s…
Source: official_site
Engineering Blog | Cactus Cactus Hybrid Cloud On-device AI Docs Blog Talk to us ⌘K Sign in Get Started [NEW] | Needle: a 26M tool-calling model distilled from Gemini Read more Cactus Blog Deep dives into on-device AI, inference optimization, and the engineeri…
Summer 2025 batch; low-latency on-device AI engine for mobile and wearables.
One inference engine for on-device AI across hardware targets.
Source: Blog / news
A simplified offline variant of TurboQuant using Hadamard rotation and per-group Lloyd-Max codebooks — 4× compression of per-layer embeddings in Gemma 4 E2B at +0.06 PPL.
Source: Blog / news
Review of NVIDIA's Parakeet-CTC-1.1B model running locally on Mac with Cactus. Architecture breakdown, benchmarks, and transcription use cases.
Source: Blog / news
Benchmarking Liquid's LFM-2.5-350m across seven devices with Cactus. INT8 quantization, single-core CPU decode, zero-copy loading, and why this configuration makes on-device inference practical.
Source: Blog / news
Review of LiquidAI's LFM2-24B-A2B mixture-of-experts model running locally on Mac with Cactus. Architecture breakdown, benchmarks, and coding agent use cases.
Source: Blog / news
How Cactus combines on-device and cloud inference for real-time speech transcription with sub-150ms latency and automatic cloud handoff for noisy audio.
Source: Blog / news
Gemma 4 runs natively on your device with real-time voice, vision, and audio, and routes hard problems to the cloud when it should.
Source: Blog / news
Deep dives into on-device AI, inference optimization, and running models on smartphones, laptops, and edge hardware.
Source: npm_registry
Create a new Stock Management System (Cactus)
Source: npm_registry
Cactus (needle-rs) tool-calling provider for @workglow/ai.
Source: npm_registry
Universal library used by both front end and back end components of Cactus. Aims to be a developer swiss army knife.
Source: npm_registry
Allows Cactus nodes to create DLT views using Cactus connectors
Source: npm_registry
Allows Cactus nodes to connect to a Fabric ledger.
Source: hackernews
Source: hackernews
Open-source low-latency mobile AI engine.
4 row(s)
The company's own site — the authoritative description of what they sell and to whom. Marketing-controlled, so treat claims as positioning rather than verified traction.
One inference engine for on-device AI across smartphones, laptops, and edge hardware. Run LLMs, transcription, and embeddings locally with automatic cloud fallback.
Why it matters: Primary source — the company's own positioning; best read for what they sell and to whom, not for traction claims.
Open source ↗Cactus - On-device AI for Smartphones, Laptops & Edge Cactus Hybrid Cloud On-device AI Docs Blog Talk to us ⌘K Sign in Get Started [NEW] | Needle: a 26M tool-calling model distilled from Gemini Read more Backed by On-device AI with cloud fallback Deploy s…
Why it matters: Primary source — the company's own positioning; best read for what they sell and to whom, not for traction claims.
Open source ↗Engineering Blog | Cactus Cactus Hybrid Cloud On-device AI Docs Blog Talk to us ⌘K Sign in Get Started [NEW] | Needle: a 26M tool-calling model distilled from Gemini Read more Cactus Blog Deep dives into on-device AI, inference optimization, and the engineeri…
Why it matters: Primary source — the company's own positioning; best read for what they sell and to whom, not for traction claims.
Open source ↗One inference engine for on-device AI across hardware targets.
Why it matters: Primary source — the company's own positioning; best read for what they sell and to whom, not for traction claims.
Open source ↗2 row(s)
Third-party press coverage. Independent reporting corroborates company claims; repeated coverage across outlets is a momentum signal.
Why it matters: Independent coverage — third-party corroboration of company claims; recurring coverage indicates rising visibility.
Open source ↗Why it matters: Independent coverage — third-party corroboration of company claims; recurring coverage indicates rising visibility.
Open source ↗1 row(s)
Funding announcements and investor-database records. The strongest public signal of capitalization: round, amount, and syndicate quality when disclosed.
Summer 2025 batch; low-latency on-device AI engine for mobile and wearables.
Why it matters: Funding signal — Pre-Seed per this source; verify against the linked original before relying on it.
Open source ↗8 row(s)
Public engineering activity. Sustained commits, releases, and stars indicate real product development and, for dev tools, developer adoption.
Create a new Stock Management System (Cactus)
Why it matters: Engineering signal — public repo activity evidences active development and possible developer adoption.
Open source ↗Cactus (needle-rs) tool-calling provider for @workglow/ai.
Why it matters: Engineering signal — public repo activity evidences active development and possible developer adoption.
Open source ↗Universal library used by both front end and back end components of Cactus. Aims to be a developer swiss army knife.
Why it matters: Engineering signal — public repo activity evidences active development and possible developer adoption.
Open source ↗Allows Cactus nodes to create DLT views using Cactus connectors
Why it matters: Engineering signal — public repo activity evidences active development and possible developer adoption.
Open source ↗Allows Cactus nodes to connect to a Fabric ledger.
Why it matters: Engineering signal — public repo activity evidences active development and possible developer adoption.
Open source ↗Why it matters: Engineering signal — public repo activity evidences active development and possible developer adoption.
Open source ↗Why it matters: Engineering signal — public repo activity evidences active development and possible developer adoption.
Open source ↗Open-source low-latency mobile AI engine.
Why it matters: Engineering signal — public repo activity evidences active development and possible developer adoption.
Open source ↗8 row(s)
Company blog and newsletters. Shipping cadence and technical depth of posts hint at product velocity and team quality.
Deep dives into on-device AI, inference optimization, and running models on smartphones, laptops, and edge hardware.
Why it matters: Company publishing — post cadence and depth hint at product velocity.
Open source ↗A simplified offline variant of TurboQuant using Hadamard rotation and per-group Lloyd-Max codebooks — 4× compression of per-layer embeddings in Gemma 4 E2B at +0.06 PPL.
Why it matters: Company publishing — post cadence and depth hint at product velocity.
Open source ↗Review of NVIDIA's Parakeet-CTC-1.1B model running locally on Mac with Cactus. Architecture breakdown, benchmarks, and transcription use cases.
Why it matters: Company publishing — post cadence and depth hint at product velocity.
Open source ↗Benchmarking Liquid's LFM-2.5-350m across seven devices with Cactus. INT8 quantization, single-core CPU decode, zero-copy loading, and why this configuration makes on-device inference practical.
Why it matters: Company publishing — post cadence and depth hint at product velocity.
Open source ↗Review of LiquidAI's LFM2-24B-A2B mixture-of-experts model running locally on Mac with Cactus. Architecture breakdown, benchmarks, and coding agent use cases.
Why it matters: Company publishing — post cadence and depth hint at product velocity.
Open source ↗How Cactus combines on-device and cloud inference for real-time speech transcription with sub-150ms latency and automatic cloud handoff for noisy audio.
Why it matters: Company publishing — post cadence and depth hint at product velocity.
Open source ↗Gemma 4 runs natively on your device with real-time voice, vision, and audio, and routes hard problems to the cloud when it should.
Why it matters: Company publishing — post cadence and depth hint at product velocity.
Open source ↗Deep dives into on-device AI, inference optimization, and running models on smartphones, laptops, and edge hardware.
Why it matters: Company publishing — post cadence and depth hint at product velocity.
Open source ↗Sign in as an active team member to view private notes, watchlist controls, transcript evidence, and interaction history.