HawkTalk
Voice AI — Competitive Landscape Brief
Where cheap, reliable voice function-calling tokens on commodity NPU wins, in a market being reshaped by the shift from training to inference.
CONFIDENTIAL · 2026
hawktalk.ai/brief
01 · The market
Big, fast, and its cost center is moving to exactly where we play.
$18B → $82B
conversational AI, 2026 → 2034 (~21% CAGR)
$3.5B → $47.5B
AI voice agents, 2026 → 2034 (~35% CAGR) — our bullseye
⅔ → 80%
share of AI compute that is inference, 2026 → 2027
02 · The shift we ride
The "inference economy" is here, and non-GPU inference just got validated at the highest level — a tailwind, and proof the thesis is real:
- Nvidia licensed Groq's inference tech + core team for ~$20B (Dec 2025) — the "cheap fast non-GPU inference" story is now blessed by the incumbent itself.
- Cerebras IPO'd (May 2026), raising $5.55B — largest US tech IPO since Uber; OpenAI 750MW deal, AWS integrating.
- Inference specialization rewards purpose-built silicon over GPUs. The market is validating our exact bet in real time.
03 · The field — three rings
Voice AI platforms / APIsRing 1
| Player | Position |
| Deepgram | voice-infra incumbent · $130M C @ $1.3B |
| Cartesia | efficient real-time voice models (Sonic ~40ms) · closest comp |
| ElevenLabs | TTS giant |
| Vapi / Retell / Bland | agent orchestration on GPU · customers, not rivals |
| OpenAI Realtime | the default / price benchmark |
Inference-API providersRing 3
Together · Fireworks · Baseten · DeepInfra · Groq Cloud — cheap tokens, mostly GPU, general-purpose. We are the voice-specialized, NPU version.
Inference siliconRing 2
| Player | 2026 status |
| Groq (LPU) | → Nvidia, ~$20B |
| Cerebras | IPO, $5.55B, OpenAI 750MW |
| SambaNova | Intel acquiring |
| d-Matrix / Etched | production-stage ASICs |
| Qualcomm AI-100 | our silicon — commodity, deployed, cheap |
43 chip startups, ~$17.7B raised. We are not a chip company — we're the software+models layer that turns one cheap NPU into voice tokens.
04 · Where we sit — the wedge
We're not head-on with any one leader — we're the intersection nobody owns:
voice + function-calling + grounded/grammar-safe output + the economics of commodity NPU (AI-100). The chip giants chase frontier-LLM inference at 750 MW hyperscale; the voice-agent platforms are GPU-riding orchestration; Cartesia plays the model/latency layer, not the silicon-cost layer. The "cheap, reliable voice-FC tokens on commodity NPU" square is open.
05 · Why the giants leave it open
Our position
- Groq / Cerebras are built for frontier-LLM inference at hyperscale on their own exotic silicon. Voice-FC on commodity NPU is too small for them to prioritize — and a great seed company.
- Voice-agent platforms (Vapi/Retell/Deepgram) ride GPU and would buy cheaper tokens under them. Channel, not competition.
- Reliability moat: grammar-grounded decode → 100% valid output, 0 hallucinated tools, at any model size — a hard property GPU LLM APIs can't guarantee.
The real threats — and the answer
- A chip cloud verticalizes into voice (e.g. Groq Cloud). → We're NPU-portable + voice-FC-specialized; we move faster in a niche they under-serve.
- Cartesia shifts onto cheap silicon. → We're the silicon-cost layer, not the model-latency layer; different center of gravity, and we can partner.
- Don't win on "cheaper than GPU" alone — that headline now belongs to $20B+ giants. Win on the voice-FC-on-commodity-NPU intersection, or not at all.