HawkTalk

Voice AI — Competitive Landscape Brief

Where cheap, reliable voice function-calling tokens on commodity NPU wins, in a market being reshaped by the shift from training to inference.

CONFIDENTIAL · 2026
hawktalk.ai/brief
01 · The market

Big, fast, and its cost center is moving to exactly where we play.

$18B → $82B
conversational AI, 2026 → 2034 (~21% CAGR)
$3.5B → $47.5B
AI voice agents, 2026 → 2034 (~35% CAGR) — our bullseye
⅔ → 80%
share of AI compute that is inference, 2026 → 2027
02 · The shift we ride

The "inference economy" is here, and non-GPU inference just got validated at the highest level — a tailwind, and proof the thesis is real:

03 · The field — three rings

Voice AI platforms / APIsRing 1

PlayerPosition
Deepgramvoice-infra incumbent · $130M C @ $1.3B
Cartesiaefficient real-time voice models (Sonic ~40ms) · closest comp
ElevenLabsTTS giant
Vapi / Retell / Blandagent orchestration on GPU · customers, not rivals
OpenAI Realtimethe default / price benchmark

Inference-API providersRing 3

Together · Fireworks · Baseten · DeepInfra · Groq Cloud — cheap tokens, mostly GPU, general-purpose. We are the voice-specialized, NPU version.

Inference siliconRing 2

Player2026 status
Groq (LPU)→ Nvidia, ~$20B
CerebrasIPO, $5.55B, OpenAI 750MW
SambaNovaIntel acquiring
d-Matrix / Etchedproduction-stage ASICs
Qualcomm AI-100our silicon — commodity, deployed, cheap

43 chip startups, ~$17.7B raised. We are not a chip company — we're the software+models layer that turns one cheap NPU into voice tokens.

04 · Where we sit — the wedge

We're not head-on with any one leader — we're the intersection nobody owns: voice + function-calling + grounded/grammar-safe output + the economics of commodity NPU (AI-100). The chip giants chase frontier-LLM inference at 750 MW hyperscale; the voice-agent platforms are GPU-riding orchestration; Cartesia plays the model/latency layer, not the silicon-cost layer. The "cheap, reliable voice-FC tokens on commodity NPU" square is open.

05 · Why the giants leave it open

Our position

  • Groq / Cerebras are built for frontier-LLM inference at hyperscale on their own exotic silicon. Voice-FC on commodity NPU is too small for them to prioritize — and a great seed company.
  • Voice-agent platforms (Vapi/Retell/Deepgram) ride GPU and would buy cheaper tokens under them. Channel, not competition.
  • Reliability moat: grammar-grounded decode → 100% valid output, 0 hallucinated tools, at any model size — a hard property GPU LLM APIs can't guarantee.

The real threats — and the answer

  • A chip cloud verticalizes into voice (e.g. Groq Cloud). → We're NPU-portable + voice-FC-specialized; we move faster in a niche they under-serve.
  • Cartesia shifts onto cheap silicon. → We're the silicon-cost layer, not the model-latency layer; different center of gravity, and we can partner.
  • Don't win on "cheaper than GPU" alone — that headline now belongs to $20B+ giants. Win on the voice-FC-on-commodity-NPU intersection, or not at all.
Sources: Fortune Business Insights · Research & Markets · FutureAGI · Retell · TrendForce · Fortune · Value Add VC  ·  Market figures are third-party estimates; ranges vary by firm.