Voice AI belongs on NPU silicon, not commodity GPU.
A cloud voice-AI token API running on purpose-built NPU inference silicon — voice at economics GPU can't touch. We move inference off the datacenter's default hardware and onto the chips built for it.
Prepared for Ubiquity Ventures · hawk, Founder · hawktalk.ai
↓ scroll · → arrow keys
02The short version
Voice AI on NPU — a fraction of GPU cost.
HawkTalk sells voice-AI inference tokens on Qualcomm AI-100 NPU silicon, not commodity cloud GPU — the reliability production voice needs, at a cost curve GPU can't match. Raising a seed round to own the cost floor of voice AI.
NPU
the silicon — beyond the screen
Hawk Alpha
the models
$/token
the moat
That's the whole company in fifteen seconds. The rest is why it works, why now, and why me.
03The problem — the pain you know
Voice AI runs on the wrong silicon.
Today, voice gets billed on the most expensive, most contended chips in the cloud — frontier GPUs built for training and long-form reasoning. For voice, that's a structural mismatch:
Cost — tokens are metered on GPU pricing; the bill scales with usage, unbounded.
Margin — token cost eats the unit economics of the products built on top.
Capacity — GPUs are scarce and volatile; you fight the rest of AI for the same racks.
Voice companies are subsidizing a GPU maker's margin on the calls they serve. It doesn't have to cost that — the workload just needs the right hardware.
04The insight
Voice is function-calling — and it's over-served by GPUs.
Voice turns are short and structured: "do X," "set Y," "find Z." That's function-calling, not essay generation — it doesn't need a frontier model on a frontier GPU. Purpose-built voice models on NPU inference silicon deliver the same answer for a fraction of the compute.
This is software beyond the screen — reaching past commodity cloud GPU onto specialized silicon most teams can't even build for. The non-consensus hardware frontier, hiding in plain sight.
05The window
It just opened — and the market proved it.
Supply arrived, demand peaking
NPU inference silicon (Qualcomm AI-100) is mature, available, cheap — and largely idle
GPU cost & scarcity at all-time highs; voice/agent volume climbing onto the meter
Non-GPU inference — just validated
Nvidia bought Groq's inference tech (~$20B), Dec 2025
Cerebras IPO'd (May 2026, $5.55B) on inference
Inference is now ⅔ of AI compute → 80% by 2027
The industry just blessed "inference on specialized silicon" at the $20B level — while voice on commodity NPU is still wide open. Nerdy, early, non-consensus — and investable right now.
06What we do
A voice-AI token API, on NPU.
HawkTalk is a drop-in cloud inference API for voice + function-calling — you send audio/text, you get structured tool calls back — served from our NPU fleet at NPU economics, in our cloud or dedicated in yours.
API
Voice tokens
Usage-based inference over an OpenAI-shaped API. Point your voice stack at us, cut the bill.
Capacity
Committed NPU
Reserve dedicated AI-100 capacity — flat, predictable cost with headroom for volume.
Private
In-VPC deploy
Regulated / high-trust buyers run the same models on NPU inside their own perimeter.
07The edge
NPU-native, where others only know CUDA.
The moat is depth on silicon almost no one builds for. Three layers:
We compile for AI-100 (QPC) — the specialized-hardware skill that unlocks the whole cost curve. Most teams can't leave the GPU.
Purpose-built voice models (Hawk Alpha) — tuned for function-calling to match models 5×+ larger, at a fraction of the compute.
Grounded, grammar-constrained decode — 100% valid structured output, 0 hallucinated tools, provably. The reliability production voice needs, that raw GPU LLMs can't guarantee.
The models are lean enough to run even on a phone — that's how efficient they are. On NPU, that efficiency becomes a cost curve no GPU stack can match.
08Unit economics · the point
Flat NPU capacity vs. a metered GPU bill.
GPU token APIs
Metered per token, forever
Priced on the scarcest, most expensive silicon
Cost scales 1:1 with usage — unbounded
HawkTalk on AI-100
Flat capacity — pennies-per-Mtoken at utilization
Inference-purpose silicon, idle and cheap
Margin widens as the fleet fills
Single AI-100 node
$329/mo
128 GB node
$549/mo
Quad node
$1,259/mo
Octo node
$2,499/mo
Qualcomm Cloud AI-100 list pricing. A flat node running lean voice models drives cost-per-Mtoken far below metered GPU — the whole business is the gap between those curves.
09Market
Big, fast — and the cost center is moving to us.
$3.5B → $47.5B
AI voice agents, 2026 → 2034 (~35% CAGR) — the bullseye
$18B → $82B
conversational AI, 2026 → 2034
⅔ → 80%
of AI compute is inference, 2026 → 2027
Inference is the recurring cost line under all of it. Whoever makes voice tokens cheapest sets the floor — and wins the segments where cost-per-token is the P&L: agent/call-center voice, voice-at-scale, regulated/private, devices & automotive.
10Landscape · defensibility
Why the giants leave this open.
The field
Chip giants (Groq→Nvidia, Cerebras) — frontier-LLM inference at 750 MW hyperscale, on their own exotic silicon.
Cartesia — efficient voice models; the model/latency layer.
Why it's ours
Voice-FC on commodity NPU is too small for the giants to prioritize — a perfect seed wedge.
The platforms would buy cheaper tokens under them → channel, not competition.
We're the silicon-cost layer, not the model-latency layer — a different center of gravity we can even partner from.
We don't win on "cheaper than GPU" — that headline now belongs to $20B+ giants. We win the intersection nobody owns: cheap, reliable voice-function-calling tokens on commodity NPU.
11Business model
Infra-margin economics.
Usage
Token API
Billed per million tokens over a flat-cost NPU base. Gross margin widens with utilization.
In-VPC inference for regulated buyers, at a premium. Land-and-expand.
Fixed-cost silicon, variable-priced tokens: the classic infrastructure gross-margin curve, compounding as the fleet fills.
12Why me · why now
I've shipped voice AI in production since 2024.
Why me
Since 2024 I've built & run production voice AI for PurlPal — outbound / agent voice at volume, in insurance: regulated, cost-sensitive, unforgiving.
I didn't read about the GPU-token bill — I paid it, at scale. That's what drove the NPU answer.
And the rare skill behind the cost curve: I compile for Qualcomm AI-100 (QPC), not just CUDA.
Why now
Voice AI is exploding — agents ~35% CAGR toward $47.5B; Deepgram just raised at $1.3B; the category is a funding frenzy.
Inference moved to specialized silicon (Nvidia→Groq $20B, Cerebras IPO) — the supply side arrived too.
The demand curve and my experience curve peak at the same moment.
Founder-market fit isn't a claim here — it's years of production scar tissue in the exact vertical (high-volume agent voice) where the cost problem hurts most, hitting a market that just went vertical.
13The durable moat
NPU depth + efficient-model IP + total focus.
NPU-native by default — the AI-100 / QPC compilation depth that is the cost curve, where the rest of AI only knows CUDA.
Efficient-model IP — Hawk Alpha: function-calling voice models tuned to punch far above their size, on lean silicon.
Reproducible pipeline — new verticals and languages are a training run, not a rewrite.
Solo, on purpose — one technical decision-maker, no committee, moving at the speed the window demands. Building the first hires around a proven wedge.
14What has to be true
The three bets — and how we de-risk them.
NPU supply & partnership scale. → AI-100 capacity is available today via multiple hosts; we design for portability across inference silicon, so we're not single-vendor-locked.
Model quality holds at production scale. → Grounded + grammar-constrained decode makes correctness a hard property, not a hope — it holds regardless of model size.
We win design-partner volume. → We land where cost-per-token is already the P&L (agent/voice-at-scale), where a real cost delta closes deals on its own.
These are the real risks. Naming them is the point — the whole plan is built to retire them in order.
15The ask
Own the cost floor of voice AI.
Raising a seed round to scale the NPU inference fleet, harden the token API, and land the first high-volume voice customers off the cost advantage.