HawkTalk
01 / 16
01HawkTalk · Software beyond the screen

Voice AI belongs on NPU silicon,
not commodity GPU.

A cloud voice-AI token API running on purpose-built NPU inference silicon — voice at economics GPU can't touch. We move inference off the datacenter's default hardware and onto the chips built for it.

Prepared for Ubiquity Ventures · hawk, Founder · hawktalk.ai
↓ scroll · → arrow keys
02The short version

Voice AI on NPU — a fraction of GPU cost.

HawkTalk sells voice-AI inference tokens on Qualcomm AI-100 NPU silicon, not commodity cloud GPU — the reliability production voice needs, at a cost curve GPU can't match. Raising a seed round to own the cost floor of voice AI.

NPU
the silicon — beyond the screen
Hawk Alpha
the models
$/token
the moat

That's the whole company in fifteen seconds. The rest is why it works, why now, and why me.

03The problem — the pain you know

Voice AI runs on the wrong silicon.

Today, voice gets billed on the most expensive, most contended chips in the cloud — frontier GPUs built for training and long-form reasoning. For voice, that's a structural mismatch:

Voice companies are subsidizing a GPU maker's margin on the calls they serve. It doesn't have to cost that — the workload just needs the right hardware.

04The insight

Voice is function-calling — and it's over-served by GPUs.

Voice turns are short and structured: "do X," "set Y," "find Z." That's function-calling, not essay generation — it doesn't need a frontier model on a frontier GPU. Purpose-built voice models on NPU inference silicon deliver the same answer for a fraction of the compute.

This is software beyond the screen — reaching past commodity cloud GPU onto specialized silicon most teams can't even build for. The non-consensus hardware frontier, hiding in plain sight.

05The window

It just opened — and the market proved it.

Supply arrived, demand peaking

  • NPU inference silicon (Qualcomm AI-100) is mature, available, cheap — and largely idle
  • GPU cost & scarcity at all-time highs; voice/agent volume climbing onto the meter

Non-GPU inference — just validated

  • Nvidia bought Groq's inference tech (~$20B), Dec 2025
  • Cerebras IPO'd (May 2026, $5.55B) on inference
  • Inference is now ⅔ of AI compute → 80% by 2027

The industry just blessed "inference on specialized silicon" at the $20B level — while voice on commodity NPU is still wide open. Nerdy, early, non-consensus — and investable right now.

06What we do

A voice-AI token API, on NPU.

HawkTalk is a drop-in cloud inference API for voice + function-calling — you send audio/text, you get structured tool calls back — served from our NPU fleet at NPU economics, in our cloud or dedicated in yours.

API

Voice tokens

Usage-based inference over an OpenAI-shaped API. Point your voice stack at us, cut the bill.

Capacity

Committed NPU

Reserve dedicated AI-100 capacity — flat, predictable cost with headroom for volume.

Private

In-VPC deploy

Regulated / high-trust buyers run the same models on NPU inside their own perimeter.

07The edge

NPU-native, where others only know CUDA.

The moat is depth on silicon almost no one builds for. Three layers:

The models are lean enough to run even on a phone — that's how efficient they are. On NPU, that efficiency becomes a cost curve no GPU stack can match.

08Unit economics · the point

Flat NPU capacity vs. a metered GPU bill.

GPU token APIs

  • Metered per token, forever
  • Priced on the scarcest, most expensive silicon
  • Cost scales 1:1 with usage — unbounded

HawkTalk on AI-100

  • Flat capacity — pennies-per-Mtoken at utilization
  • Inference-purpose silicon, idle and cheap
  • Margin widens as the fleet fills
Single AI-100 node
$329/mo
128 GB node
$549/mo
Quad node
$1,259/mo
Octo node
$2,499/mo

Qualcomm Cloud AI-100 list pricing. A flat node running lean voice models drives cost-per-Mtoken far below metered GPU — the whole business is the gap between those curves.

09Market

Big, fast — and the cost center is moving to us.

$3.5B → $47.5B
AI voice agents, 2026 → 2034 (~35% CAGR) — the bullseye
$18B → $82B
conversational AI, 2026 → 2034
⅔ → 80%
of AI compute is inference, 2026 → 2027

Inference is the recurring cost line under all of it. Whoever makes voice tokens cheapest sets the floor — and wins the segments where cost-per-token is the P&L: agent/call-center voice, voice-at-scale, regulated/private, devices & automotive.

10Landscape · defensibility

Why the giants leave this open.

The field

  • Chip giants (Groq→Nvidia, Cerebras) — frontier-LLM inference at 750 MW hyperscale, on their own exotic silicon.
  • Voice platforms (Deepgram, Vapi, Retell, Bland) — GPU-riding orchestration.
  • Cartesia — efficient voice models; the model/latency layer.

Why it's ours

  • Voice-FC on commodity NPU is too small for the giants to prioritize — a perfect seed wedge.
  • The platforms would buy cheaper tokens under them → channel, not competition.
  • We're the silicon-cost layer, not the model-latency layer — a different center of gravity we can even partner from.

We don't win on "cheaper than GPU" — that headline now belongs to $20B+ giants. We win the intersection nobody owns: cheap, reliable voice-function-calling tokens on commodity NPU.

11Business model

Infra-margin economics.

Usage

Token API

Billed per million tokens over a flat-cost NPU base. Gross margin widens with utilization.

Recurring

Reserved capacity

Committed NPU nodes for high-volume customers — predictable revenue, predictable cost.

Enterprise

Private NPU deploy

In-VPC inference for regulated buyers, at a premium. Land-and-expand.

Fixed-cost silicon, variable-priced tokens: the classic infrastructure gross-margin curve, compounding as the fleet fills.

12Why me · why now

I've shipped voice AI in production since 2024.

Why me

  • Since 2024 I've built & run production voice AI for PurlPal — outbound / agent voice at volume, in insurance: regulated, cost-sensitive, unforgiving.
  • I didn't read about the GPU-token bill — I paid it, at scale. That's what drove the NPU answer.
  • And the rare skill behind the cost curve: I compile for Qualcomm AI-100 (QPC), not just CUDA.

Why now

  • Voice AI is exploding — agents ~35% CAGR toward $47.5B; Deepgram just raised at $1.3B; the category is a funding frenzy.
  • Inference moved to specialized silicon (Nvidia→Groq $20B, Cerebras IPO) — the supply side arrived too.
  • The demand curve and my experience curve peak at the same moment.

Founder-market fit isn't a claim here — it's years of production scar tissue in the exact vertical (high-volume agent voice) where the cost problem hurts most, hitting a market that just went vertical.

13The durable moat

NPU depth + efficient-model IP + total focus.

14What has to be true

The three bets — and how we de-risk them.

These are the real risks. Naming them is the point — the whole plan is built to retire them in order.

15The ask

Own the cost floor of voice AI.

Raising a seed round to scale the NPU inference fleet, harden the token API, and land the first high-volume voice customers off the cost advantage.

NPU
the silicon — beyond the screen
Hawk Alpha
the models
$/token
the moat
hawk · hawktalk.ai · hawktalktech@gmail.com
16Launch plan

One gate between us and live: NPU capacity.

✓ Built — models · QPC compile · token API ▶ GATE — secure Cirrascale AI-100 capacity live — serving tokens scale — + AWS spot burst

The one hurdle

  • Securing Cirrascale AI-100 capacity is the critical path — the single gating item to launch.
  • Everything else is already built: the Hawk Alpha models, the QPC compile, the token API.
  • An execution gate, not a research risk — and exactly what the round unlocks.

The capacity plan

  • Baseline — reserved Cirrascale NPU: flat cost, the cost advantage the business rides on.
  • Burst — AWS spot absorbs demand spikes: elastic, pay-only-when-needed, so we never over-buy NPU.
  • Two sources → no single point of failure, capital-efficient from day one.

Reserved NPU for the floor, cloud spot for the peaks. Fund the capacity and we serve voice tokens on day one — nothing left to invent.