State of the Hawk · benchmark & status report

HawkTalkLive is live, sovereign, and running at SOTA speed — measured, not projected.

A real product serving on api.hawktalk.ai over TLS — authenticated, metered, full-duplex voice AI on owned silicon. Every number on this page came off a run artifact from the overnight build, dated 2026-08-11. Where something is still roadmap, it says so.

How to read this page: anything tagged measured is artifact-backed from real runs; live means deployed and verified on the public endpoint; next and roadmap are the honest plan, not claims.
The headline benchmark

Voice-on-NPU: 6× the seats, 6× cheaper, still under the SOTA latency floor

The overnight work moved the whole voice cascade onto the NPU. One box went from 2 concurrent voice seats to a measured knee of 12 — while keeping felt first-audio at ~590ms, inside the 450–600ms floor the best cloud cascades operate in. This is the number that changes the economics.

2 → 12
Voice seats per box (measured knee, was CPU-cascade baseline)
measured load-knee run
590 ms p50
Felt first-audio at the 12-seat knee — inside the SOTA cascade floor (450–600ms)
measured at-knee latency run
$0.063 /seat-hr
Was $0.38 — 6× cheaper per concurrent voice seat
measured derived from knee run
$7.80 /Mtok
Live-surface cost per million tokens — was $132, 17× cheaper
measured derived from knee run
The measured numbers

Every stage of the loop, on the clock

The full pipeline — hear, feel, think, speak, interrupt — measured stage by stage on the serving box.

StageResultWhat it meansStatus
Speech-to-text
gemma-native audio tower, on NPU
436ms warm TTFT
375ms p95, flat 1→16 seats
Whisper eliminated from the loop entirely; 42.6× faster than host-CPU decode, and latency stays flat as seats load up. measured
Emotion sensing (SER)
wav2vec2, on NPU
44.7ms (vs 364ms CPU)
66ms end-to-end via /v1/ser
8.1× faster on a dedicated NPU core partition — live in production, with a CPU fallback armed. measured live
Text-to-speech — fast tier
Matcha (non-autoregressive)
RTF 0.175 Renders speech ~5.7× faster than realtime — the volume/realtime tier. measured
Text-to-speech — mid tier
Kokoro
RTF 1.31 The balanced quality tier. measured
LLM
gemma-E2B
536ms TTFT
29–70 tok/s
Fast enough that the model is not the bottleneck in the loop. measured
Barge-in (interrupt the AI mid-sentence) ~0ms server decision 12/12 hermetic checks green — the AI stops the instant you speak. measured live
Echo cancellation (AEC) 11.7dB ERLE
0.93 user-voice correlation in double-talk
Speakerphone-style conversation without the AI hearing itself — built from scratch for this stack. measured live
KV-paging (memory headroom) ~2× KV capacity (INT8) Design complete, off-box tests green; on-box validation queued behind a free accelerator core. next
Why the STT number matters most: replacing Whisper with the gemma-native audio tower on the NPU is what collapsed the front of the loop — 42.6× faster than host-CPU decode, and it holds flat under concurrent load instead of degrading.
The spectrum

One sovereign stack, every rung of the ladder

The same pipeline scales from a phone in your pocket to a dense cloud box. Two rungs are measured today; the rest is the same code with more silicon under it.

L0Edge
Phone NPU (Hexagon). One seat per device, fully private, works offline. Zero marginal serving cost — this rung is the moat.
$0 marginalroadmap port in progress
L1Baseline
One inf2.xlarge, CPU cascade. Where the night started: 2 seats at $0.38/seat-hour.
2 seats · $0.38measured baseline
L2Voice-on-NPU
Same box, cascade moved onto the NPU. 12 seats at 590ms first-audio — the rung we serve on today.
12 seats · $0.063measured tonight
L3Batched
+ decoder continuous-batching + KV-paging. The next lever, on the same box.
16–32+ seats · ~$0.03next projected
L4Dense
One inf2.24xlarge, data-parallel ×12. Replicate the proven 12-seat pipeline — no exotic parallelism required.
~144 seatsroadmap projected
Live now

Deployed and verified on the public endpoint

All of this is serving today on api.hawktalk.ai — not a demo build, the production surface.

The honest roadmap

What's next — and what isn't ready yet

The discipline that produced tonight's numbers is the same one that keeps unfinished things off the "done" list. Here is the plan, in order, and the one honest caveat.

  1. Decoder continuous-batching — the single biggest lever left on the box: takes the measured 12 seats toward 16–32 on the same hardware. next
  2. KV-paging on-box validation — design and off-box tests are already green; validating live buys ~2× memory headroom. next
  3. Morpheus premium voice tier — see the caveat below. roadmap
  4. Edge port to phone NPU — the fused voice loop onto Hexagon: the sovereign, offline, zero-marginal-cost rung of the spectrum. roadmap
  5. Dense scaling on inf2.24xlarge — replicate the proven pipeline data-parallel to ~144 seats per box. roadmap
The honest caveat — Morpheus (premium voice) is not demo-ready. The gold-tier neural voice currently renders at 0.52× realtime with stability issues, so it stays labeled "in development" rather than shipping half-working. The measurements isolated the actual cause (a parallelism-latency bound, not compute), so the fix path is known: a re-export with the right parallelism plus stability work. Meanwhile Matcha and Kokoro are the reliable production tiers — and emotion detection ships today (45ms SER), while emotion-expressive voice is roadmap. Finding this before demo day is the system working, not failing.
The moat

Structural, not incremental

The big labs are structurally barred from this position: they can't run on your phone, can't hand you the weights, and can't let you own the stack that serves your traffic. HawkTalk is the same sovereign pipeline on cloud silicon and edge silicon, with ownable weights at every rung. That's the pitch behind the numbers above — a running product asking for fuel, not a vision asking for a chance.