Why we trained our entire voice stack from scratch to run on CPUs

by•

We're a small team in Madrid. When we started Lokutor, the usual stack - rented STT, rented LLM, rented TTS, all GPU-bound - didn't work for what our customers needed: unit economics that scale, and data residency for regulated sectors in Spain and LatAm.

So we trained everything ourselves:

- Vela: turn-taking

- Psst: noise suppression

- Conv: STT (2.0: 9.6% WER, 89x realtime on CPU)

- Our own fine-tuned LLM

- Versa: TTS (2.0: zero-shot voice cloning at 44.1kHz, first audio in ~100-280ms on a single CPU thread)

50,000 GPU hours on MareNostrum 5 (EuroHPC) went into training - so that inference never needs a GPU again. The result: ~12 concurrent voice sessions on a plain 8-vCPU x86 node (~20 on ARM64), sub-450ms conversations, deployable on-prem.

The Go orchestrator is open source and there's a LiveKit Agents plugin. Happy to go deep on any part - training models specifically for CPU targets was the hard, fun part.

2 views

Add a comment

Replies

Be the first to comment