Why we trained our entire voice stack from scratch to run on CPUs
We're a small team in Madrid. When we started Lokutor, the usual stack - rented STT, rented LLM, rented TTS, all GPU-bound - didn't work for what our customers needed: unit economics that scale, and data residency for regulated sectors in Spain and LatAm.
So we trained everything ourselves:
- Vela: turn-taking
- Psst: noise suppression
- Conv: STT (2.0: 9.6% WER, 89x realtime on CPU)
- Our own fine-tuned LLM
- Versa: TTS (2.0: zero-shot voice cloning at 44.1kHz, first audio in ~100-280ms on a single CPU thread)
50,000 GPU hours on MareNostrum 5 (EuroHPC) went into training - so that inference never needs a GPU again. The result: ~12 concurrent voice sessions on a plain 8-vCPU x86 node (~20 on ARM64), sub-450ms conversations, deployable on-prem.
The Go orchestrator is open source and there's a LiveKit Agents plugin. Happy to go deep on any part - training models specifically for CPU targets was the hard, fun part.

Replies