I Built a Fully Local Voice AI Stack on a CPU With No GPU: Here Are the Real Numbers
A complete voice AI stack, speech to text, a 2B language model, and text to speech, running fully local on a 2013-era CPU server with no GPU. I benchmarked every stage, audited the marketing claims, and the "150 tokens per second" pitch came out 13.6x slower in reality. Here is what a local voice assistant actually costs and does.
