The author built a voice agent from scratch, achieving sub-500ms latency, and discusses the technical challenges and solutions involved. The project highlights the importance of orchestration and latency optimization in voice agents. The author used a combination of services, including Deepgram's Flux for turn detection and Groq's llama-3.3-70b for language modeling. The custom implementation outperformed off-the-shelf platforms like Vapi in terms of latency.