The article discusses the implementation of Nvidia's PersonaPlex 7B model on Apple Silicon, enabling full-duplex speech-to-speech conversation in native Swift. The model processes audio tokens directly, allowing for faster-than-real-time conversation. The library, speech-swift, also supports streaming voice processing and has been optimized for performance. The model has been quantized to 4-bit, reducing its size from 16.7 GB to 5.3 GB.