The article discusses how ChatGPT can serve 700 million users weekly despite the difficulty of running a single GPT-4 model locally. The author wonders what engineering tricks make this possible at such a massive scale while keeping latency low. The discussion revolves around the use of huge GPU clusters, model optimizations, and custom hardware. The author is curious to hear insights from people who have built large-scale ML systems.