news.volyx.in

Ask HN: How can ChatGPT serve 700M users when I can't run one GPT-4 locally?

574 points by superasn · 354 days ago · 365 comments on HN

Article summary

The article discusses how ChatGPT can serve 700 million users weekly despite the difficulty of running a single GPT-4 model locally. The author wonders what engineering tricks make this possible at such a massive scale while keeping latency low. The discussion revolves around the use of huge GPU clusters, model optimizations, and custom hardware. The author is curious to hear insights from people who have built large-scale ML systems.

Main themes

  • Large-scale ML systems
  • GPU clusters and custom hardware
  • Model optimization
  • Cloud computing and infrastructure
  • AI and machine learning

What commenters say

  • The key to running large-scale ML systems lies in the use of huge GPU clusters and custom hardware, which allows for massive parallel processing and faster computation.
  • Trade secrets and proprietary technology play a significant role in the success of companies like OpenAI and Google in building large-scale ML systems.
  • Google's custom TPUs are a major factor in their ability to efficiently run large-scale ML models, and they have a significant advantage over other companies in this regard.
  • The use of custom hardware like TPUs and Inferentia chips can provide a significant boost to the performance of ML models, but it also requires significant investment and expertise.
  • Some argue that Google's reputation for abandoning products and lack of customer support may hinder their ability to successfully host businesses on their cloud platform, despite their technical advantages.
  • Others believe that Google's talent and resources will ultimately give them an edge in the AI and ML market, despite potential challenges and setbacks.
  • The development of custom hardware for ML is a complex and costly process, and even large companies like Intel have struggled to succeed in this area.
  • The ability to efficiently run large-scale ML models is not just a matter of hardware, but also requires significant advances in software and model optimization.