Petals is a system that allows users to run large language models at home, similar to BitTorrent, by loading a small part of the model and joining a network of people serving the other parts. This approach enables fine-tuning and inference up to 10x faster than offloading. The system is designed for interactive applications such as chatbots and can process sensitive data in a private swarm. Petals relies on people sharing their GPUs to serve the models.