Ferret is a multimodal large language model that can understand spatial referring and grounding in images. It uses a hybrid region representation and spatial-aware visual sampler to achieve fine-grained and open-vocabulary referring and grounding. The model is trained on a large-scale dataset and can be used for various tasks such as image description and question answering. Ferret is released as an open-source project with pre-trained models and evaluation benchmarks.