ViperGPT is a framework that leverages code-generation models to compose vision-and-language models into subroutines to produce a result for any query. It utilizes a provided API to access available modules and composes them by generating Python code that is later executed. This approach requires no further training and achieves state-of-the-art results across various complex visual tasks. ViperGPT can perform logical operations, spatial understanding, and knowledge access, among other capabilities.