The article discusses a project that uses GPT-4 Vision with Vimium to browse the web. The project aims to explore the possibility of using multimodal models to interact with the web. The author provides a script that uses Vimium to give the model a way to interact with the web. The project is still in its early stages, and the author has listed several potential next steps.