The article discusses a demo that uses GPT-4 to process video and audio input, similar to Google's Gemini demo, but with actual functionality. The demo's creator used a simple technique to feed images and text into GPT-4, achieving impressive results. The cost of using GPT-4 for this demo was relatively low, at $0.47 for 77 requests. The demo's success has sparked discussion about the potential of GPT-4 and other language models for multimodal tasks.