A new model, Llama 3-V, is claimed to match the performance of GPT-4-V with a significantly smaller model size and a fine-tuning cost of $500. The model is built on top of Llama3 and adds a vision encoder. However, the article's claims are met with skepticism by some commenters, who question the validity of the comparison and the potential for overfitting. The model's performance is also compared to other models, such as InternVL and CogVLM.