news.volyx.in

Gemini 2.5 Flash Image (developers.googleblog.com)

1093 points by meetpateltech · 335 days ago · 475 comments on HN

Article summary

Google has introduced Gemini 2.5 Flash Image, a state-of-the-art image generation and editing model that enables users to blend multiple images, maintain character consistency, and make targeted transformations using natural language. The model is available via the Gemini API and Google AI Studio, and is priced at $30.00 per 1 million output tokens. Gemini 2.5 Flash Image also includes native world knowledge, allowing it to understand and generate images based on real-world concepts. The model is currently in preview and will be stable in the coming weeks.

Main themes

  • Image Generation
  • AI Models
  • Google Gemini
  • Image Editing
  • AI Ethics
  • Model Availability

What commenters say

  • The model's safety filtering seems inconsistent, as it rejects certain prompts while allowing others that may be considered problematic.
  • The rejection message for certain prompts does not accurately reflect the model's capabilities or limitations.
  • Some users have successfully generated images of humans using the model, despite the claimed safety restrictions.
  • The model's availability and accessibility vary by region, with some users experiencing issues accessing it from certain locations.
  • The model's performance and quality differ when used through different platforms, such as AI Studio and Fal.ai.
  • The distinction between Gemini Flash Image and Imagen models is not clearly understood, with some users seeking clarification on their differences.
  • The model's potential for misuse, such as generating deepfakes or non-consensual imagery, is a concern for some users.
  • The use of language models to generate content, such as the article itself, can result in a perceived lack of quality or authenticity.