news.volyx.in

Riffusion – Stable Diffusion fine-tuned to generate music (riffusion.com)

2421 points by MitPitt · 1361 days ago · 465 comments on HN

Article summary

Riffusion is a fine-tuned version of Stable Diffusion, a model that generates music based on text prompts. The model works by generating spectrograms, which are then converted into audio. The results are surprisingly good, with some examples sounding like real music. The model can be used to create full-length songs with rich musicality and dynamic vocals.

Main themes

  • AI-generated music
  • Stable Diffusion
  • Spectrograms
  • Music production
  • Copyright implications
  • Innovation in music creation
  • Technical details of the model
  • Potential applications of the model

What commenters say

  • The model's ability to generate music from text prompts is impressive and has the potential to revolutionize music production.
  • The use of spectrograms as an intermediate representation is a key factor in the model's success.
  • Some commenters are concerned about the potential copyright implications of generating music based on existing songs.
  • Others are excited about the possibilities of using the model to create new and innovative music.
  • The model's limitations, such as its inability to handle vocals as well as text-to-speech models, are also discussed.
  • There is debate about whether the model's output is truly 'music' or just a clever imitation.
  • Some commenters are interested in exploring the technical details of the model and how it can be improved.
  • The potential applications of the model, such as generating music for videos or video games, are also discussed.