news.volyx.in

VibeVoice: Open-source frontier voice AI (github.com)

386 points by tosh · 124 days ago · 181 comments on HN

Article summary

VibeVoice is an open-source frontier voice AI model developed by Microsoft, which includes both Text-to-Speech (TTS) and Automatic Speech Recognition (ASR) models. The model can handle long-form audio and supports multiple speakers, languages, and streaming text input. However, the original TTS model was removed from the repository due to security and safety concerns. The current models available include VibeVoice-ASR, VibeVoice-TTS, and VibeVoice-Streaming.

Main themes

  • Voice AI
  • Open-source models
  • Text-to-Speech
  • Speech Recognition
  • Long-form audio processing
  • Multilingual support

What commenters say

  • The VibeVoice model has limitations, such as being slow and heavy, and may not be suitable for commercial or real-world applications.
  • Some users have expressed disappointment with the model's performance, particularly with the 0.5B real-time model, which can add music randomly and struggle with special characters.
  • The model's ability to handle long-form audio and support multiple speakers is impressive, but its accuracy and reliability are still questionable.
  • There are concerns about the potential misuse of the model for creating deepfakes and disinformation, and users are advised to use it responsibly.
  • The removal of the original TTS model from the repository has raised questions about the model's security and safety, and some users are skeptical about Microsoft's intentions.
  • Some users have noted that the model's performance is not significantly better than other existing models, such as Parakeet and Whisper, and that it may not be worth the hype.
  • The use of stars on GitHub can be misleading, as they may not necessarily indicate the quality or trustworthiness of a repository, but rather serve as bookmarks or indicators of popularity.
  • The model's training data and algorithms may be flawed, leading to inconsistent and unreliable results, and some users have reported better experiences with other models, such as Chatterbox Turbo and Qwen TTS.