
Suno Bark: Suno's transformer-based text-to-audio model (MIT, 36k stars)
Suno's transformer-based text-to-audio model for generating speech, music, sound effects, and non-verbal vocalizations from text prompts.
Explore our latest articles and technical deep dives.

Suno's transformer-based text-to-audio model for generating speech, music, sound effects, and non-verbal vocalizations from text prompts.

A deep learning toolkit for text-to-speech supporting 1100+ languages with voice cloning, fine-tuning, and real-time inference.

Supporting 8 languages with voice cloning, emotion control, and 4x faster than real-time inference — a state-of-the-art TTS model with permissive licensing.

A few-shot voice cloning and TTS system achieving natural speech synthesis with just 1 minute of reference audio.

Retrieval-based Voice Conversion for real-time voice conversion that preserves intonation and emotion while changing the speaker identity.

A style-based TTS model achieving human-level naturalness with expressive speech synthesis and zero-shot voice cloning.

A token-based neural codec language model achieving state-of-the-art zero-shot TTS and voice editing with edit capability.

Transcribing and translating speech with 96.8% word accuracy across 99 languages — OpenAI's state-of-the-art speech recognition system.

A motion module for Stable Diffusion that turns any SD checkpoint into an animation generator — no specialized video models needed.