Speech Beta Adds Natural Voice Generation
Suno expanded its platform with a dedicated speech generation feature currently in beta. This update allows creators to generate expressive voiceovers and dialogue directly within the tool.
Suno, the AI audio generation platform, launched a speech generation feature in beta to expand its capabilities beyond music production. The update allows users to generate natural-sounding dialogue and narration with specific control over tone and emotion. This shift positions Suno as a more comprehensive audio suite for creators who need both soundtracks and vocal performances.
What's new
Suno Speech (beta) introduces a dedicated model designed specifically for the nuances of human talk rather than singing. While the platform previously allowed for some vocal manipulation through its song-generation engine, this update provides a focused interface for text-to-speech tasks. The beta release emphasizes expressive performance, aiming to capture the subtle inflections, pauses, and emotional shifts found in natural conversation.
As of October 2024, the feature is rolling out to Suno Pro and Premier subscribers. Users can input text and select from various vocal styles to generate clips that function as standalone narration or dialogue. The model handles different accents and pacing, providing a toolset for creators who require high-fidelity vocal tracks without the melodic structure of a song.
How it fits your workflow
Suno Speech (beta) serves as a direct alternative to ElevenLabs or OpenAI’s voice tools for creators already using the Suno ecosystem. For filmmakers and video editors, this feature simplifies the process of creating scratch vocals or temp tracks. Instead of jumping between a music generator for the score and a separate text-to-speech platform for the narration, editors can manage all audio assets within the Suno interface.
The expressive nature of the Suno Speech model makes it particularly useful for character dialogue in animation or video game prototyping. By providing more than just a flat, robotic read, it helps directors visualize the pacing of a scene before hiring voice actors. In the context of social media content, it allows for quick turnaround on narrated videos, competing with the built-in voice tools found in apps like TikTok or CapCut, but with a higher degree of tonal control. This development suggests Suno is moving toward a unified audio workstation model, where music, sound effects, and speech are generated in a single environment.
What it costs / how to try it
Suno Speech is currently available in beta for Pro and Premier members. Subscribers can access the feature through the main Suno web interface. There is no word yet on when the feature will move out of beta or if it will be available to free-tier users in the future.
Read the original announcement on Suno ↗