Complete AI Training

AI news ·

Suno introduces Speech, an audio model that generates voice and music together

Suno's new Speech feature generates spoken-word audio and background music together from text, now in public beta. Early testers saw unpredictable accents and pauses, but the tool lets marketers, writers, and event planners produce scored audio without a studio.

Share

Suno, the AI music generation company, is rolling out a new feature called Speech in public beta today. It combines spoken-word audio with original background music in a single track, letting users turn text-poems, stories, pep talks, or any written idea-into a produced audio piece. For marketers drafting ad voiceovers, event planners creating themed invitations, or writers hearing their work set to music, it opens a direct line from text to finished audio without a recording studio or production team.

What Speech does

Speech is built directly into Suno's existing platform. A user types or pastes text, then describes the voice and musical style they want. The model generates both elements together as one cohesive track. The company describes it as "the first audio model that generates voice and music together as one cohesive track."

Suno has always centered on music creation, but the team sees Speech as a natural extension of how people already use the tool. "Every day, people make songs for birthdays, weddings, inside jokes, faith and worship, their kids, their friends, and moments that might mean absolutely nothing to anyone else but mean everything to them," the company said. Speech adds a spoken-word layer to that same personal, often private creativity.

Beta quirks and what the company learned

The beta label is not a formality. Early testers encountered unpredictable results. "Occasionally, British accents can wander off to Australia and back. Dramatic pauses may be very dramatic," Suno said. The company spent the past month testing with a small group and is now opening access to all users.

During development, the team found themselves using Speech in ways that ranged from absurd to unexpectedly affecting. They turned friends' text messages into over-the-top dramatic readings. They scored ordinary voice notes with epic backing tracks. They made meditations, bedtime stories, and personal pep talks. "That creativity is a natural extension of how people already use Suno," the company said.

Why this matters for creatives, marketers, and writers

For professionals who need audio content but lack production resources, Speech lowers the barrier to creating finished pieces. A copywriter can hear how a tagline lands when spoken. An event producer can generate themed audio invitations without hiring voice talent. A sales team can turn a pitch script into a scored walkthrough for a client. The output may not replace a professional studio session, but it provides a fast way to prototype, test, or deliver personal audio where polish matters less than speed and sentiment.

Suno frames this as part of a broader shift toward "creative entertainment"-the idea that making something, however small, carries its own fulfillment. AI Music Production Courses can help vocal artists and songwriters build on tools like this, while Text to Speech AI Courses cover the wider landscape of AI-generated voice. Speech is available now in beta to all Suno users.

Share