Suno Now Speaks, and It Brings Its Own Soundtrack

On 1 October, Suno opened Speech in public beta for every user on web, iOS, and Android. You type an idea, a poem, or a script you already wrote, then describe the voice and the musical mood you want underneath it. Speech returns the two as a single track.
Suno's chief product officer, Jack Brody, described it in the announcement as the first audio model that generates voice and music together as one cohesive track. That framing is the whole product. Plenty of tools produce a synthetic voice, and plenty produce music. The pitch here is that rhythm and mood are matched to the meaning while both are being generated, rather than bolted together afterward in an editor, which is where every amateur podcast intro goes to die.
The mechanics are deliberately simple. Simple mode takes a natural-language prompt, something like "a sea captain rallying his crew before a storm," and generates both the speech and the score. Advanced mode takes an exact script and exposes controls for voice gender, vocal style, pacing, and how much variety each generation should have. There is a toggle to drop the music entirely if you only want clean narration. The beta caps a generation at roughly eight minutes.
What beta means here
Suno is unusually candid about the state of the thing. Brody's note says a British accent can occasionally wander off to Australia and back, and that dramatic pauses may be very dramatic. That kind of disclosure is rarer than it should be in launch posts, and it tells you something about where the voice modelling still is.
The company also named no supported languages, which means performance outside English is unverified until you try it. Given that Suno's existing audience is global, the absence of a language list in a speech launch is a signal rather than an oversight. Voice cloning is hard in English and harder in languages with richer tonal systems, and Suno is shipping before it can promise much.
Nothing was said about pricing tiers, generation limits, or commercial rights for the audio either. Those live in the plan details, and anyone planning to use Speech for client work should read them carefully, given Suno's history.
The licensing turn
Suno's last year has been a slow pivot from scraping to signing. On 9 September it released v6, a generation of music models built with Warner Music Group, BMG, and Believe. The family has three parts: the flagship v6 and an exploration-oriented v6-wild for Pro and Premier subscribers, and v6-mini for everyone. Suno says v6 can edit part of an existing song using plain language, build a mashup from several sources in one request, isolate audio to build a new beat, create music from text, audio, images, and video, and rewrite a single lyric without rebuilding the whole song. As v6 rolls out, the company plans to retire its older models and move the platform entirely onto the new generation.
That is a real capability upgrade, and it is also a legal strategy. The licensing deals turn Suno from a defendant into a partner, and the opt-in artist experiences the company says it is building, where artists choose to participate and get paid, are the shape of a business that wants to stop being sued.
The lawsuits have not stopped. Sony Music and Universal Music Group sued Suno for copyright infringement, and in September 2026 the labels expanded the case with an accusation of model laundering: using older, allegedly infringing models to generate synthetic audio for training newer ones. That is a sharper claim than training on scraped songs, because it describes a chain of infringement that continues after a company says it has cleaned up.
Meanwhile, CEO Mikey Shulman told Bloomberg the same day as the Speech launch that Suno is well past the 2 million subscribers and $300 million in revenue it last reported. A company that big expanding into speech is doing something more than adding a feature. It is building a portfolio of generative audio tools under one subscription so that a customer has less reason to leave.
Who this compresses work for
The right question for a new tool is not what it does but whose work it replaces. Judging by the uses Suno reported from its month of private testing, four groups feel it immediately.
Short-form video creators have the most obvious fit. A musical signature intro, transitions between segments, a closing sting: assets that used to require a composer and a session now come out of one generation. The compression is real, because the old pipeline had a hard step where a human matched a track to a script, and that step is now inside the model.
Podcasters get a version too. A narrated segment with original music underneath is a two-part problem that just became one, and the toggle for music off means a clean read is available when a client needs one.
The third group is less obvious and probably larger. Suno's reported test cases lean hard on personal and semi-personal content: dramatic readings of friends' text messages, voice notes with epic scores, guided meditations, pep talks, and bedtime stories for the team's own children. That is the same pattern Suno's song generator followed. It became a fixture of birthdays and inside jokes before it became a production tool, and the company is positioning Speech in the same lane, calling the category creative entertainment.
Competitors are already in the room. ElevenLabs has dominated voice synthesis since 2023, Adobe has a text-to-speech tool, and Google DeepMind has been working on speech for a decade. Suno's edge is not voice quality, which the beta caveats suggest is behind the leaders. It is the pairing: one prompt, one take, spoken words and original music that already fit each other.
The unresolved part
The gap between Suno's product story and its legal story is where this gets interesting. Speech broadens what the platform can make, and it also broadens what the platform can be accused of making. A voice tool that produces narration in a user's chosen style sits close to impersonation, and a model that generates music alongside speech inherits every question about training data the music side is already fighting over.
Suno's answer so far is deals and disclosure. It has signed the labels, added screening for uploaded audio and lyrics, and promised artist opt-ins with payment. That is more than most generative audio companies have done. It is also not a settled answer, because the model laundering claim, if it holds up, describes a problem inside Suno's own training history rather than at its edges.
What Speech does well is remove a step that used to require two specialists and a session booking. What it does not do is resolve whether the models behind it were built the way the company now says it builds them. Those two questions will keep running in parallel, and only one of them will be answered by using the product.
Related articles
Agility Digit 5 Ships With a Safety Case, Not Just a Spec Sheet
A warehouse floor is not a lab. Certification is the gate, not the demo.
Frontier Agents Finished 30 Percent of a Research Workflow. That Is the Number.
Agents can run research. Inventing the procedure is still out of reach.
Figure AI Locked In $3.5 Billion of Compute Before It Has a Product to Sell
The bet is that generalisation is a compute problem. The field has not settled that.
OpenAI Finally Put Transparent Backgrounds in the Image API
A small feature that deletes a whole step from the pipeline.