Radar · 03/08/2026 · happened on 29/07/2026 · models

Lyria 3.5 in Google Flow Music: music generation becomes a creative tool

Google DeepMind has released Lyria 3.5 inside Google Flow Music. The music generation model improves across four dimensions: richer and more natural melodic structures, lyrics with better adherence to prompts and awareness of structure (chorus, verse, bridge), more expressive voices with improved pronunciation, and controls over tempo and duration of the output.

For anyone working in audio, the signal is clear. Music generation is moving out of the phase where it produced «something that sounds like music» and becoming a tool with creative controls. You can set the tempo, define duration, write lyrics that the model respects structurally. The vocality is emotionally expressive, a step ahead of the flat text-to-speech of a few years ago.

The distinction that matters is between functional generation (reading text aloud) and creative generation (producing a song with structure, emotion, and control). Lyria 3.5 aims at the latter. For those building content, jingles, music demos, or audio prototypes, this means less time fighting tools that ignore your intentions and more time iterating on results.

The limitation is in the data: the announcement is sparse on numbers. No comparable benchmarks, no verifiable samples, no documented API. It’s a closed product within Flow Music, accessible as a creative tool, not as a downloadable model.

If you want to try it: open Flow Music and generate a track with your own lyrics, setting tempo and duration. Verify how much your choices actually come through in the result.

In detail

Context helps explain why this release matters, even though technical details are limited.

What came before. AI music generation so far has been divided into two separate fields. On one side, text-to-speech models (like those from ElevenLabs or voice systems from OpenAI and Anthropic), which read text in a believable voice but without musical intent. On the other, actual music generation models (Meta’s MusicGen, Suno, Udio), which produce complete tracks but with limited control: you ask for a genre, get something, and then hope the bridge lands where you wanted it.

Lyria 3.5 seeks to bridge that gap. The novelty is it generates music with granular controls: tempo (BPM), output duration, text structure (the model recognizes the difference between a chorus and a verse and treats them accordingly). Vocality improves on two axes: emotional expression and pronunciation. If you’ve tried having a model from a year ago sing lyrics, you know pronunciation was the most annoying weak point.

What we don’t know. The DeepMind blog announcement is brief: four bullet points, no numbers, no reproducible audio samples, no comparison with previous or competing models. We don’t know the maximum track length, or how well the controls actually work on complex genres (polyphony, tempo changes, extended instrumental sections). There’s no documented API: the model lives inside Flow Music, a consumer product, not a developer platform.

Practical implications. For those building audio content today, Lyria 3.5 is a tool to try on work where prototype quality matters more than perfect reproducibility: jingles for campaigns, demos for music pitches, background layers for video. It’s not yet a final production tool, and the absence of an API means you can’t integrate it into an automated workflow. For those hoping creative voice generation would become accessible, this signals the direction is right, but the path remains open.

Stated limitations. Google provides no information on output usage licenses or the origin of training data. For anyone producing commercially, the lack of copyright clarity is a practical blocker.

Type to search across course, playbooks, skills, papers…