Google Lyria 3.5 Moves AI Music From Prompting to Production

Rohit Ramachandran avatarRohit Ramachandran
Lyria 3.5 waveform timeline showing vocals, lyrics, and variable-length music generation in Google Flow Music

Google Lyria 3.5 Moves AI Music From Prompting to Production

Thirty seconds can fake brilliance. Three minutes exposes whether an AI music model understands arrangement, lyrical continuity, vocal character, transitions, and—most revealingly—where a song should end.

That is the useful way to read Google Lyria 3.5, which Google launched on July 29 inside Google Flow Music. Google leads with richer melodies, better lyrics, more expressive vocals, improved pronunciation, and easier control over tempo and duration. Those improvements matter. But the strategically important detail is narrower: Lyria can now target the length a project actually needs, from a short clip to a cohesive three-minute track.

Exact duration turns generated music from an entertaining demo into a production input. A video editor does not merely need “a good electronic song.” They need a 73-second bed that lifts at 0:42 and resolves before the final card. A game team needs a loop that survives repetition. A podcast producer needs an intro that lands on the spoken handoff. An advertiser needs fifteen-, thirty-, and sixty-second variants that still sound related.

Lyria 3.5 is Google trying to own that layer.

There is an equally important caveat. This is a Flow Music product launch, not yet a documented Lyria 3.5 developer-platform launch. Google has not published a 3.5 API model ID, dedicated model card, benchmark table, price, rate limit, latency figure, or audio-format specification. The public developer contract still names the older lyria-3-clip-preview and lyria-3-pro-preview models.

So the honest verdict is two-sided: Lyria 3.5 may be Google’s strongest music model, and Google may have the most formidable distribution stack in AI music. But builders should separate what they can hear in Flow from what they can safely promise in production.

The sixty-second song is the clue

Google’s launch post lists four areas of improvement:

  • Musicality: richer, more complex melodic structures that sound more natural.
  • Lyrics: better quality, prompt adherence, and structural awareness.
  • Vocals: more realistic, emotionally nuanced delivery and better pronunciation.
  • Creative control: easier control over tempo and output duration.

The updated Lyria model page makes the duration claim concrete. It says creators can ask for a quick 60-second clip, a half-length track, or a full three-minute song. It also continues to show text-to-music, image-to-music, user-written or generated lyrics, vocal-style direction, global genres, and detailed acoustic prompting.

“Up to three minutes” is not itself new to Google’s music family. Lyria 3 Pro already generated approximately three-minute songs. What is new in the 3.5 launch is the emphasis on variable, requested duration alongside quality improvements.

That distinction matters. Maximum length answers:

How long can the model keep generating?

Duration control answers:

Can the model compose toward a deadline?

The second question is harder and more useful. A model can fill three minutes with plausible sound yet still miss the requested length, repeat a verse, rush a bridge, drift away from the lyric theme, or end like someone pulled the power cable. Production music needs an arc, not just more audio.

Duration is a composition constraint
The same model has to budget structure differently for every finish line
0:30
Hook or compact cue
1:00
Short-form narrative
1:30
Half-length arrangement
3:00
Full verse–chorus arc

What shipped—and what did not

Google’s announcement is unusually concise. It tells us what changed in the creator experience but leaves most developer and technical questions open. Here is the clean fact ledger as of launch day.

QuestionVerified answerWhat is still missing
What is the model?Lyria 3.5, Google DeepMind’s newest music-generation model.No parameter count, architecture delta, training cutoff, or technical report specific to 3.5.
Where is it available?Rolling out in Google Flow Music on July 29.No confirmed 3.5 rollout to Gemini, Vids, Dream Track, AI Studio, Vertex AI, or the Gemini API.
What improved?Musicality, lyrics, structural awareness, vocals, pronunciation, tempo control, and variable duration.No public side-by-side scores, blind-listening results, error rates, or independent evaluation.
How long are tracks?Up to three minutes, with requested-length examples including 60 seconds and a half-length song.No published tolerance for duration error or guarantee that every requested length is hit exactly.
What inputs work?Text and image prompting are shown; Google documents generated or user-provided lyrics and detailed vocal direction.No 3.5-specific API schema, input limits, upload rules, or native audio-reference contract.
What does it cost?It is available through Flow Music’s credit-based product.No dedicated Lyria 3.5 credit cost, API price, free quota, region list, or tier matrix.
How is it marked?Google says all Lyria tracks carry imperceptible SynthID watermarks.C2PA is documented for older Cloud Lyria 3 endpoints, but not separately confirmed for every Flow Music 3.5 export.

This is not a reason to dismiss the launch. It is a reason to name it correctly: a creator-product rollout with a partially inherited technical disclosure, not a fully specified platform release.

Three dates explain the Lyria name

Lyria 3.5 makes more sense as the third step in a fast product rollout than as a standalone model announcement.

February 18
Lyria 3 beta

Thirty-second songs arrive in Gemini from text, photos, or videos.

March 25
Lyria 3 Pro + APIs

Full songs expand into Gemini API, AI Studio, Vertex AI, Vids, and Flow Music.

July 29
Lyria 3.5 in Flow

Quality, pronunciation, tempo, and requested-duration control improve in the creator product.

The chronology prevents two common mistakes: treating the original 30-second Lyria 3 beta as today’s news, or assuming every surface that received Lyria 3 Pro in March already received 3.5 in July.

One name now hides several different Lyria products

“Lyria” is becoming a family label, and mixing the products together produces bad advice.

Map of the Lyria family separating Flow Music, API preview models, real-time generation, and the open Magenta model

The practical map looks like this:

  • Lyria 3.5 in Flow Music: the July 29 model and creator experience, with the new quality and duration claims.
  • Lyria 3 Pro Preview: the existing developer model for full songs up to roughly three minutes.
  • Lyria 3 Clip Preview: the existing developer model for fixed 30-second clips.
  • Lyria RealTime Experimental: an interactive, continuously steerable instrumental music stream with weighted prompts and controls such as BPM, key, density, brightness, bass, and drums.
  • Magenta RealTime: Google’s open-model experiment for interactive music creation, distinct from the closed Lyria 3.5 product.

The existing API IDs remain:

lyria-3-clip-preview
lyria-3-pro-preview

Google’s current Gemini API music guide documents those two preview models, including supplied or generated lyrics, structured song output, and the formats available on the existing developer surface. It does not document a 3.5 endpoint.

The Gemini API pricing page lists the Clip preview at $0.04 per 30-second song and the Pro preview at $0.08 per full song. Those are useful baseline numbers, but they are not confirmed Lyria 3.5 API prices because no public 3.5 endpoint exists.

There is also a documentation mismatch worth knowing. Google’s March developer announcement describes Lyria 3 output as 48 kHz stereo. The current Google Cloud model specification lists 44.1 kHz, 192 kbps MP3 and a 184-second Pro maximum. These may reflect different serving surfaces or updated implementations, but Google does not reconcile them. If output format matters to your pipeline, inspect the generated file instead of repeating a marketing number.

Flow Music may be the real product

If audio quality keeps improving across the market, one-shot generation becomes less defensible. The valuable product is the system that gets from idea to an accepted, editable, publishable track with the fewest dead ends.

That is why the Flow Music launch surface matters.

Flow Music’s public product already advertises a conversational Producer, full-length song creation, projects, remixing, audio effects, stem separation, downloads, publishing, playlists, social discovery, taste personalization, and music videos powered by Google’s video models. Some of those are product-layer tools rather than Lyria 3.5 model outputs. That is precisely the point: users buy the workflow, not the checkpoint.

The moat Google is assembling looks like:

idea or image
    ↓
Lyria generation
    ↓
local edits + lyrics + stems
    ↓
Veo / Gemini visual creation
    ↓
YouTube / Shorts / Vids distribution
    ↓
feedback, taste, provenance, moderation

Suno and Udio proved that people want to generate songs. Google’s opportunity is larger: make music a native supply layer for every video, presentation, ad, short, game prototype, creator channel, and AI-generated experience inside its ecosystem. That fits a pattern RohitAI has also tracked in Google’s specialized audio strategy: the model becomes more valuable when it is attached to a distribution and workflow surface.

That is a better business than selling isolated songs.

It also explains the product-first launch. Flow can collect the telemetry Google needs before exposing a broader API:

  • Which requested durations fail?
  • Where do people regenerate?
  • Which sections get edited?
  • Which languages produce pronunciation complaints?
  • Which prompts trigger artist or lyric filters?
  • How often are stems used?
  • What percentage of tracks are downloaded, paired with video, or published?

An API tells Google whether a request returned 200. A creative environment tells Google whether the result survived contact with taste.

The global sample reel is a distribution signal

Google’s eleven official 3.5 samples are not a random playlist. They include Spanish dance pop and reggaeton, French alternative pop, Brazilian funk carioca, Chilean neoperreo, Hindi singer-songwriter material, Afro-pop and Caribbean jazz fusion, as well as several English-language electronic, funk, R&B, and glitch styles.

Pair that with Google’s explicit claim of improved pronunciation and the strategy becomes clearer.

The first wave of AI music products was visibly English-first. Yet YouTube’s creator economy is global, and localized music is not just translation. Rhythm, vocal phrasing, instrumentation, accent, genre history, and lyrical idiom all matter. A model that pronounces words correctly but flattens every region into a familiar American pop arrangement has not solved localization.

Lyria’s multilingual push could support:

  • Localized ad variants with the same campaign structure.
  • Region-specific versions of a creator’s intro or theme.
  • Video soundtracks that reflect the visual setting rather than defaulting to generic cinematic music.
  • Educational songs and mnemonic content in local languages.
  • Draft demos for multilingual songwriters.
  • Cross-market testing before a team commissions final human production.

My prediction is that music localization arrives before a broad 3.5 API. It fits Google’s strongest assets—YouTube distribution, translation, Gemini, Ads, and regional creator networks—and it generates the kind of product feedback Flow can capture.

Where Lyria 3.5 sits against Suno, ElevenLabs, and open models

There is no honest universal leaderboard for generated music. Taste is subjective, product features differ, launch samples are curated, and the same prompt rarely maps cleanly across systems. The useful comparison is not “which makes the best song?” but “which problem has each company chosen to solve?”

Google Lyria 3.5
Best strategic fit: cross-media production

Strongest story for exact-length soundtracks, image-led creation, multilingual creator workflows, provenance, and eventual integration across Flow, Gemini, Vids, YouTube, and Cloud. The 3.5 developer contract is still missing.

Suno v5.5
Best strategic fit: personal identity

Voices, Custom Models, and My Taste, plus a mature song community and generative studio, make Suno stronger when the objective is a persistent sound that feels recognizably yours.

ElevenLabs Music v2
Best strategic fit: rights-conscious production

ElevenLabs Music v2 emphasizes licensed-only training, commercial clearance, reference tracks, inpainting, composition plans, private fine-tunes, and developer or brand workflows.

Open music models
Best strategic fit: control and deployment

Open weights and adaptation matter when you need local inference, LoRA-style specialization, inspectable infrastructure, or freedom from a single hosted product’s policy and lifecycle.

The gap in Google’s current offer is persistent sonic identity. Lyria 3.5 may make polished, expressive music, but Google has not announced a private voice model, catalog-trained creator model, or brand-specific sound model comparable to the identity features rivals are building.

That gap will become more important as raw fidelity improves. When everyone can make a competent pop song, “sounds good” stops being scarce. What remains scarce is:

sounds like me
sounds like this brand
fits this scene
can be edited precisely
has a defensible rights trail
reaches an audience

Identity, editability, rights, and distribution are the next benchmark suite.

What the model card does—and does not—tell us

Google’s existing Lyria 3 model card says it covers Lyria 3 “and subsequent versions,” so it is the best technical disclosure available for 3.5. It describes:

  • A latent-diffusion system operating on temporal audio latents.
  • Audio training data annotated with text at multiple levels of detail.
  • Deduplication, safety filtering, and quality filtering.
  • Training on Google TPUs using JAX and ML Pathways.
  • Human and automated evaluations for musical quality, aesthetics, vocals, audio fidelity, and prompt adherence.
  • Supervised fine-tuning, reinforcement learning from human and critic feedback, red teaming, product filtering, and SynthID.

The card says Lyria 3 significantly improved audio fidelity and lyric prompt adherence over Lyria 2. It publishes no numerical results. It also does not disclose a 3.5 architectural change, parameter count, corpus size, named datasets, training cutoff, language distribution, generation speed, benchmark prompts, or comparative win rate.

That leaves a major evidence gap. “Significant advancements” is a vendor claim until Google publishes a repeatable evaluation or independent listeners test the same prompts across models.

A serious music-model evaluation should measure at least five layers:

  1. Constraint accuracy: requested duration, BPM, key, instrumentation, language, vocal register, and supplied lyrics.
  2. Long-range structure: whether verses, choruses, bridges, builds, transitions, and endings feel intentional.
  3. Audio integrity: vocal artifacts, transient smearing, phase issues, clipping, stem bleed, and compression quality.
  4. Editability: whether a local change stays local or destroys everything that already worked.
  5. Outcome efficiency: retries, time to accepted track, human cleanup, export friction, and effective cost per used minute.

Without those measurements, we are comparing demos.

SynthID is provenance, not permission

Google says every Lyria track includes an imperceptible SynthID watermark. SynthID is designed to remain detectable after common transformations, and Gemini can inspect uploaded audio for Google’s watermark. Google also warns that a negative result does not prove a file is human-made and that heavy editing can weaken detection.

For the existing Cloud Lyria 3 endpoints, Google documents C2PA Content Credentials as well. C2PA can record origin and editing history in a signed manifest. It does not certify that a song is true, original, non-infringing, commercially cleared, or copyrightable.

Those distinctions are easy to blur:

SynthID -> “Google AI likely touched this audio”
C2PA -> “Here is a signed provenance record”
copyright -> “A legal right may exist in human-authored expression”
commercial clearance -> “Your use fits the contract and rights chain”
indemnity -> “Someone contractually agrees to cover defined claims”

They are not interchangeable.

Google says Lyria is intended for original expression rather than direct artist imitation. Named creators are treated as broad stylistic inspiration, and the system uses output-recitation and vocal-likeness filters. Those mitigations are valuable, but Google itself does not claim they are foolproof.

The training-data question remains contentious. Google says Lyria was trained using material YouTube and Google have a right to use under terms of service, partner agreements, and applicable law. That language appears in Google’s Lyria 3 Pro launch explanation, and it is broader than “every recording was affirmatively licensed for AI training.”

In March, independent musicians filed a proposed class action, Kogon v. Google, alleging unauthorized use of music from YouTube to train Lyria. Google moved to dismiss in June, arguing among other things that uploaders granted a broad license through YouTube’s terms. The complaint and dismissal arguments are summarized here. Those are the plaintiffs’ allegations and Google’s defense; the dispute is unresolved.

The case matters beyond one model. If Google’s interpretation of platform terms prevails, companies that host enormous user-media catalogs gain a structural training-data advantage over standalone AI labs. If it fails, the legal foundation beneath that advantage becomes much more expensive.

Creators and brands should keep their own provenance package:

  • Original lyrics and dated drafts.
  • The exact prompt and source images.
  • Generation timestamps and model or product surface.
  • Downloaded C2PA manifests where available.
  • Stems and edit-session history.
  • Human arrangement, performance, selection, and modification notes.
  • Clearance records for uploaded material and reference audio.

That record does not answer every copyright question, but it documents human contribution far better than a final MP3 alone.

Who should use Lyria 3.5 now?

Lyria 3.5 looks immediately useful for:

  • YouTube and short-form creators who need music matched to an edit.
  • Marketing teams producing many localized or duration-specific variants.
  • Product teams prototyping sonic branding, onboarding, and demo experiences.
  • Game developers creating temporary cues, mood studies, and adaptive-content prototypes.
  • Podcasters producing intros, transitions, and beds.
  • Songwriters testing arrangements, genres, or vocal treatments before a studio session.

I would wait—or run a serious comparison—if you need:

  • A fixed personal singing voice.
  • A private brand or artist model trained on an authorized catalog.
  • Long-form songs beyond three minutes.
  • Open weights, local inference, or on-device generation.
  • A stable Lyria 3.5 API, SLA, lifecycle commitment, or deterministic integration.
  • A fully enumerated licensed-training catalog or explicit enterprise indemnity.
  • Precise inpainting and edit locality proven on your own material.

The word prototype matters. Generated music can move very quickly from sketch to customer-facing asset, but the governance burden changes the moment it ships.

The 20-prompt Lyria 3.5 evaluation I would run
01Use the same prompts across Lyria, Suno, ElevenLabs, and your current music source
02Split tests between mainstream genres and the niche genres your product actually needs
03Measure requested-versus-actual duration and BPM instead of judging by ear alone
04Test supplied lyrics separately from model-generated lyrics
05Use multilingual native speakers to score pronunciation, phrasing, and cultural fit
06Check verse, chorus, bridge, transition, and ending coherence over the full track
07Repeat each prompt to measure variance and generations required per acceptable result
08Test whether section edits stay local and whether stems contain audible bleed
09Record moderation blocks and false positives near artist, lyric, and vocal constraints
10Inspect exported sample rate, bitrate, format, metadata, SynthID, and C2PA behavior
11Calculate cost and human edit time per accepted minute of music
12Have legal or rights owners approve the workflow before customer-facing distribution

Three predictions from this launch

1. Lyria 3.5 reaches APIs after Flow supplies the failure data

Google followed a product-to-platform sequence with Lyria 3: consumer launch in February, then Pro, API, AI Studio, Vertex, Vids, and broader surfaces in March. I expect 3.5 to expand too, but the Flow-only launch suggests Google wants creator telemetry first. The staged rollout also resembles the broader control-plane shift discussed in RohitAI’s analysis of Google’s managed model contracts.

When a developer endpoint appears, watch for more than a model ID. The meaningful contract will include duration tolerances, formats, quotas, moderation behavior, rights terms, lifecycle status, and whether editing or stems are model-native or separate services.

2. AI music becomes machine-readable media infrastructure

The existing Lyria developer flow can return audio alongside lyrics and song structure. That second output is quietly powerful. Once a song has machine-readable sections and text, another system can align captions, build karaoke, cut visuals on boundaries, translate lyrics, create beat-aware edits, or generate multiple campaign lengths.

Music stops being an opaque file and becomes a structured object:

audio
+ lyrics
+ section boundaries
+ tempo
+ provenance
+ edit history
= programmable soundtrack

That is a larger platform opportunity than song generation alone.

3. Commodity music gets cheaper; recognizable identity gets more valuable

At the existing Lyria 3 API prices—four cents for a clip and eight cents for a full song—the marginal generation cost is already collapsing. Flow’s subscription tiers push the perceived price lower still. As with real-time voice agents, the right unit is not raw inference cost but the cost of an accepted audio outcome.

The first market under pressure is not great artists. It is interchangeable production music: generic corporate beds, podcast transitions, rough demo cues, stock jingles, and low-stakes social soundtracks.

As supply explodes, value moves to taste, identity, trustworthy rights, cultural specificity, editing craft, fan relationships, and distribution. The model makes sound abundant. It does not make attention abundant.

The practical verdict

Lyria 3.5 is a big release, but not because Google added another decimal point to a music model.

It matters because Google is joining model quality to a production surface and a distribution empire. Requested duration, better pronunciation, Flow’s editing loop, Veo-powered visuals, YouTube reach, SynthID, and Cloud infrastructure can become one vertically integrated media pipeline.

That is the advantage competitors should fear.

The missing pieces are equally important. There is no public 3.5 developer contract, no quantitative evaluation, no version-specific technical report, no dedicated API price, and no complete training-data accounting. Lyria 3.5 may sound better; Google has not yet made it measurable or programmable.

So use it now if Flow Music fits your creative workflow. Test it aggressively if music is going into customer-facing work. Keep your prompts, stems, edits, and provenance records. And if you are a developer, do not build against a model name that does not yet exist in the docs.

The music-model race is no longer about who can make the most impressive first thirty seconds.

It is about who can turn generated sound into a finished, editable, rights-aware, distributed product.

Lyria 3.5 is Google’s strongest attempt yet.

Frequently asked questions

Is Lyria 3.5 the same as Lyria 3 Pro?

No. Lyria 3 Pro is the existing preview model for full songs and developer access. Lyria 3.5 is the July 29 upgrade launched in Flow Music with claimed improvements to musicality, lyrics, vocals, pronunciation, tempo, and duration control. Google has not published a 3.5 API model ID.

Where can I use Lyria 3.5?

Google explicitly announced Lyria 3.5 in Google Flow Music. The company has not yet confirmed that Gemini, Vids, Dream Track, AI Studio, Vertex AI, or the Gemini API are serving 3.5.

How long can Lyria 3.5 songs be?

Google says up to three minutes and gives examples of requesting a 60-second clip, a half-length track, or a full three-minute song. It has not published a formal duration-error tolerance.

Can Lyria 3.5 use images and custom lyrics?

The updated Lyria page and official prompt guide document image-to-music, generated lyrics, user-provided lyrics prefixed with Lyrics:, vocal profiles, instruments, tempo, dynamics, genre blending, language, and backing-vocal instructions.

Is there a Lyria 3.5 API?

Not publicly documented at launch. The current developer models remain lyria-3-clip-preview and lyria-3-pro-preview.

What does Lyria cost?

For the older developer previews, Google lists $0.04 per 30-second Clip generation and $0.08 per Pro full song. Flow Music has free and paid credit tiers. Neither is a dedicated Lyria 3.5 API price.

Does Lyria 3.5 generate stems?

Flow Music offers stem tools, but Google has not described stem generation as a native Lyria 3.5 model output. Product-layer editing and model capability should not be treated as the same thing.

Does SynthID make a track safe for commercial use?

No. SynthID identifies likely Google AI provenance. Commercial use depends on the applicable product terms, your inputs, the output, local law, platform policy, and any third-party rights. SynthID is not clearance or indemnity.

Is Lyria 3.5 trained only on licensed music?

Google has not made that narrow claim. It says it used material YouTube and Google have a right to use under terms, partner agreements, and applicable law. The company has not published a catalog-level licensing breakdown.

What should developers watch next?

A 3.5 model ID, API and Vertex availability, pricing, rate limits, audio formats, latency, lifecycle status, regional access, quantitative evaluations, updated model-card details, and explicit migration guidance.