Lyria 3.5 is Google DeepMind's flagship music model, accepting a 131,072-token input and generating full songs up to three minutes long with tuned vocals in 8 languages. It launched in Flow Music on July 29, 2026 and reached the Gemini API and Gemini app on September 4, 2026.
Lyria 3.5 is Google DeepMind's music generation model, producing full-length structured songs up to three minutes long from text or up to 10 images. It launched in Flow Music on July 29, 2026, and reached the Gemini API on September 4, 2026, priced at $0.08 per generated song.
Provider: Google DeepMind · Family: Lyria 3
More about Google DeepMind on HokAI
Context window: 131,072 tokens
Input modalities: text, image · Output: audio, text
About Lyria 3.5
Lyria 3.5 is Google DeepMind's text-and-image-to-music model, released first in Google Flow Music on July 29, 2026, then extended to the Gemini app, the Gemini API in public preview, Google AI Studio and Google Vids on September 4, 2026. It uses a latent diffusion architecture on temporal audio latents, trained on Google's TPUs with JAX and ML Pathways, and sits as the flagship of DeepMind's Lyria line above the shorter Lyria 3 Clip model and 2025's instrumental-only Lyria 2. As a music model rather than a language model, Lyria 3.5 does not publish scores on reasoning benchmarks. Google's announcement states vocals, lyric adherence and melodic complexity all improved over Lyria 3. On the Artificial Analysis Music Arena, which ranks models by blind listener votes, predecessor Lyria 3 Pro placed 4th on the instrumental leaderboard at an Elo of 1119, behind Suno V5.5 (1190), Mureka V8 (1167) and Suno V5 (1159); Lyria 3.5 had not been separately arena-tested as of September 2026. The Gemini API lists a 131,072-token input limit for lyria-3.5, covering the text prompt, lyric tags and any image data. DeepMind's product page describes a full song arc rather than a single loop: an opening builds into verses, a hook-driven chorus, a bridge, and a close, running up to three minutes. Input is text and up to 10 images per request; output is 44.1 kHz stereo audio, MP3 by default with WAV optional, plus the generated lyric text. Users supply their own lyrics marked by section with timing cues like 0:00 to 0:10, or let the model draft lyrics from a one-line theme. The API's unsupported-feature list rules out function calling, structured outputs, caching, batch and the Live API here; generation is single-turn, so no follow-up call edits one section of an already-produced track. Lyria 3.5 is billed at a flat $0.08 per generated song through the Gemini API, the same rate as predecessor Lyria 3 Pro, with no free API tier; a short jingle costs the same as a full track since price is per request, not per second. Use inside Google's own apps (Flow Music, Vids, the Gemini app) is bundled in, not billed separately. The model is reachable through the Gemini API's interactions endpoint and Google AI Studio with an API key. As of September 2026 it had not appeared in Vertex AI's Model Garden, where Lyria 2, Lyria 3 and Lyria 3 Pro are already listed; enterprise Cloud users work with one of those instead. DeepMind builds safety in at pretraining, post-training and the product layer: dataset filtering, supervised fine-tuning, RLHF, and a safety filter at inference. Every prompt passes that filter, which refuses to imitate a specific singer or reproduce existing song lyrics. Every track carries an imperceptible SynthID watermark. Lyria 3.5 fits teams needing a finished, royalty-clear backing track, jingle or campaign soundtrack quickly, or developers who'd rather pay per song than run a subscription. It is weaker for anyone needing to edit one section of a track, clone an artist's voice, hum a melody, or work from an existing stem. Suno and Udio, which offer stem export and multi-turn editing, are named most often as substitutes. DeepMind's model card says Lyria 3.5 trains on audio paired with variable-detail text captions, filtered for safety and quality, and that training data is not linked back to individual recordings, catalogs or rights holders. Google says the material comes from sources it has rights to via YouTube's terms, partner agreements and applicable law. The paid Gemini API does not train on submitted prompts or audio; logs are kept briefly for abuse review, with zero-data-retention available for enterprise projects. Lyria 3.5 replaces Lyria 3 Pro (released March 25, 2026 alongside the shorter Lyria 3 Clip), adding richer melodic structure, better lyric adherence, more expressive vocals, and direct tempo and duration control. Lyria 2, GA on Vertex AI since May 2025, remains the instrumental-only predecessor two generations back.
Pricing
Billed per generated song rather than per token, regardless of length up to three minutes or how many images were sent as a prompt. No free API tier exists. Lyria 3 Pro carries the identical rate; the shorter Lyria 3 Clip runs half that for a fixed 30-second clip. Access through the consumer products (Flow Music, the Gemini app, Vids) carries no separate per-song charge.
Key Features
- Full-Length Song Generation: Builds a complete song arrangement, opening through close, up to three minutes long, rather than the fixed 30-second loop the sibling Lyria 3 Clip model returns.
- Structural Lyric Tags: Follows section labels (opening, hook, middle-eight) and timestamp cues in a supplied lyric sheet, or writes original lyrics on request from a one-line theme.
- Image-to-Music Prompting: Takes up to 10 images alongside a text prompt so a visual mood board can steer the composition, not just adjectives in a text prompt.
- 44.1 kHz Stereo Output: Renders MP3 by default with WAV available, matching streaming-quality audio rather than a low-bitrate preview.
- 8-Language Vocal Tuning: Pronunciation and vocal style are tuned for English, German, Spanish, French, Hindi, Japanese, Korean and Portuguese.
- SynthID Watermarking: Every generated track carries an imperceptible SynthID audio watermark so platforms can detect AI-made music.
Pros
- Generates a complete, structured 3-minute song for one flat per-request price through the Gemini API, with no separate charge for vocals, lyrics or the images used as a prompt.
- Supports section-tagged lyrics with timestamps, giving more compositional control over song structure than a single free-text prompt.
- Vocal pronunciation is tuned across 8 languages rather than English only, widening use beyond English-language pop and hip-hop prompts.
- Every output carries a SynthID watermark, letting platforms and listeners verify AI provenance without a separate detection tool.
Cons
- Single-turn generation only: there is no multi-turn editing to refine one section of an already-generated track.
- No melody, hum or reference-track input and no voice cloning; every generation starts from a text or image prompt with no audio conditioning.
- Listed as a preview model in the Gemini API as of September 2026, with no published rate limits and no confirmed availability on Vertex AI Model Garden, where its Lyria 3 and Lyria 3 Pro predecessors already are.
- Refuses prompts that name a specific singer to imitate or ask for existing song lyrics verbatim, which rules out cover-style or parody requests.
Frequently Asked Questions
What are Lyria 3.5's pricing plans in 2026?
Lyria 3.5 costs a flat $0.08 per generated song through the Gemini API, the same per-song price as its predecessor Lyria 3 Pro, regardless of length up to three minutes. The shorter Lyria 3 Clip model costs $0.04 per 30-second clip. Consumer access through Flow Music, the Gemini app and Google Vids is bundled into those products rather than billed per song.
Is Lyria 3.5 free to use?
There is no free tier on the Gemini API for Lyria 3.5; every generated song costs $0.08. It is available at no separate charge to users of Google Flow Music, the Gemini app and Google Vids, where the cost is bundled into those products rather than billed per song.
What are the best alternatives to Lyria 3.5?
Suno and Udio are the alternatives named most often, both offering multi-turn editing and, for Suno, stem export that Lyria 3.5 does not support. Riffusion is a lighter-weight option for real-time, jam-style generation rather than a single finished track per request. Mureka is another option that placed ahead of Lyria 3.5's predecessor on independent listening tests.
Lyria 3.5 or Suno v5.5: which should you pick?
Suno v5.5 led the Artificial Analysis instrumental Music Arena at an Elo of 1190 and supports full stem export plus multi-turn editing, ahead of the predecessor Lyria 3 Pro's Elo of 1119 in the same test. Lyria 3.5 has not been separately arena-tested, but it wins on API pricing simplicity ($0.08 flat) and native image-to-music prompting, which Suno does not offer.
What does it take to start using Lyria 3.5?
Developers call the lyria-3.5 model through the Gemini API's interactions endpoint with a Google AI Studio API key, sending a text prompt and up to 10 optional images. Consumer users can generate a song directly inside Google Flow Music, the Gemini app or Google Vids with no API setup at all.
Top Alternatives
- Suno: Pick Suno if stem export and multi-turn editing of a generated track matter more to you than a flat per-song API price.
- Udio: Pick Udio if you want fast iteration on short clips rather than a single structured 3-minute song per request.
- Riffusion: Pick Riffusion if you want a real-time, jam-style generation loop instead of a single-turn finished song.