LYRICS TO MUSIC
Write the words, then hear them sung
Paste a lyric sheet and get a finished song back. Structure markers tell the model where the chorus lifts, the duration slider runs to four minutes, and leaving the lyrics field empty returns a clean instrumental instead. Nothing to install.
Updated:
What lyrics to music does
Lyric to Music takes words you wrote and sets them to music. You keep control of the song, and the model handles the melody, the arrangement and the performance.
The name undersells it for filmmakers. On the default model the lyrics field is optional, and an empty one returns a scored instrumental at a length you set. That combination, a duration slider and no vocal, is what the other generation type cannot do, and it is what a cue under dialogue usually needs.
The lyrics to music form, control by control
Ten controls across the two models, and the one you pick decides which four of them you get.

The generation type dropdown, with Lyric to Music selected.
Select the generation type
The music workspace opens on a dropdown holding Text to Music and Lyric to Music. Each carries its own model roster, so the choice here decides which models the picker below will show you.
The name understates what this generation type does. Lyrics are optional on one of the two models here, so this is also where you come for an instrumental with a duration you control.
Choose between the two models
ACE-Step 1.5 carries a Default and Fast label and is already selected. MiniMax Music 2.5 carries High Quality and Latest. The picker is searchable, which matters less with two entries than it will later.
This choice rewrites the bottom half of the form. ACE-Step 1.5 gives you a duration slider and a seed. MiniMax Music 2.5 replaces both with sample rate, bitrate and audio format. There is no overlap between the two sets.

Two models, and the labels beside them describe the trade rather than the ranking.

The style field, whose label changes with the selected model.
Describe the style
The label reads Style Tags on ACE-Step 1.5 and Style Prompt on MiniMax Music 2.5. That is not cosmetic. The first wants comma separated tags, the second accepts a fuller description, and both hold 3,000 characters.
The buttons below append without typing, covering rap, k-pop, hard house, synth-pop, classical, jazz, country, alternate-rock, european, rock, R&B, EDM, reggae, blues, folk, metal, punk, disco, soul and funk. Adding a tempo in beats per minute alongside them gives the model a target that a genre tag cannot imply.
Write the lyrics, or leave them out
On ACE-Step 1.5 this field is optional and the label says so. Leaving it empty returns an instrumental with no vocal line at all, which is the single most useful behaviour on this page for anyone scoring picture.
Eight buttons sit underneath. Intro, Verse, Chorus, Bridge and Outro insert a lowercase structure marker in square brackets. New Line inserts a single break, Pause inserts a double break that reads as a rest, and Accompaniment inserts a ## marker used in pairs to wrap an instrumental direction.

The lyrics field with its eight markers. Note the optional label on ACE-Step 1.5.

Duration, on ACE-Step 1.5 only. Marks at 1s, 60s, 120s and 240s.
Set the length, on ACE-Step 1.5
The slider runs from 1 second to 240 seconds and starts at 60. Four minutes covers a full song, and the low end is short enough to generate a sting or a transition element.
Set this to match the cut rather than generating long and trimming. The model paces an arrangement across whatever length it is given, so a 30 second setting resolves at 30 seconds instead of producing a four minute piece cut short.
Lock the seed once something works
Seed defaults to -1, which the helper text under the field confirms means random. The dice button on the right fills in a number for you.
A fixed seed is what makes iteration possible. Hold the seed and change one style tag, and you hear what that tag did. With a random seed every run is a different piece and comparison tells you nothing.

Seed, on ACE-Step 1.5 only. The dice button fills in a random value.

Sample Rate, on MiniMax Music 2.5 only. 44100 Hz is the default.
Set the sample rate
Switching to MiniMax Music 2.5 removes the duration slider and the seed field, and puts three output controls in their place. The first is Sample Rate, offering 16000 Hz, 24000 Hz, 32000 Hz and 44100 Hz.
44100 Hz is the default and the standard for streaming and broadcast delivery, so it drops into a timeline without conversion. The lower rates exist for speech style output and small file delivery.
Set the bitrate
Bitrate offers 32, 60, 64, 128 and 256 kbps, and defaults to the top of that range. It controls compression, so it decides both the file size and how much detail survives in the mix.
Stay at 256 kbps for anything heading into an edit. The lower settings are for previewing an idea quickly, and the difference becomes audible once a track sits under dialogue.

Bitrate, on MiniMax Music 2.5 only. 256 kbps is the default.

Audio Format, on MiniMax Music 2.5 only. Three containers.
Choose the output format
The last of the three controls offers MP3, PCM and FLAC, with MP3 as the default. This decides the container rather than anything about the music itself.
MP3 is compressed and opens anywhere. FLAC is lossless compression, so it keeps full quality at roughly half the size of raw audio. PCM is uncompressed and produces the largest files.
What comes back
The generated track appears on the right with a player, a download button and a delete button. Each run returns a single file rather than a pair, which is where this generation type differs from Text to Music.
Track length on MiniMax Music 2.5 follows from the lyrics and structure you supplied, since it has no duration control. The example here runs 2 minutes 48 seconds from a lyric sheet with marked sections.

The finished result. One file, with a player, a download and a delete.
Models available for lyrics to music
Two models that split cleanly. One controls length and can drop the vocal, the other controls the output file.
| Model | Best for | Notes |
|---|---|---|
| ACE-Step 1.5 Default | Instrumental cues, and any track that needs an exact length | The default. Lyrics are optional, so an empty field returns an instrumental. Duration runs 1 to 240 seconds and a seed field makes runs repeatable. |
| MiniMax Music 2.5 | Complete songs where the delivery format matters | Lyrics are required. Exposes sample rate to 44100 Hz, bitrate to 256 kbps and MP3, PCM or FLAC output. No duration or seed control. |
The two models expose mutually exclusive controls rather than overlapping ones. Selecting ACE-Step 1.5 shows a duration slider and a seed field. Selecting MiniMax Music 2.5 removes both and shows sample rate, bitrate and audio format instead. Confirm the setting a workflow depends on is present on the model you plan to use.
Settings reference
Every control in the lyrics to music form. Several are present on one model and absent on the other.
Style Tags / Style Prompt
Genre, mood, instrumentation and tempo. The label changes with the model. ACE-Step 1.5 wants comma separated tags, MiniMax Music 2.5 accepts a fuller description. Tag buttons append rather than replace.
Up to 3,000 characters
Lyrics
The words sung, and the structure of the arrangement. Optional on ACE-Step 1.5, where an empty field returns an instrumental. Required on MiniMax Music 2.5.
Up to 3,000 characters
Structure markers
Inserted by the buttons under the lyrics field. Intro, Verse, Chorus, Bridge and Outro place a marker in square brackets. They change the arrangement rather than only the vocal delivery.
Five markers, inserted at the cursor
Pause and Accompaniment
Pause inserts a double line break, which reads as a rest. Accompaniment inserts a ## marker, used in pairs to wrap a passage the model should treat as an instrumental direction rather than words to sing.
Inserted at the cursor
Duration
Track length. Present on ACE-Step 1.5 only. Set it to match the cut, since the model paces the arrangement across whatever length it is given.
1 to 240 seconds, default 60
Seed
Makes a run repeatable. Present on ACE-Step 1.5 only. Hold it fixed to hear what a single style tag changed.
Any integer, or -1 for random
Select Sample Rate
Frequency detail in the output file. Present on MiniMax Music 2.5 only. 44100 Hz is the standard for streaming and broadcast delivery.
16000, 24000, 32000 or 44100 Hz
Select Bitrate
Compression, so it drives file size and how much detail survives in the mix. Present on MiniMax Music 2.5 only.
32, 60, 64, 128 or 256 kbps
Select Audio Format
Container and compression type. Present on MiniMax Music 2.5 only. FLAC is lossless, PCM is uncompressed, MP3 opens anywhere.
MP3, PCM or FLAC
Structure markers, and why they matter more than they look
The five markers under the lyrics field look like formatting. They are closer to an arrangement instruction. A verse marker and a chorus marker tell the model where the energy should sit low and where it should lift, and the instrumentation follows that shape.
The same words with no markers produce a flatter result. The model still sings them, but without a marked hook it has no reason to return to one, so the track reads as a continuous passage rather than a song with a shape.
The Accompaniment marker is the least obvious and the most useful. It inserts a ## symbol, and a passage wrapped in ## on both sides is treated as a direction rather than words to sing. That is how you describe an instrumental intro, a break or a texture change inside the lyrics field without any of it being sung aloud.
Choosing between the two models
The two models here do not sit on a quality ladder. They expose different controls, and the right one depends on which of those controls the job needs.
ACE-Step 1.5 owns length and repeatability. The duration slider means a cue can be generated at the length of the shot rather than trimmed to fit, and the seed field means a promising result can be refined one tag at a time instead of regenerated from scratch. Optional lyrics make it the only route to an instrumental on this generation type.
MiniMax Music 2.5 owns the output file. Sample rate to 44100 Hz, bitrate to 256 kbps and a choice of MP3, PCM or FLAC give you delivery control the other model does not offer. It requires lyrics and has neither duration nor seed, so it suits a finished song more than an iterative scoring pass.
Model tutorials
Deeper coverage of both models on this generation type, with sample settings and where each one struggles.
Lyrics to music questions
Why did the duration slider disappear?
You switched to MiniMax Music 2.5, which has no duration control. Track length there follows from the lyrics and structure you supply. The slider belongs to ACE-Step 1.5, along with the seed field, and both return when you select it again.
What do the structure markers actually do?
They segment the arrangement, not just the vocal. A lyric sheet with a marked chorus produces a track that returns to a recognisable hook, while the same words with no markers produce something flatter. The buttons insert them in square brackets at the cursor position.
What is the ## marker for?
The Accompaniment button inserts it, and it works in pairs. Wrap a passage in ## on both sides and the model treats the text inside as an instrumental direction rather than words to sing, which is how you describe an intro or a break without it being sung aloud.
Which of the two models should I pick?
ACE-Step 1.5 for anything where length matters or where you want an instrumental, because it is the only one with a duration slider and optional lyrics. MiniMax Music 2.5 when you have finished lyrics and the delivery format matters, since it is the only one exposing sample rate, bitrate and file format.
Can I get the same track twice?
On ACE-Step 1.5, yes. Fix the seed rather than leaving it at -1 and the run becomes repeatable. MiniMax Music 2.5 has no seed field, so two runs with identical settings will not return identical audio. Download anything you want to keep before running again.
Do I get separated stems?
No. Both models return a single mixed file, so the vocal cannot be rebalanced against the instrumentation afterwards. Plan the mix through the style field instead, naming the instruments that should carry the melody and the bass.
What does a lyrics to music generation cost?
The exact cost is shown in the form before you submit. On ACE-Step 1.5 it moves with the duration you set, so a short cue costs less than a four minute track. Failed tasks are refunded automatically during credit reconciliation.
Other ways to make music
Lyrics to Music is one of two routes into the music workspace.
Video & LipSync
- Video Generator
- Text to Video
- Image to Video
- Start-End Frame to Video
- Draw to Video
- Motion Control
- Video Enhancer
- Video Upscaler
- Video to Video LipSync
- Audio to Video LipSync
- Image to Video LipSync
- Video FaceSwap
- Seedance 2
- FLUX 3
- Minimax H3
- Vidu Q3 Pro
- Grok Imagine Video 1.5
- Gemini Omni
- Google Veo 3.1
- Kling 3.0 Pro
- Luma Ray 3.2
- LTX 2.3
- Happy Horse 1.1
- WAN 2.7
- Kling 3.0 Motion
- ByteDance Upscaler
- InfiniteTalk
- InsightFace