TEXT TO MUSIC
Describe the sound, then hear it
Write a handful of style descriptors and get two finished tracks back from the same run. Seven model versions, three sliders that shape how closely the result follows you, and a switch that strips the vocal out entirely for underscore. No DAW and nothing to install.
Updated:
What text to music does
Text to Music writes both the composition and the words from a written description. You supply the style and the model decides the melody, the arrangement, the instrumentation and, unless you switch it off, the vocal.
That makes it the faster of the two music generation types and the one to reach for when you need a cue rather than a song. The Instrumental switch is what turns it into a scoring tool, because it removes the singer and leaves the arrangement, which is what a scene under dialogue actually needs.
The text to music form, control by control
Eleven controls, and three of them are hidden behind switches until you turn those on.

The generation type dropdown. It decides which models the picker below will offer.
Pick the generation type first
The music workspace opens on a dropdown with two entries, Text to Music and Lyric to Music. They are not two routes to the same place. Each carries its own model roster, and the form beneath rebuilds when you switch.
Text to Music is the one that writes both the words and the music for you. If you already have lyrics and want them sung as written, the other entry is the one you want.
Choose the model
Text to Music runs a single model, Suno Music, and the label beside it states what a run gives you back. Two tracks, style guided.
Two tracks per generation is the detail worth planning around. Every run returns a pair of variations built from the same settings, so comparison is built into the workflow rather than something you pay for twice.

One model on this generation type, labelled with what it returns.

Seven versions. v5.5 is the default and the one most work should use.
Then choose the version
A second dropdown holds seven versions, running v5.5, v5, v4.5 All, v4.5 Plus, v4.5, v4 and v3.5. This control is separate from the model picker and easy to miss.
v5.5 is the default. The older versions stay available because their sonic character differs, which is useful when you want a retro or lo fi quality that the current version cleans away.
Describe the style
Style is the only required field on this form and it holds 120 characters, which is deliberately short. It wants a list of descriptors rather than a sentence.
The buttons underneath append a tag to whatever is already there, covering cinematic, orchestral, epic, emotional, ambient, electronic, rock, pop, jazz, classical, hip hop, EDM, country, folk, metal and punk. Combining two or three buttons with your own instrument or tempo terms gives tighter control than either method on its own.

The Style field, with tag buttons that append rather than replace.

One field, two jobs, decided by the Instrumental switch further down.
Add words, or a theme
This field runs to 4,100 characters and changes meaning depending on a switch lower in the form. With Instrumental off, whatever you write here is treated as lyrics and sung.
With Instrumental on, it reads as a thematic description instead. For purely instrumental work the Style field still carries more weight, so this one is best used for the mood or narrative the piece should follow.
List what to keep out
Negative Tags works the way a negative prompt does on an image model. It names the qualities you do not want to hear, and it arrives already carrying distortion and low quality.
Exclusions belong here rather than in the Style field. Asking the main description to avoid something tends to introduce it, while listing it as a negative tag works as intended.

Negative Tags arrives pre-filled rather than empty.

Title and Persona ID, both revealed by Custom Mode.
Name the track and set a persona
Title takes 200 characters and does more than label the file. The name frames how the model approaches the composition, so a descriptive title is worth writing rather than leaving blank.
Persona ID offers Style or Voice, and it applies to v5.5 through v5 only. Both fields appear only when Custom Mode is switched on.
Style Weight decides how literally it reads you
Style Weight runs 0 to 1 and starts at 0.65. High values hold the output tightly inside the genre and texture you described. Low values give the model room to interpret.
When a generation keeps drifting away from the brief, this is the first slider to raise. When every run sounds like the same safe idea, lower it.

Style Weight. All three sliders share the same range and default.

Weirdness. Higher is not better, it is stranger.
Weirdness controls how conventional it stays
Weirdness also runs 0 to 1 from a 0.65 default. Below 0.5 you get familiar chord movement and predictable arrangement. Above 0.8 the model reaches for unusual intervals and odd time signatures.
For anything sitting under dialogue, keep it between 0.5 and 0.65. A raised Weirdness value is a creative choice rather than a quality improvement, and it makes a cue harder to cut against picture.
Audio Weight balances polish against range
Audio Weight is the last of the three sliders, on the same 0 to 1 scale and the same 0.65 default. Higher values favour a clean, defined mix. Lower values allow a rawer character.
Leaving all three at 0.65 is a reasonable starting position. Moving one at a time is the only way to hear what each is actually doing, because their effects overlap.

Audio Weight, the third of the three, trading polish against variety.

Two switches. Each one changes which other controls exist.
Custom Mode and Instrumental
Custom Mode unlocks the Title and Persona ID fields. With it off, those controls are not on the form at all, which is why they can appear to be missing.
Instrumental removes the vocal layer entirely, and it is the correct setting for underscore and background beds where a singer would fight the dialogue. Switching it off reveals a Vocal Gender selector offering Male or Female.
Models available for text to music
One model on this generation type, with the version chosen separately. The other music models live on Lyric to Music.
| Model | Best for | Notes |
|---|---|---|
| Suno Music Default | Complete songs and instrumental cues from a style description | The only model on this generation type. Seven selectable versions from v3.5 to v5.5, and every run returns two tracks rather than one. |
The Model Version dropdown is a separate control from the model picker, and it is the one that changes output quality. v5.5 outputs at 44.1 kHz, the standard for streaming and broadcast delivery. The older versions are kept for their sonic character rather than as a fallback.
Settings reference
Every control in the text to music form, what it changes and the values it accepts. Three of them appear only once a switch is turned on.
Style
The genre, mood and instrumentation. Required, and the shortest field on the form, so it wants a list of descriptors rather than prose. Tag buttons below it append rather than replace.
Up to 120 characters
Model Version
Which Suno version runs. Separate from the model picker and easy to overlook. Older versions are kept for their sonic character rather than as a fallback.
v5.5, v5, v4.5 All, v4.5 Plus, v4.5, v4, v3.5
Prompt / Lyrics
Read as lyrics when Instrumental is off, and as a thematic description when it is on. One field whose meaning is set by a switch further down the form.
Up to 4,100 characters
Negative Tags
Qualities to keep out of the result. Pre-filled with distortion and low quality. Exclusions work here and tend to backfire when written into the Style field instead.
Free text, pre-filled by default
Title
Names the track and influences how the model frames the composition. Appears only with Custom Mode switched on.
Up to 200 characters
Persona ID
Applies a saved style or voice character. Applies to v5.5 through v5 only, and appears only with Custom Mode switched on.
Style or Voice
Style Weight
How literally the model follows the Style field. Raise it when generations drift from the brief, lower it when every run sounds the same.
0 to 1, default 0.65
Weirdness
How unconventional the harmony and structure become. Above 0.8 produces unusual intervals and odd time signatures on purpose.
0 to 1, default 0.65
Audio Weight
Trades production polish against variety. Higher values give a cleaner, more defined mix.
0 to 1, default 0.65
Custom Mode
Reveals the Title and Persona ID fields. With it off those controls are absent from the form rather than disabled.
Switch, on by default
Instrumental (No Vocals)
Removes the vocal layer. The correct setting for underscore and background beds. Switching it off reveals a Vocal Gender selector.
Switch
Vocal Gender
The voice character for the generated singer. Present only when Instrumental is switched off.
Male or Female
Writing a style that produces something specific
The Style field holds 120 characters. That limit is the clearest signal about how it wants to be used, because it rules out a paragraph and forces a list. Every word has to earn its place.
A single genre term produces the average of that genre. "Cinematic" returns generic orchestral wash. "Cinematic, sweeping strings, low brass undertone, slow build, war film" names an instrument section, a register, an arc and a reference point, and the output narrows accordingly.
The most useful additions are the ones a genre tag cannot imply. Tempo, the specific instruments carrying the melody and the bass, the production era, and whether the piece should build or hold steady. Three or four of those alongside a genre gives the model more to work with than six genre tags stacked together.
Exclusions belong in Negative Tags rather than here. Writing "no drums" into the Style field tends to introduce drums, because the field is read as a description of what the piece contains.
Using the two tracks you get back
Every run returns two complete tracks generated from identical settings. They are not a draft and a final. They are two independent readings of the same brief, and they often differ more than a slider adjustment would produce.
That changes how to iterate. The instinct after a disappointing result is to reach for a slider, but with two outputs in front of you the more informative question is whether they failed in the same way. Two tracks that miss the brief identically point at the Style field. Two tracks that differ wildly point at Style Weight being too low.
Running the same settings a second time is also a reasonable move rather than a wasted one, since it gives four variations to choose between. For a cue that has to sit under a specific edit, having four candidates is usually faster than refining a single generation toward an exact target.
Model tutorials
Deeper coverage of the music models, with the controls each one exposes and where they differ.
Text to music questions
Which Suno version should I use?
v5.5 unless you have a reason not to. It is the default and outputs at 44.1 kHz, the standard for streaming and broadcast delivery, so it drops into a timeline without conversion. The older versions remain because their character differs, which is useful for retro or lo fi work.
Where did the Title field go?
It is hidden behind the Custom Mode switch, along with Persona ID. With Custom Mode off both fields are absent from the form rather than greyed out, which is why they can look like they were removed.
How do I generate music with no singing?
Switch on Instrumental (No Vocals). That removes the vocal layer entirely and is the right setting for underscore, background beds and anything that sits beneath dialogue. The Prompt / Lyrics field then reads as a thematic description instead of words to sing.
What do the three sliders actually change?
Style Weight sets how literally the model follows your Style text. Weirdness sets how conventional the harmony and structure stay. Audio Weight trades production polish against variety. All three run 0 to 1 and start at 0.65, and their effects overlap, so move one at a time.
Should I raise Weirdness to get a better result?
No. Higher Weirdness produces stranger output rather than better output. Above 0.8 the model reaches for unusual intervals and odd time signatures on purpose. For a cue that has to sit under dialogue, keep it between 0.5 and 0.65.
What is the difference between Text to Music and Lyric to Music?
Text to Music writes both the words and the music from a style description, and runs Suno Music. Lyric to Music expects you to supply the words, and runs a different pair of models with their own controls including a duration slider. They are separate generation types with separate rosters.
What does a text to music generation cost?
The exact cost is shown in the form before you submit, and it is the same whichever Suno version you select. Failed tasks are refunded automatically during credit reconciliation.
Other ways to make music
Text to Music is one of two routes into the music workspace.
Video & LipSync
- Video Generator
- Text to Video
- Image to Video
- Start-End Frame to Video
- Draw to Video
- Motion Control
- Video Enhancer
- Video Upscaler
- Video to Video LipSync
- Audio to Video LipSync
- Image to Video LipSync
- Video FaceSwap
- Seedance 2
- FLUX 3
- Minimax H3
- Vidu Q3 Pro
- Grok Imagine Video 1.5
- Gemini Omni
- Google Veo 3.1
- Kling 3.0 Pro
- Luma Ray 3.2
- LTX 2.3
- Happy Horse 1.1
- WAN 2.7
- Kling 3.0 Motion
- ByteDance Upscaler
- InfiniteTalk
- InsightFace