AI MUSIC GENERATOR
Two ways to score a scene
Describe a style in a handful of words and let the model write the song, or paste a lyric sheet and control the arrangement yourself. Either route can drop the vocal entirely and return a clean instrumental at the length your cut needs. All of it runs in the browser.
Updated:
One workspace, one type dropdown
The music workspace is a single surface with a generation type dropdown at the top holding two entries, Text to Music and Lyric to Music. You pick the type, the form below rebuilds itself with the controls that type needs, and you generate.
The two entries carry separate model rosters, which is the thing most worth knowing before you start. A model available on one is not available on the other, and the controls differ in kind rather than degree. Duration and seed exist on exactly one model in the workspace. Sample rate and bitrate exist on exactly one other.
Results from both types land in the same place. The Artifact Gallery on the My Creations page holds them, and a finished track can be passed into a later generation as input. That is what makes scoring a sequence and then syncing it to picture a single workflow rather than two disconnected ones.
The two ways in
Each has its own page covering the models it offers, the settings it exposes and the work it is genuinely good at.
Choosing between writing the words and describing the sound
The decision is simpler than it looks. If you have lyrics, use Lyric to Music. If you have a mood and no words, use Text to Music. Everything else follows from that.
Text to Music is the faster route and returns two complete tracks from every run, built from the same settings. That pair is useful in itself, because two outputs that fail in the same way point at the style description while two that differ wildly point at the style weight being set too low.
Lyric to Music gives up the second track and gains control. Structure markers place the verse and the chorus, the duration slider sets an exact length, and the seed field makes a promising result repeatable so it can be refined one tag at a time.
Scoring a scene rather than writing a song
Most film work needs a cue, not a song. A track with a singer competes with dialogue, and a track of the wrong length has to be faded out mid phrase, which an audience hears even when they cannot name what is wrong.
Both problems have a control. The vocal comes off with the Instrumental switch on Text to Music, or by leaving the lyrics field empty on ACE-Step 1.5. Length is settled by the duration slider on that same model, and setting it to the length of the shot produces a piece that resolves properly rather than one trimmed to fit.
That combination, an instrumental at an exact duration with a fixed seed, is the reason the Lyric to Music page is worth reading even if you never intend to write a lyric. It is the only route in the workspace that offers all three.
Music in the Nodes Graph Editor
Every music model in the workspace is also available as a node. The Text To Music node accepts a Prompt node as its input, holds the model selector and the model specific controls, and feeds an Audio Viewer node for playback.
The node route earns its extra setup when the track is one step in a longer chain rather than the finished deliverable. Generating a cue and passing it into a lipsync or video node keeps the whole sequence in one graph, which means a change to the style tags re-runs everything downstream instead of starting a new manual pass.
Music tutorials
Walkthroughs of each model in the workspace, plus the open weight releases behind them.
Common questions
Do I need a DAW or any local install to generate music?
No. Both generation types run in the browser. There is no model download, no local inference requirement and no dependency on the hardware in your machine. You sign in and generate from the music workspace.
How do I generate an instrumental with no singing?
Two routes. On Text to Music, switch on Instrumental (No Vocals). On Lyric to Music, select ACE-Step 1.5 and leave the lyrics field empty. The second route is the one that also lets you set an exact duration, which matters when scoring to picture.
How long can a generated track be?
Up to four minutes. ACE-Step 1.5 on the Lyric to Music generation type carries a duration slider running from 1 to 240 seconds. The other two models have no duration control, so their length follows from the lyrics and structure you supply.
Can I get the same track twice?
Only on ACE-Step 1.5, which has a seed field. Fix the seed rather than leaving it at -1 and the run becomes repeatable, which is what makes it possible to change one style tag and hear what that tag did. The other two models have no seed control.
Do I get separated stems for mixing?
No. Every model returns a single mixed file, so the vocal cannot be rebalanced against the instrumentation afterwards. Plan the mix through the style field instead, naming which instruments carry the melody and the bass.
Can I use generated music in a commercial film?
Tracks you generate are yours to use in your productions. Rights and usage terms are covered in the terms and conditions rather than here, so read those before a commercial delivery.
What does a music generation cost?
The exact cost is shown in the form before you submit, and it updates as you change settings. On ACE-Step 1.5 it moves with the duration, so a short cue costs less than a four minute track. Failed tasks are refunded automatically during credit reconciliation.
Can music generation be chained into other tools?
Yes. The Nodes Graph Editor exposes music generation as a Text To Music node that accepts a Prompt node as input and feeds an Audio Viewer node. A generated track can be passed straight into a lipsync or video node, which keeps the whole sequence in one graph.
Video & LipSync
- Video Generator
- Text to Video
- Image to Video
- Start-End Frame to Video
- Draw to Video
- Motion Control
- Video Enhancer
- Video Upscaler
- Video to Video LipSync
- Audio to Video LipSync
- Image to Video LipSync
- Video FaceSwap
- Seedance 2
- FLUX 3
- Minimax H3
- Vidu Q3 Pro
- Grok Imagine Video 1.5
- Gemini Omni
- Google Veo 3.1
- Kling 3.0 Pro
- Luma Ray 3.2
- LTX 2.3
- Happy Horse 1.1
- WAN 2.7
- Kling 3.0 Motion
- ByteDance Upscaler
- InfiniteTalk
- InsightFace