Wan-Dancer-14B: Alibaba's Music Driven Dance Video Model Reaches Minute Scale

Share this post:
Wan-Dancer-14B: Alibaba's Music Driven Dance Video Model Reaches Minute Scale
Alibaba's video research team has released Wan-Dancer-14B, a 14 billion parameter model that generates dance video from a music audio input. The model produces continuous output at minute scale and offers two control mechanisms not commonly available in dance generation tools: outfit customization from a reference image, and motion style derived from a reference dancer video. It supports Chinese Classical Dance, K-pop, Street, Tap, and Latin styles.
The 14B parameter scale places Wan-Dancer above the 5 to 8B range common in publicly available motion generation models. It is built on the Wan video backbone, the same foundation used by several recent Alibaba video releases.
Wan-Dancer-14B: music-driven dance video generation
Outfit Customization From a Reference Image
Wan-Dancer-14B's first control mechanism lets a user specify the dancer's appearance through a reference image. The model generates the dance performance with the outfit and visual style of the reference applied to the output character. The capability allows generated videos to feature a specific costume, set, or visual theme without requiring text description alone.
Outfit Customization: the model applies a reference image's visual style to the generated dancer.
Outfit reference conditioning is a meaningful addition to generation driven by music because it separates the choreographic output from the appearance output. The same dance performance, synchronized to the same music, can be rendered with different visual styles by swapping the reference image.
For production applications like music video previs or concept development, this means the same motion sequence does not have to be regenerated for each visual variant. The model handles the appearance transfer from the reference while keeping the choreography consistent with the audio input.
Outfit Change: appearance transfer from a reference image applied to the generated dancer
Movement Reference Control
The second control mechanism takes a reference dancer video as input and uses the motion patterns from that video to shape the choreography generated for the new output. The model extracts movement style from the reference and applies it to a new performance, driven by the specified music track.
Movements Control: motion style transferred from a reference dancer to the new output
Motion reference control is distinct from motion capture transfer. Motion capture requires attaching sensors or markers to a performer and recording their body. Wan-Dancer-14B derives motion characteristics from existing video without any instrumented capture session. A reference performance from any available source can inform the movement style of a newly generated video, synchronized to a different music track.
Dance Style Coverage
Wan-Dancer-14B natively supports five dance styles across two broad categories: traditional and contemporary. Each style has distinct movement vocabularies that the model handles within the same architecture.
Chinese Classical Dance
K-pop Dance
Street Dance
Tap Dance
Latin Dance
The style coverage spans a wider cultural range than most Western dance generation tools have demonstrated publicly. Chinese Classical Dance in particular involves sleeve and fabric movement dynamics that require a different set of temporal patterns from street or K-pop styles. Including it natively in the same model reflects that the training data and the architecture were built for global style coverage.
Single Music, Multiple References
Wan-Dancer-14B supports a generation mode where a single music track drives multiple reference configurations simultaneously, producing a set of distinct output videos that share the same audio but differ in dancer appearance or motion style.
Single music track, multiple reference outputs (variant A)
Single music track, multiple reference outputs (variant B)
Single music track, single reference
The multiple reference mode is a useful production feature for music video concept work. When a director or choreographer needs to evaluate how different performers or costume styles read against the same track, generating a batch of variants from one input avoids the need to reconfigure and regenerate each output separately.
What This Adds to the Music Video Pipeline
Wan-Dancer-14B sits at a specific point in a music video production workflow: it takes a finished audio track and generates a performance at minute scale, with visual control over appearance and motion source. The music and the visual do not need to be developed in sequence. A director with a locked audio mix can generate dance video previs while the production design is still in development, using reference images to test different visual directions against the choreographed output.
For music generation itself, tools like ACE-Step 1.5 produce audio tracks that can serve directly as Wan-Dancer-14B inputs. The combination covers the full pipeline from music creation through visual performance generation without requiring live performers at either stage.
Alibaba has released several video AI tools in the same development window. The world model ABot-World-0 from the AMAP Computer Vision Lab targets real-time scene navigation for previs. BlockVid, from Alibaba DAMO Academy, addresses coherent minute scale video generation from text. Wan-Dancer-14B extends that output duration into a music-driven domain that the other two tools do not target.
Generate text-to-video and image-to-video with the latest AI models at AI FILMS Studio.
Sources
Alibaba Research
Continue Reading
Video & LipSync
- Video Generator
- Text to Video
- Image to Video
- Start-End Frame to Video
- Draw to Video
- Motion Control
- Video Enhancer
- Video Upscaler
- Video to Video LipSync
- Audio to Video LipSync
- Image to Video LipSync
- Video FaceSwap
- Seedance 2
- Vidu Q3 Pro
- Gemini Omni
- Google Veo 3.1
- Kling 3.0 Pro
- Luma Ray 3.2
- LTX 2.3
- Happy Horse 1.1
- Kling 3.0 Motion
- ByteDance Upscaler
- InfiniteTalk
- InsightFace

