Wan 3.0 Video
Provided by Fal — Learn More — Wan 3.0 Video on Fuser
Wan 3.0 generates video from text, start and end frames, or multimodal references with native audio and coherent motion. Both Standard and Prime support output from 2 to 30 seconds at 480p, 720p, or 1080p. Prime accelerates generation for creative iteration, while image, video, and audio references guide subjects, scenes, and sound.
Vidu Reference
Vidu Q3 Mix generates 1 to 16 second videos with native audio from one to four reference images at up to 1080p. Q1 and the original Vidu reference models remain available. Vidu Reference is a video generation model that enables seamless interaction between multiple subjects, including characters, props, objects, and environments in the same scene. This model is ideal for creating videos with complex scenes and multiple characters, where characters interact naturally within the same scene. Vidu supports feature fusion, allowing elements from different subjects—such as the front of Character A and the back of Character B—to merge seamlessly into a new character or object.
Wan-2.1
Wan-2.1 is an advanced and powerful visual generation model developed by Tongyi Lab. It can generate videos based on text, images and other control signals. Wan-2.1 excels at generating realistic videos featuring extensive body movements, complex rotations, dynamic scene transitions, and fluid camera motions, and can accurately simulate real-world physics and realistic object interactions, while offering movie-like visuals with rich textures and a variety of stylized effects. It can also create text and dynamic text effects in videos directly from text prompts.