Gemini Omni Video
Provided by Google — Gemini Omni Video on Fuser
Gemini Omni 1.1 Flash is Google's multimodal video model for text-to-video, image-guided video, and video editing with synchronized audio, grounded in Gemini's real-world knowledge and physics understanding. It generates from 360p up to 4K, accepts reference images, and applies conversational edits that preserve character and scene consistency across turns.
FLUX.3
FLUX.3 is Black Forest Labs' frontier video model, generating 5 to 20 second clips with optional synchronized audio at up to 1080p. Describe your scene for text-to-video, provide a start image to animate it, add an end image to interpolate a smooth transition between two frames, or provide a video to extend it. Draft mode trades resolution for much faster, cheaper 720p generations — ideal for iterating on a shot before a final render.
Grok Imagine
Grok Imagine Video generates video with synchronized audio from text prompts, single images, or multiple reference images using xAI's Grok model. In text-to-video mode, describe your scene to produce realistic footage up to 15 seconds. Provide a single image to animate it as the first frame. Supply two to seven reference images and cite them in your prompt as @Image1 through @Image7 to blend their visual features into a single coherent video. Version 1.5 adds 1080p output for text-to-video and image-to-video.