Alibaba Wan/wan-v2-6-image-to-video/Released Dec 2025

WAN 2.6 - Image to Video

Image-to-video with audio and lip-sync

Commercial use
Video

About WAN 2.6 - Image to Video

WAN 2.6 image-to-video is a Large Vision Model designed for animating static images into short videos. It excels in preserving subject identity, facial features, proportions, textures, and overall composition from the input image, producing outputs up to 15 seconds at 1080p resolution with native audio and lip-sync.

Some other noteworthy features of WAN 2.6 image-to-video include multi-shot sequence generation with consistent details across shots and support for reference videos to guide appearance, style, and voice.

Ready to build with WAN 2.6 - Image to Video?

Try WAN 2.6 - Image to Video in the Workbench to prompt it, compare outputs, and iterate on prompts without writing any code. When you're ready to ship, call the same model from our API and build your own apps on top of it.

curl -sSf -X POST https://hub.oxen.ai/api/ai/videos/generate \
    -H "Content-Type: application/json" \
    -H "Authorization: Bearer $OXEN_API_KEY" \
    -d '{
  "model": "wan-v2-6-image-to-video",
  "prompt": "Zoom into the ox, while it is walking forward on the road, in cinematic fashion",
  "input_image": "https://hub.oxen.ai/api/repos/ox/Oxen-AI-Assets/file/main/images/ox_zoom_out_1926_1076.png",
  "duration": 5
}'

Request parameters

Every field WAN 2.6 - Image to Video accepts in the request body. See the API reference for response formats.

POSThttps://hub.oxen.ai/api/ai/videos/generate
FieldTypeRequiredDefaultDescription
promptstringYes
—
Text description of what you want to generate, or the instruction on how to edit the given image.
input_imagestring
uri
Yes
—
Image to use as reference. Must be jpeg, png, gif, or webp.
audio_urlstring
uri
No
—
URL of the audio to use as the background music. If the audio duration exceeds the duration value, the audio is truncated to the first N seconds, and the rest is discarded.
durationintegerNo5
Range: 5 – 15