A landscape still is used as the reference for a short generated clip.
An image-to-video model uses an image as conditioning input. Details vary: a system may use a first frame or support additional reference or ending frames. Read its accepted inputs rather than assuming all tools work alike. Hugging Face documents video-generation pipelines.
This differs from transferring an existing performance video's motion to a character. That task needs a suitable motion-conditioned or animation workflow. A still-image input alone does not supply the choreography of a separate clip.
Check required image size, supported frame counts and export format before setup. Inspect the whole output for changing features and unwanted movement, not just its first frame. Generated frames and encoded playback are separate stages; record the playback rate when comparing durations. No output or speed benchmark is asserted by this definition.
The official Wan image-to-video pipeline accepts an input image and documents an optional final-frame image (last_image) for supported models. This supports the distinction between image conditioning and a text prompt; individual tools can expose different controls.
Sources
Primary references checked 11 October 2026. This explanation is not a benchmark of a particular model or computer.