Realtime TTS-2 is live. Built for realtime conversation that feels human. Read the Realtime TTS-2 announcement

Specialized models

Image-to-video

Start a video from a first frame, end it on a last frame, or keep subjects consistent with reference images

The Video API can start from images as well as a text prompt. Send them in the same POST /v1/videos request, then poll and download the job as usual. There are two ways to use images, and a request uses one or the other:

FieldWhat the images doUse it to
frame_imagesBecome the video's first frame, and optionally its last frameAnimate a keyframe you already made, or move between two keyframes
input_referencesShow the model subjects, characters or objects to include, without fixing any frameKeep a character or product consistent across several videos

What the API accepts for each model:

Wan 3.0 (alibaba/wan-3.0)MiniMax H3 (minimax/minimax-h3)Grok Imagine Video 1.5 Lite (xai/grok-imagine-video-1.5-lite)
First frameYesYesYes
Last frameYes, with a first frameYes, with a first frameNo
Reference imagesUp to 10Up to 9Up to 7

Each image is an image_url item, given by an HTTPS URL. The steps below reuse the INWORLD_API_KEY and VIDEO_REQUEST_KEY variables from Generate your first video. Use a new VIDEO_REQUEST_KEY for each new video. The Python examples need the requests package (python -m pip install requests). The Node.js examples use top-level await, so save them as .mjs files and run them with Node.js 20 or later.

Prepare your images

Upload each image, then replace the example URLs on this page with your own. Each URL must be HTTPS and download without your browser session or additional authentication headers, because the model provider downloads the image itself when generation starts. The API checks the URL when you create the job, but not the image behind it.

URL rules checked at creation

The API rejects the create request with 400 if an image URL breaks these rules:

  • It must use HTTPS and name a public host. HTTP URLs, private-network IP addresses, and hosts such as localhost or names ending in .internal or .local are rejected.
  • It can be at most 4,096 bytes long.
  • It must be a link, not inline data. Base64 data: URLs are rejected. Upload the image somewhere the provider can download it, such as your own storage bucket.

Recommendations

These aren't checked when you create the job. If the provider can't download or use an image, the job fails after it was created:

  • Keep signed URLs valid for at least an hour. The provider downloads the image when generation starts, not when you create the job. One hour is a recommendation, not an API limit.
  • Check the URL before you send it. Confirm it opens and returns an image file. The API doesn't download or inspect the image.

The model provider downloads each image directly from the URL you send, so a signed URL, including its signature, is shared with that provider. Use short-lived, read-only URLs.

Start from a first frame

Put the image in frame_images with frame_type: "first_frame". The prompt describes what happens next:

bash
curl -X POST https://api.inworld.ai/v1/videos \
  -H "Authorization: Basic $INWORLD_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $VIDEO_REQUEST_KEY" \
  -d '{
    "model": "alibaba/wan-3.0",
    "prompt": "The fox lifts its head, looks toward the camera, then trots out of frame",
    "seconds": 5,
    "resolution": "720p",
    "frame_images": [
      {
        "type": "image_url",
        "image_url": { "url": "https://storage.example.com/keyframes/fox-start.png" },
        "frame_type": "first_frame"
      }
    ]
  }'

The job comes back like any other, with aspect_ratio: "adaptive":

json
{
  "id": "video_3c59dc048e8850243be8079a5c74d079",
  "object": "video",
  "model": "alibaba/wan-3.0",
  "status": "queued",
  "created_at": 1790000000,
  "seconds": 5,
  "resolution": "720p",
  "aspect_ratio": "adaptive"
}

The video takes its shape from the first frame, so leave aspect_ratio out or set it to "adaptive". Any other value is rejected with 400. seconds and resolution work as they do for text-only videos.

Each example prints the job's id, status and aspect ratio. To poll and download the job, use the scripts in Video generation with this request body, or poll as in Generate your first video. The job object does not list the images you sent.

End on a last frame

To control where the video ends, add a second item with frame_type: "last_frame". The model generates the motion between the two images:

json
"frame_images": [
  {
    "type": "image_url",
    "image_url": { "url": "https://storage.example.com/keyframes/door-closed.png" },
    "frame_type": "first_frame"
  },
  {
    "type": "image_url",
    "image_url": { "url": "https://storage.example.com/keyframes/door-open.png" },
    "frame_type": "last_frame"
  }
]
  • A last frame needs a first frame. frame_images with only a last_frame is rejected.
  • Not every model accepts a last frame; see the table above. A last_frame for a model that doesn't accept one is rejected with 400.
  • Send at most one first_frame and one last_frame. Their order in the array doesn't matter.

Keep subjects consistent with reference images

Put images of the subjects in input_references. Unlike frame images, references don't fix any frame: the model uses them to decide what the subjects look like, and the prompt decides what happens.

bash
curl -X POST https://api.inworld.ai/v1/videos \
  -H "Authorization: Basic $INWORLD_API_KEY" \
  -H "Content-Type: application/json" \
  -H "Idempotency-Key: $VIDEO_REQUEST_KEY" \
  -d '{
    "model": "alibaba/wan-3.0",
    "prompt": "The woman in the red coat walks into a busy cafe and sits by the window",
    "seconds": 5,
    "resolution": "720p",
    "aspect_ratio": "16:9",
    "input_references": [
      {
        "type": "image_url",
        "image_url": { "url": "https://storage.example.com/characters/red-coat.png" }
      }
    ]
  }'
  • The number of references a model accepts is in the table above. Sending more is rejected with 400.
  • aspect_ratio works as it does for text-only videos.
  • A request can't combine input_references with frame_images. Choose one.
  • References must be images. Video and audio references are rejected.

Refer to more than one image

The API passes references to the model in the order you send them. In the prompt, refer to each image by its position, Image 1 for the first item and Image 2 for the second, and describe what it shows, so the model can tell the subjects apart:

json
"prompt": "The woman in Image 1 carries the teapot in Image 2 into a sunlit kitchen",
"input_references": [
  {
    "type": "image_url",
    "image_url": { "url": "https://storage.example.com/characters/red-coat.png" }
  },
  {
    "type": "image_url",
    "image_url": { "url": "https://storage.example.com/products/teapot.png" }
  }
]

Billing

You pay the same per-second rate as for a text-only video at the same model and resolution. Some models also charge for each input image, frame or reference. Current rates are on the models page; see Billing for how jobs are billed. A failed job costs nothing, including its images.

Errors

Image fields are validated when you create the job. A rejected request returns 400 with code: "invalid_request" and a param that points to the item at fault. The message never repeats the URL.

paramCause
frame_images[0].type, input_references[0].typeThe item's type isn't image_url
frame_images[0].frame_typeframe_type is missing or not first_frame or last_frame, two items have the same type, or the model doesn't support a last frame
frame_imagesA last_frame was sent without a first_frame
frame_images[0].image_url.urlThe URL is missing, isn't HTTPS, names a private host, or is longer than 4,096 bytes
input_referencesToo many references for the model, or combined with frame_images
aspect_ratioSet to a value other than adaptive together with frame_images

If the job fails after creation

A valid URL doesn't guarantee a usable image. The job can still fail after it was created, with status: "failed". To recover:

  1. Read error.code on the job. An image problem usually shows up as invalid_request (the provider rejected the request or the image) or generation_failed, and an image the provider's content policy blocks as content_filtered. See Job failures.
  2. Check that each URL still downloads without a browser session or extra headers, and that it returns an image file.
  3. Submit the corrected request with a new Idempotency-Key. Resending the old key returns the failed job.