# Create video generation task

Submit a video generation task; returns a task ID immediately. Video generation is an async operation --
poll task status via `GET /v1/videos/generations/{taskId}`.

Billing: estimated amount is frozen at submission based on duration; settled by actual usage on success. Frozen amount is refunded on failure.

Supports multiple input modes (specified via the `content` array):
- **Text-to-video**: Only pass a `type: text` prompt
- **Image-to-video (first frame / first+last frame)**: Pass images with `role` set to `first_frame` / `last_frame`
- **Multi-modal reference video** (Seedance 2.0 / MiniMax-H3): reference images, videos, and audio
- **Audio-enabled video** (Seedance 2.0, 1.5 pro): Set `generate_audio: true`. MiniMax-H3 does not support this field.
- **Draft mode** (Seedance 2.5 only): `draft: true` first renders a 480p preview, billed the same as a normal 480p video. After review, submit again with only a `draft_task` item in `content` and the returned task ID to render the 1080p final video. Do not resend the prompt, references, duration, aspect ratio, seed, or `generate_audio`. The two steps are billed separately. The final video's price follows whether the draft step included a reference video; the draft itself is not an input video. The draft task ID is valid for 7 days. Idle models do not support this.

MiniMax-H3 (China / international): see "MiniMax H3 request parameters". `resolution` is `768P` or `2K`, `duration` is an integer 4–15, text-to-video requires a concrete `ratio`, and first/last frames cannot mix with reference assets.

**Image / media reference rules (important):**
- **Regular images**: `image_url.url` must be a **publicly reachable https URL** (object storage / CDN). The upstream engine fetches the URL directly; the proxy does not buffer the image.
- **Reference images that contain real people**: must use `asset://<asset_id>`, pointing to a verified real-person asset. Complete the H5 liveness verification flow via `POST /v1/asset-groups/validate-session` first to obtain a `group_id`/`asset_id`. Passing a real-person image as a URL or inline base64 will be rejected upstream with `InputImageSensitiveContentDetected.PrivacyInformation`.
- **Do not inline base64**: `data:image/...;base64,...` data URIs bloat the request body to MB scale. **The request body is hard-capped at 5 MB; exceeding it returns `413 payload_too_large`.** Upload to object storage and pass a URL, or use `asset://` from the asset library.

Set `video_url_mode` to `upstream` to receive the native output without waiting for archiving, `auto` to permit a saved-copy fallback, or `tos` to wait for a verified archive. Omitting it preserves the account default. This setting applies to the new task and its callbacks; it does not disable background saving. Native expiry may be unknown, and idle enhanced models reject `upstream`.

To combine an explicit mode with `callback_url`, configure the account Webhook Secret first; otherwise creation returns `400 webhook_secret_required`. Existing callbacks without a secret keep their previous provider delivery path when the mode is omitted.

## POST /v1/videos/generations

> Create video generation task

Submit a video generation task; returns a task ID immediately. Video generation is an async operation --
poll task status via `GET /v1/videos/generations/{taskId}`.

Billing: estimated amount is frozen at submission based on duration; settled by actual usage on success. Frozen amount is refunded on failure.

Supports multiple input modes (specified via the `content` array):
- **Text-to-video**: Only pass a `type: text` prompt
- **Image-to-video (first frame / first+last frame)**: Pass images with `role` set to `first_frame` / `last_frame`
- **Multi-modal reference video** (Seedance 2.0 / MiniMax-H3): reference images, videos, and audio
- **Audio-enabled video** (Seedance 2.0, 1.5 pro): Set `generate_audio: true`. MiniMax-H3 does not support this field.

MiniMax-H3 (China / international): see "MiniMax H3 request parameters". `resolution` is `768P` or `2K`, `duration` is an integer 4–15, text-to-video requires a concrete `ratio`, and first/last frames cannot mix with reference assets.

**Image / media reference rules (important):**
- **Regular images**: `image_url.url` must be a **publicly reachable https URL** (object storage / CDN). The upstream engine fetches the URL directly; the proxy does not buffer the image.
- **Reference images that contain real people**: must use `asset://<asset_id>`, pointing to a verified real-person asset. Complete the H5 liveness verification flow via `POST /v1/asset-groups/validate-session` first to obtain a `group_id`/`asset_id`. Passing a real-person image as a URL or inline base64 will be rejected upstream with `InputImageSensitiveContentDetected.PrivacyInformation`.
- **Do not inline base64**: `data:image/...;base64,...` data URIs bloat the request body to MB scale. **The request body is hard-capped at 5 MB; exceeding it returns `413 payload_too_large`.** Upload to object storage and pass a URL, or use `asset://` from the asset library.

### Authentication

`Authorization: Bearer tr-xxx`

### Request Body

Content-Type: `application/json`

- **video_url_mode** ``auto` | `upstream` | `tos``  
  Delivery address: auto prefers a usable native URL with archive fallback; upstream strictly preserves the native URL; tos waits for a verified archive. Omitted uses the account/task default. Existing accounts retain legacy behavior; new accounts default to auto. Does not disable archiving. Idle enhanced models reject upstream. Combining callback_url with an explicit mode requires a configured webhook secret; otherwise creation returns 400 webhook_secret_required.
- **model** `string` **(required)**  
  Video generation model ID. Filter by `type: video` via `GET /v1/models`.
- **content** `object[]` **(required)**  
  Input content for the model, supporting text, images, video, and audio. Supported combinations:
- **content[].type** ``text` | `image_url` | `video_url` | `audio_url` | `draft_task`` **(required)**  
  Input content type:
- **content[].text** `string`  
  Text prompt (used when `type: text`). Supports Chinese and English; recommended max 500 Chinese characters
- **content[].image_url** `object`  
  Image object (used when `type: image_url`)
- **content[].role** ``first_frame` | `last_frame` | `reference_image` | `reference_video` | `reference_audio``  
  Role or purpose of the image/video/audio:
- **content[].video_url** `object`  
  Video object (used when `type: video_url`, Seedance 2.0 only)
- **content[].audio_url** `object`  
  Audio object (used when `type: audio_url`, Seedance 2.0 only)
- **content[].draft_task** `object`  
  Draft task reference (used when `type: draft_task`). `id` is the task ID returned when the draft was created.
- **resolution** ``480p` | `720p` | `1080p` | `768P` | `2K`` (default: `720p`)  
  Output video resolution.
- **ratio** ``16:9` | `4:3` | `1:1` | `3:4` | `9:16` | `21:9` | `adaptive`` (default: `adaptive`)  
  Output video aspect ratio.
- **duration** `integer` (default: `5`)  
  Output video duration (seconds), integer. Set to `-1` to let the model choose an appropriate duration (note: duration affects billing).
- **seed** `integer` (default: `-1`)  
  Random seed for controlling generation randomness. Range: [-1, 2^32-1].
- **generate_audio** `boolean` (default: `true`)  
  Whether to generate audio-enabled video. The model generates matching voice, sound effects, and background music based on the prompt and visual content.
- **draft** `boolean` (default: `false`)  
  Generate a draft preview. Seedance 2.5 only. Idle models do not support it.
- **return_last_frame** `boolean` (default: `false`)  
  Whether to return the last frame of the video as an image (PNG format, no watermark, same resolution as the video).
- **camera_fixed** `boolean` (default: `false`)  
  Whether to fix the camera position.
- **watermark** `boolean` (default: `false`)  
  Whether the generated video includes a watermark
- **service_tier** ``default` | `flex`` (default: `default`)  
  Service tier (cannot be changed for submitted tasks):
- **callback_url** `string`  
  Callback URL for terminal task states. When a task reaches `succeeded` / `failed`, XRToken sends a `POST` request to this URL.
- **safety_identifier** `string`  
  End-user unique identifier for safety auditing. It is recommended to pass a hash of the user ID, up to 64 characters.

### Response

- **id** `string` **(required)**  
  Platform internal task ID (used for status polling)
- **upstream_id** `string`  
  Upstream provider task ID (for reference only)
- **model** `string` **(required)**  
  Model ID used
- **status** ``queued`` **(required)**  
  Initial task status, fixed value `queued`
- **created_at** `string` **(required)**  
  Task creation time (ISO 8601 format)

### Error Codes

- `400`: 
- `401`: 
- `402`: 
- `429`: 
- `502`:
