# 创建视频生成任务

提交视频生成任务，立即返回任务 ID。视频生成为异步操作，需通过
`GET /v1/videos/generations/{taskId}` 轮询任务状态。

计费方式：提交时按预估时长冻结费用，视频生成成功后按实际用量结算；生成失败则退还冻结金额。

支持多种输入模式（通过 `content` 数组指定）：
- **文生视频**：仅传入 `type: text` 的提示词
- **图生视频（首帧/首尾帧）**：传入图片，`role` 设为 `first_frame` / `last_frame`
- **多模态参考生视频**（Seedance 2.0 / MiniMax-H3）：传入参考图片、参考视频、参考音频
- **有声视频**（Seedance 2.0、1.5 pro）：设置 `generate_audio: true`，模型自动生成同步音频。MiniMax-H3 不支持此字段。
- **样片模式**（仅 Seedance 2.5）：`draft: true` 先出 480p 预览，计费与直接生成 480p 相同。确认后用返回的任务 ID 再提交一次，`content` 只放 `draft_task`，生成 1080p 成片。成片不要再传提示词、参考素材、时长、宽高比、种子和 `generate_audio`。两步分开计费；成片单价看样片那一步有没有参考视频，样片本身不算输入视频。样片任务 7 天内有效。闲时模型不支持。

MiniMax-H3（国内站 / 国际站）参数标准见文档「MiniMax H3 请求参数」：`resolution` 为 `768P` 或 `2K`，`duration` 为 4–15 的整数，文生必须给具体 `ratio`，首尾帧不能和参考素材混用。

**图片 / 媒体引用规范（重要）：**
- **普通图片**：`image_url.url` 请传 **公网可访问的 https URL**（如对象存储 / CDN）。上游引擎按 URL 直接拉取，不经过本服务内存。
- **含真人的参考图**：必须用 `asset://<asset_id>` 引用一个经过验证的真人素材（先调 `POST /v1/asset-groups/validate-session` 完成 H5 真人活体验证，拿到 `group_id`/`asset_id`）。直接传真人图 URL 或 base64 都会被上游 `InputImageSensitiveContentDetected.PrivacyInformation` 拒绝。
- **不要 inline base64**：`data:image/...;base64,...` 这种内联编码会让请求体迅速膨胀到 MB 级。**请求体硬上限 5 MB，超过返回 `413 payload_too_large`**。需要传图请上传到对象存储后传 URL，或用 `asset://` 引用素材库。

创建时可指定 `video_url_mode`：`upstream` 直接返回原生地址，不等待保存；`auto` 优先原生地址，必要时回退到已保存版本；`tos` 等待我方确认保存完成。省略时使用账号默认策略。该设置会固化到新任务和回调，不影响后台保存；原生有效期可能未知，闲时升档模型不支持 `upstream`。

显式模式与 `callback_url` 同时使用前须配置账号 Webhook Secret，否则创建返回 `400 webhook_secret_required`。未配置 secret 且省略模式的原有回调客户继续沿用旧链路。

## POST /v1/videos/generations

> Create video generation task

Submit a video generation task; returns a task ID immediately. Video generation is an async operation --
poll task status via `GET /v1/videos/generations/{taskId}`.

Billing: estimated amount is frozen at submission based on duration; settled by actual usage on success. Frozen amount is refunded on failure.

Supports multiple input modes (specified via the `content` array):
- **Text-to-video**: Only pass a `type: text` prompt
- **Image-to-video (first frame / first+last frame)**: Pass images with `role` set to `first_frame` / `last_frame`
- **Multi-modal reference video** (Seedance 2.0 / MiniMax-H3): reference images, videos, and audio
- **Audio-enabled video** (Seedance 2.0, 1.5 pro): Set `generate_audio: true`. MiniMax-H3 does not support this field.

MiniMax-H3 (China / international): see "MiniMax H3 request parameters". `resolution` is `768P` or `2K`, `duration` is an integer 4–15, text-to-video requires a concrete `ratio`, and first/last frames cannot mix with reference assets.

**Image / media reference rules (important):**
- **Regular images**: `image_url.url` must be a **publicly reachable https URL** (object storage / CDN). The upstream engine fetches the URL directly; the proxy does not buffer the image.
- **Reference images that contain real people**: must use `asset://<asset_id>`, pointing to a verified real-person asset. Complete the H5 liveness verification flow via `POST /v1/asset-groups/validate-session` first to obtain a `group_id`/`asset_id`. Passing a real-person image as a URL or inline base64 will be rejected upstream with `InputImageSensitiveContentDetected.PrivacyInformation`.
- **Do not inline base64**: `data:image/...;base64,...` data URIs bloat the request body to MB scale. **The request body is hard-capped at 5 MB; exceeding it returns `413 payload_too_large`.** Upload to object storage and pass a URL, or use `asset://` from the asset library.

### Authentication

`Authorization: Bearer tr-xxx`

### Request Body

Content-Type: `application/json`

- **video_url_mode** ``auto` | `upstream` | `tos``  
  Delivery address: auto prefers a usable native URL with archive fallback; upstream strictly preserves the native URL; tos waits for a verified archive. Omitted uses the account/task default. Existing accounts retain legacy behavior; new accounts default to auto. Does not disable archiving. Idle enhanced models reject upstream. Combining callback_url with an explicit mode requires a configured webhook secret; otherwise creation returns 400 webhook_secret_required.
- **model** `string` **(required)**  
  Video generation model ID. Filter by `type: video` via `GET /v1/models`.
- **content** `object[]` **(required)**  
  Input content for the model, supporting text, images, video, and audio. Supported combinations:
- **content[].type** ``text` | `image_url` | `video_url` | `audio_url` | `draft_task`` **(required)**  
  Input content type:
- **content[].text** `string`  
  Text prompt (used when `type: text`). Supports Chinese and English; recommended max 500 Chinese characters
- **content[].image_url** `object`  
  Image object (used when `type: image_url`)
- **content[].role** ``first_frame` | `last_frame` | `reference_image` | `reference_video` | `reference_audio``  
  Role or purpose of the image/video/audio:
- **content[].video_url** `object`  
  Video object (used when `type: video_url`, Seedance 2.0 only)
- **content[].audio_url** `object`  
  Audio object (used when `type: audio_url`, Seedance 2.0 only)
- **content[].draft_task** `object`  
  Draft task reference (used when `type: draft_task`). `id` is the task ID returned when the draft was created.
- **resolution** ``480p` | `720p` | `1080p` | `768P` | `2K`` (default: `720p`)  
  Output video resolution.
- **ratio** ``16:9` | `4:3` | `1:1` | `3:4` | `9:16` | `21:9` | `adaptive`` (default: `adaptive`)  
  Output video aspect ratio.
- **duration** `integer` (default: `5`)  
  Output video duration (seconds), integer. Set to `-1` to let the model choose an appropriate duration (note: duration affects billing).
- **seed** `integer` (default: `-1`)  
  Random seed for controlling generation randomness. Range: [-1, 2^32-1].
- **generate_audio** `boolean` (default: `true`)  
  Whether to generate audio-enabled video. The model generates matching voice, sound effects, and background music based on the prompt and visual content.
- **draft** `boolean` (default: `false`)  
  Generate a draft preview. Seedance 2.5 only. Idle models do not support it.
- **return_last_frame** `boolean` (default: `false`)  
  Whether to return the last frame of the video as an image (PNG format, no watermark, same resolution as the video).
- **camera_fixed** `boolean` (default: `false`)  
  Whether to fix the camera position.
- **watermark** `boolean` (default: `false`)  
  Whether the generated video includes a watermark
- **service_tier** ``default` | `flex`` (default: `default`)  
  Service tier (cannot be changed for submitted tasks):
- **callback_url** `string`  
  Callback URL for terminal task states. When a task reaches `succeeded` / `failed`, XRToken sends a `POST` request to this URL.
- **safety_identifier** `string`  
  End-user unique identifier for safety auditing. It is recommended to pass a hash of the user ID, up to 64 characters.

### Response

- **id** `string` **(required)**  
  Platform internal task ID (used for status polling)
- **upstream_id** `string`  
  Upstream provider task ID (for reference only)
- **model** `string` **(required)**  
  Model ID used
- **status** ``queued`` **(required)**  
  Initial task status, fixed value `queued`
- **created_at** `string` **(required)**  
  Task creation time (ISO 8601 format)

### Error Codes

- `400`: 
- `401`: 
- `402`: 
- `429`: 
- `502`:
