# MiniMax H3 request parameters

Same endpoints as other video models:

- China create: `POST https://api.xrtoken.net/v1/videos/generations`
- International create: `POST https://api.xrtoken.ai/v1/videos/generations`
- Poll: `GET {same host}/v1/videos/generations/{id}`

`MiniMax-H3` is available on both the China and international sites. Billing is per generated second. The first 5 reference images are free; extras are billed per image. A reference video is billed at the same per-second rate for input-video seconds. China uses MiniMax China list prices; international uses the USD list.

| Site | 768P | 2K | Extra reference image |
|---|---|---|---|
| China | ¥0.50 / s | ¥0.80 / s | ¥0.20 / image |
| International | $0.08 / s | $0.13 / s | $0.04 / image |

## MiniMax H3 Max / Max-Turbo (China and international)

Model IDs: `MiniMax-H3-Max` and `MiniMax-H3-Max-Turbo`. Both support text, first-frame and first/last-frame generation. Max also supports image/video/audio references; Turbo does not.

| Output model / resolution | China / second | International / second |
| --- | --- | --- |
| Max 480P | ¥0.33 | $0.05 |
| Max 768P | ¥0.50 | $0.08 |
| Max-Turbo 480P | ¥0.17 | $0.025 |
| Max-Turbo 768P | ¥0.26 | $0.04 |

Max reference video costs ¥0.37 / $0.0553 per input second at 480P, or ¥0.97 / $0.143 at 768P. The first two images are free; each additional image costs ¥0.50 / $0.074. Reference audio has no separate charge. Turbo charges output only.

Max total = output seconds × output rate + input-video seconds × reference rate + max(image count − 2, 0) × image rate. Use `usage.input_seconds`; do not multiply `total_seconds` by the output rate.

Rules specific to Max / Max-Turbo:

- Integer duration **5–15 seconds**; resolution **480P / 768P**. XRToken defaults to 5 seconds / 768P when omitted.
- Exactly one non-empty text item in `content`. Media URLs must use nested objects, such as `image_url: { "url": "..." }`.
- At most one first frame and one last frame. A last frame requires a first frame. Frames cannot mix with references.
- Max: up to 9 images, 3 videos, 3 audios and 12 references total. Audio-only references are unsupported. Each video/audio is 2–15 seconds; total video and total audio duration must each be at most 15 seconds, validated upstream after download.
- Use a concrete `ratio` for text generation; `adaptive` is recommended for frames and supported for Max references.
- Optional `extra: { "prompt_expansion_mode": "disabled" | "balanced" | "quality" }`, default `balanced`, and integer `seed`.
- Use the XRToken endpoints above. `callback_url` follows XRToken's callback mechanism.

Example: Max 768P, 5 output seconds, 3 images, 10 reference-video seconds and 10 audio seconds costs ¥12.70 / $1.904.

## Spec

| Field | Required | Allowed values |
|---|---|---|
| `model` | yes | `MiniMax-H3` |
| `content` | yes | Array. Must include a non-empty `{ "type": "text", "text": "..." }` |
| `duration` | yes | Integer `4`–`15` |
| `resolution` | yes | `768P` or `2K` |
| `ratio` | required for text-to-video | One of `16:9`, `4:3`, `1:1`, `3:4`, `9:16`, `21:9` |

Media items in `content` must use these `type` + `role` pairs:

| type | role | Meaning |
|---|---|---|
| `image_url` | `first_frame` | First frame (image-to-video) |
| `image_url` | `last_frame` | Last frame (image-to-video) |
| `image_url` | `reference_image` | Reference image |
| `video_url` | `reference_video` | Reference video |
| `audio_url` | `reference_audio` | Reference audio |

Rules:

- Media URLs must be public `https`. This model does not use `asset://`.
- Images + videos + audios combined ≤ 12.
- Use one mode only: **text-to-video**, **first/last frame**, or **reference-to-video**. `first_frame` / `last_frame` must not appear with any `reference_*`.
- `duration` must be an integer. Do not send `-1` or a fraction.
- `resolution` accepts only `768P` and `2K`, with this casing.
- Do not send `ratio: "adaptive"` for text-to-video.

This model does not support: `generate_audio`, `camera_fixed`, `watermark`, `fps`, `seed`, `service_tier`.

## Examples

### Text-to-video

```bash
curl -sS https://api.xrtoken.ai/v1/videos/generations \
  -H "Authorization: Bearer $XRT_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "MiniMax-H3",
    "content": [
      { "type": "text", "text": "An orange cat running on a sunny lawn, cinematic slow motion" }
    ],
    "duration": 5,
    "resolution": "2K",
    "ratio": "16:9"
  }'
```

### Image-to-video (first frame)

```json
{
  "model": "MiniMax-H3",
  "content": [
    { "type": "text", "text": "The apple rolls gently in the sunlight" },
    {
      "type": "image_url",
      "image_url": { "url": "https://example.com/apple.jpg" },
      "role": "first_frame"
    }
  ],
  "duration": 6,
  "resolution": "768P",
  "ratio": "16:9"
}
```

Add another image with `"role": "last_frame"` for first-and-last-frame.

### Reference-to-video

```json
{
  "model": "MiniMax-H3",
  "content": [
    { "type": "text", "text": "Follow the camera move of the reference video; use the audio as BGM" },
    {
      "type": "image_url",
      "image_url": { "url": "https://example.com/style.jpg" },
      "role": "reference_image"
    },
    {
      "type": "video_url",
      "video_url": { "url": "https://example.com/cam.mp4" },
      "role": "reference_video"
    },
    {
      "type": "audio_url",
      "audio_url": { "url": "https://example.com/bgm.mp3" },
      "role": "reference_audio"
    }
  ],
  "duration": 8,
  "resolution": "2K",
  "ratio": "16:9"
}
```

## Poll

Create returns a task `id`. Poll with:

```bash
curl -sS "https://api.xrtoken.ai/v1/videos/generations/$TASK_ID" \
  -H "Authorization: Bearer $XRT_API_KEY"
```

Status is `queued` / `processing` / `succeeded` / `failed`. On success the clip is at `content.video_url`.

See [Create video generation](/docs/api/createVideoGeneration) for the full schema.
