MiniMax H3 request parameters
Parameter spec for MiniMax-H3 on POST /v1/videos/generations
Same endpoints as other video models:
- China create:
POST https://api.xrtoken.net/v1/videos/generations - International create:
POST https://api.xrtoken.ai/v1/videos/generations - Poll:
GET {same host}/v1/videos/generations/{id}
MiniMax-H3 is available on both the China and international sites. Billing is per generated second. The first 5 reference images are free; extras are billed per image. A reference video is billed at the same per-second rate for input-video seconds. China uses MiniMax China list prices; international uses the USD list.
| Site | 768P | 2K | Extra reference image |
|---|---|---|---|
| China | ¥0.50 / s | ¥0.80 / s | ¥0.20 / image |
| International | $0.08 / s | $0.13 / s | $0.04 / image |
MiniMax H3 Max / Max-Turbo (China and international)
Model IDs: MiniMax-H3-Max and MiniMax-H3-Max-Turbo. Both support text, first-frame and first/last-frame generation. Max also supports image/video/audio references; Turbo does not.
| Output model / resolution | China / second | International / second |
|---|---|---|
| Max 480P | ¥0.33 | $0.05 |
| Max 768P | ¥0.50 | $0.08 |
| Max-Turbo 480P | ¥0.17 | $0.025 |
| Max-Turbo 768P | ¥0.26 | $0.04 |
Max reference video costs ¥0.37 / $0.0553 per input second at 480P, or ¥0.97 / $0.143 at 768P. The first two images are free; each additional image costs ¥0.50 / $0.074. Reference audio has no separate charge. Turbo charges output only.
Max total = output seconds × output rate + input-video seconds × reference rate + max(image count − 2, 0) × image rate. Use usage.input_seconds; do not multiply total_seconds by the output rate.
Rules specific to Max / Max-Turbo:
- Integer duration 5–15 seconds; resolution 480P / 768P. XRToken defaults to 5 seconds / 768P when omitted.
- Exactly one non-empty text item in
content. Media URLs must use nested objects, such asimage_url: { "url": "..." }. - At most one first frame and one last frame. A last frame requires a first frame. Frames cannot mix with references.
- Max: up to 9 images, 3 videos, 3 audios and 12 references total. Audio-only references are unsupported. Each video/audio is 2–15 seconds; total video and total audio duration must each be at most 15 seconds, validated upstream after download.
- Use a concrete
ratiofor text generation;adaptiveis recommended for frames and supported for Max references. - Optional
extra: { "prompt_expansion_mode": "disabled" | "balanced" | "quality" }, defaultbalanced, and integerseed. - Use the XRToken endpoints above.
callback_urlfollows XRToken's callback mechanism.
Example: Max 768P, 5 output seconds, 3 images, 10 reference-video seconds and 10 audio seconds costs ¥12.70 / $1.904.
Spec
| Field | Required | Allowed values |
|---|---|---|
model | yes | MiniMax-H3 |
content | yes | Array. Must include a non-empty { "type": "text", "text": "..." } |
duration | yes | Integer 4–15 |
resolution | yes | 768P or 2K |
ratio | required for text-to-video | One of 16:9, 4:3, 1:1, 3:4, 9:16, 21:9 |
Media items in content must use these type + role pairs:
| type | role | Meaning |
|---|---|---|
image_url | first_frame | First frame (image-to-video) |
image_url | last_frame | Last frame (image-to-video) |
image_url | reference_image | Reference image |
video_url | reference_video | Reference video |
audio_url | reference_audio | Reference audio |
Rules:
- Media URLs must be public
https. This model does not useasset://. - Images + videos + audios combined ≤ 12.
- Use one mode only: text-to-video, first/last frame, or reference-to-video.
first_frame/last_framemust not appear with anyreference_*. durationmust be an integer. Do not send-1or a fraction.resolutionaccepts only768Pand2K, with this casing.- Do not send
ratio: "adaptive"for text-to-video.
This model does not support: generate_audio, camera_fixed, watermark, fps, seed, service_tier.
Examples
Text-to-video
curl -sS https://api.xrtoken.ai/v1/videos/generations \
-H "Authorization: Bearer $XRT_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "MiniMax-H3",
"content": [
{ "type": "text", "text": "An orange cat running on a sunny lawn, cinematic slow motion" }
],
"duration": 5,
"resolution": "2K",
"ratio": "16:9"
}'
Image-to-video (first frame)
{
"model": "MiniMax-H3",
"content": [
{ "type": "text", "text": "The apple rolls gently in the sunlight" },
{
"type": "image_url",
"image_url": { "url": "https://example.com/apple.jpg" },
"role": "first_frame"
}
],
"duration": 6,
"resolution": "768P",
"ratio": "16:9"
}Add another image with "role": "last_frame" for first-and-last-frame.
Reference-to-video
{
"model": "MiniMax-H3",
"content": [
{ "type": "text", "text": "Follow the camera move of the reference video; use the audio as BGM" },
{
"type": "image_url",
"image_url": { "url": "https://example.com/style.jpg" },
"role": "reference_image"
},
{
"type": "video_url",
"video_url": { "url": "https://example.com/cam.mp4" },
"role": "reference_video"
},
{
"type": "audio_url",
"audio_url": { "url": "https://example.com/bgm.mp3" },
"role": "reference_audio"
}
],
"duration": 8,
"resolution": "2K",
"ratio": "16:9"
}Poll
Create returns a task id. Poll with:
curl -sS "https://api.xrtoken.ai/v1/videos/generations/$TASK_ID" \
-H "Authorization: Bearer $XRT_API_KEY"
Status is queued / processing / succeeded / failed. On success the clip is at content.video_url.
See Create video generation for the full schema.