Speech-to-Text (STT)
Transcribe audio files to text. The request uses multipart/form-data format and must include file (16k/mono/16bit WAV) and model fields.
Response format: On success, returns a Volcengine-shaped JSON response with the following fields:
result.text— Recognized text contentaudio_info.duration— Recognized audio duration in milliseconds, used for settlementrequest_id— Gateway request ID (tr-req-prefix)
Billing: Charged per audio duration. Settlement uses audio_info.duration (milliseconds) to calculate the actual cost.
Authorization
BearerAuth API key authentication (OpenAI format). Pass in the Authorization header:
Authorization: Bearer tr-xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
In: header
Request Body
multipart/form-data
TypeScript Definitions
Use the request body type in TypeScript.
Response Body
application/json
application/json
application/json
application/json
application/json
application/json
curl -X post "https://api.xrtoken.ai/v1/audio/transcriptions" \ -F model="openai/whisper-1" \ -F file="string"{
"audio_info": {
"duration": 1200
},
"request_id": "tr-req-a1b2c3d4e5f6",
"result": {
"text": "Hello, welcome to XRToken."
}
}{
"error": "model field is required",
"type": "invalid_request_error"
}{
"error": "invalid or missing API key",
"type": "auth_error"
}{
"error": "insufficient balance -- please top up or upgrade your plan",
"type": "billing_error"
}{
"error": "rate limit exceeded",
"type": "rate_limit_error"
}{
"error": "upstream provider error",
"type": "server_error"
}