# Speech-to-Text (STT)

Transcribe audio files to text. The request uses `multipart/form-data` format and must include `file` (16k/mono/16bit WAV) and `model` fields.

**Response format**: On success, returns a Volcengine-shaped JSON response with the following fields:
- `result.text` — Recognized text content
- `audio_info.duration` — Recognized audio duration in **milliseconds**, used for settlement
- `request_id` — Gateway request ID (`tr-req-` prefix)

**Billing**: Charged per audio duration. Settlement uses `audio_info.duration` (milliseconds) to calculate the actual cost.

## POST /v1/audio/transcriptions

> Speech-to-Text (STT)

Transcribe audio files to text. The request uses `multipart/form-data` format and must include `file` (16k/mono/16bit WAV) and `model` fields.

**Response format**: On success, returns a Volcengine-shaped JSON response with the following fields:
- `result.text` — Recognized text content
- `audio_info.duration` — Recognized audio duration in **milliseconds**, used for settlement
- `request_id` — Gateway request ID (`tr-req-` prefix)

**Billing**: Charged per audio duration. Settlement uses `audio_info.duration` (milliseconds) to calculate the actual cost.

### Authentication

`Authorization: Bearer tr-xxx`

### Request Body

Content-Type: `multipart/form-data`

- **model** `string` **(required)**  
  STT model ID. Filter by `type: stt` via `GET /v1/models`
- **file** `string` **(required)**  
  Audio file. Supported formats: mp3, mp4, mpeg, mpga, m4a, wav, webm
- **language** `string`  
  Optional: Audio language code (BCP-47 format, e.g. `zh-CN`, `en-US`). Specifying this can improve recognition accuracy.
- **prompt** `string`  
  Prompt text to provide context and improve recognition accuracy
- **response_format** ``json` | `text` | `srt` | `verbose_json` | `vtt`` (default: `json`)  
  Response format

### Response

- **result** `object` **(required)**  
  
- **result.text** `string` **(required)**  
  Transcribed text content
- **audio_info** `object` **(required)**  
  
- **audio_info.duration** `integer` **(required)**  
  Recognized audio duration in milliseconds
- **request_id** `string` **(required)**  
  Gateway request ID

### Error Codes

- `400`: 
- `401`: 
- `402`: 
- `429`: 
- `502`:
