# 语音合成 HTTP 单向流（火山原生协议透传）

**火山原生协议透传，非 OpenAI 兼容格式**：请求体与响应体均为火山语音合成 `bigmodel`
HTTP chunked 单向流原生格式，网关只做鉴权重写与计费旁路，不改写内容。适合已有火山
SDK 集成、想直接切换域名接入的客户；OpenAI 兼容格式见 `POST /v1/audio/speech`。

**鉴权**（二选一，均为 `tr-` 前缀 key）：
- `Authorization: Bearer <tr-key>`
- `X-Api-Key: <tr-key>`（火山 SDK 原生头，零改造接入）

**选模型**：`?model=<模型ID>` 优先；不传则回退到请求头 `X-Api-Resource-Id`
（火山原生客户端天然携带）。

**响应**：`Content-Type: application/json`，body 是 chunked、换行分隔的 JSON 对象流
（非单个 JSON），每个对象可能含：
- `code`：状态码，`20000000` 表示本次合成结束
- `data`：base64 编码的音频分片（增量）
- `sentence`：分句 / 字幕信息（`audio_params.enable_subtitle=true` 时返回字级时间戳）
- `usage.text_words`：本次合成计费字符数（结算权威口径，以最后一次出现的值为准）

**计费**：按文本字符数计费（含标点）；`req_params.additions.context_texts` 中的语音
指令文本不计费；结算以上游最终返回的 `usage.text_words` 为准，网关侧字符数估算仅用于
请求前预授权与上游未回传时的兜底。

## POST /v1/audio/speech/unidirectional

> TTS HTTP unidirectional stream (Volcengine native protocol passthrough)

**Native Volcengine protocol passthrough, not OpenAI-compatible**: both the
request body and response body use Volcengine's `bigmodel` HTTP chunked
unidirectional streaming native format. The gateway only rewrites
authentication and observes billing usage on the side — request/response
content is forwarded unchanged. Intended for clients already integrated with
the Volcengine SDK who just want to swap the domain; for the OpenAI-compatible
shape, use `POST /v1/audio/speech`.

**Auth** (either works, both are `tr-` prefixed keys):
- `Authorization: Bearer <tr-key>`
- `X-Api-Key: <tr-key>` (Volcengine SDK's native header, zero code changes)

**Model selection**: `?model=<model id>` takes priority; falls back to the
`X-Api-Resource-Id` request header (naturally sent by native Volcengine
clients) if omitted.

**Response**: `Content-Type: application/json`, body is a chunked,
newline-delimited stream of JSON objects (not a single JSON document). Each
object may contain:
- `code`: status code, `20000000` marks the end of this synthesis
- `data`: base64-encoded audio chunk (incremental)
- `sentence`: sentence/subtitle info (returns word-level timestamps when
  `audio_params.enable_subtitle=true`)
- `usage.text_words`: the character count billed for this synthesis
  (authoritative settlement figure — use the last value seen)

**Billing**: charged per text character (punctuation included);
instruction text under `req_params.additions.context_texts` is not billed;
settlement uses the final `usage.text_words` returned upstream — the
gateway's own character-count estimate is only used for pre-authorization
and as a fallback if upstream never returns usage.

### Authentication

`Authorization: Bearer tr-xxx`

### Query Parameters

- **model** `string`  
  TTS model / channel ID; either this or `X-Api-Resource-Id` — query takes priority

### Request Body

Content-Type: `application/json`

- **user** `object`  
  Optional caller-defined user identifier, forwarded upstream unchanged
- **user.uid** `string`  
  
- **req_params** `object` **(required)**  
  
- **req_params.text** `string` **(required)**  
  Text to synthesize (required); billed per character count (punctuation included)
- **req_params.model** `string` (default: `seed-tts-2.0-standard`)  
  Model ID used for voice cloning scenarios, defaults to `seed-tts-2.0-standard`
- **req_params.speaker** `string` **(required)**  
  Voice ID (required)
- **req_params.ssml** `string`  
  SSML-marked text, used instead of `text` — see Volcengine docs for the rules
- **req_params.audio_params** `object`  
  
- **req_params.additions** `object`  
  
- **req_params.section_id** `string`  
  Segment ID that preserves semantic continuity across requests (used for long-text segmented synthesis)
- **req_params.tone_fidelity** `boolean`  
  Tone fidelity toggle, only supported on voice-cloning 2.0 models

### Response

### Error Codes

- `400`: 
- `401`: 
- `402`: 
- `429`: 
- `502`:
