curl --request POST \
--url https://api.minimax.io/v1/t2a_v2 \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: <content-type>' \
--data '
{
"model": "speech-2.8-hd",
"text": "Omg(sighs), the real danger is not that computers start thinking like people, but that people start thinking like computers. Computers can only help us with simple tasks.",
"stream": false,
"language_boost": "auto",
"output_format": "hex",
"voice_setting": {
"voice_id": "English_expressive_narrator",
"speed": 1,
"vol": 1,
"pitch": 0
},
"pronunciation_dict": {
"tone": [
"Omg/Oh my god"
]
},
"audio_setting": {
"sample_rate": 32000,
"bitrate": 128000,
"format": "mp3",
"channel": 1
},
"voice_modify": {
"pitch": 0,
"intensity": 0,
"timbre": 0,
"sound_effects": "spacious_echo"
}
}
'{
"data": {
"audio": "<hex encoded audio>",
"status": 2
},
"extra_info": {
"audio_length": 11124,
"audio_sample_rate": 32000,
"audio_size": 179926,
"bitrate": 128000,
"word_count": 163,
"invisible_character_ratio": 0,
"usage_characters": 163,
"audio_format": "mp3",
"audio_channel": 1
},
"trace_id": "01b8bf9bb7433cc75c18eee6cfa8fe21",
"base_resp": {
"status_code": 0,
"status_msg": "success"
}
}Synchronous Text to Speech
Submit the full text in a single HTTP request and get the synthesized audio back, with optional streaming output.
curl --request POST \
--url https://api.minimax.io/v1/t2a_v2 \
--header 'Authorization: Bearer <token>' \
--header 'Content-Type: <content-type>' \
--data '
{
"model": "speech-2.8-hd",
"text": "Omg(sighs), the real danger is not that computers start thinking like people, but that people start thinking like computers. Computers can only help us with simple tasks.",
"stream": false,
"language_boost": "auto",
"output_format": "hex",
"voice_setting": {
"voice_id": "English_expressive_narrator",
"speed": 1,
"vol": 1,
"pitch": 0
},
"pronunciation_dict": {
"tone": [
"Omg/Oh my god"
]
},
"audio_setting": {
"sample_rate": 32000,
"bitrate": 128000,
"format": "mp3",
"channel": 1
},
"voice_modify": {
"pitch": 0,
"intensity": 0,
"timbre": 0,
"sound_effects": "spacious_echo"
}
}
'{
"data": {
"audio": "<hex encoded audio>",
"status": 2
},
"extra_info": {
"audio_length": 11124,
"audio_sample_rate": 32000,
"audio_size": 179926,
"bitrate": 128000,
"word_count": 163,
"invisible_character_ratio": 0,
"usage_characters": 163,
"audio_format": "mp3",
"audio_channel": 1
},
"trace_id": "01b8bf9bb7433cc75c18eee6cfa8fe21",
"base_resp": {
"status_code": 0,
"status_msg": "success"
}
}Authorizations
HTTP: Bearer Auth
- Security Scheme Type: http
- HTTP Authorization Scheme:
Bearer API_key, can be found in Account Management>API Keys.
Headers
The media type of the request body. Must be set to application/json to ensure the data is sent in JSON format.
application/json Body
The speech synthesis model version to use.
speech-2.8-hd, speech-2.8-turbo, speech-2.6-hd, speech-2.6-turbo, speech-02-hd, speech-02-turbo, speech-01-hd, speech-01-turbo The text to be converted into speech. Must be less than 10,000 characters.
-
For texts over 3,000 characters, streaming output is recommended.
-
Paragraph breaks should be marked with newline characters.
-
Pause control: You can customize speech pauses by adding markers in the form
<#x#>, wherexis the pause duration in seconds. Valid range:[0.01, 99.99], up to two decimal places. Pause markers must be placed between speakable text segments and cannot be used consecutively. -
Inline pronunciation: Wrap Mandarin Pinyin (with tone number
1–5) or IPA symbols or Cantonese Jyutping (with tone number1–6) in half-width parentheses to override pronunciation of the target word or polyphonic character."The word live is pronounced (lɪv) as a verb and (laɪv) as an adjective.""This is (he2)平, not (huo4)面.""去街市買啲(sung3)。"
-
Interjection tags: Only supported when using
speech-2.8-hdorspeech-2.8-turbomodels. Supported interjections:(laughs),(chuckle),(coughs),(clear-throat),(groans),(breath),(pant),(inhale),(exhale),(gasps),(sniffs),(sighs),(snorts),(burps),(lip-smacking),(humming),(hissing),(emm),(sneezes).
Whether to enable streaming output. Defaults to false.
Show child attributes
Show child attributes
Show child attributes
Show child attributes
Show child attributes
Show child attributes
Show child attributes
Show child attributes
Timbre weights (legacy field)
Show child attributes
Show child attributes
Controls whether recognition for specific minority languages and dialects is enhanced. Default is null. If the language type is unknown, set to "auto" and the model will automatically detect it.
Note: The speech-01 and speech-02 series models do not currently support Persian, Filipino, or Tamil.
Chinese, Chinese,Yue, English, Arabic, Russian, Spanish, French, Portuguese, German, Turkish, Dutch, Ukrainian, Vietnamese, Indonesian, Japanese, Italian, Korean, Thai, Polish, Romanian, Greek, Czech, Finnish, Hindi, Bulgarian, Danish, Hebrew, Malay, Persian, Slovak, Swedish, Croatian, Filipino, Hungarian, Norwegian, Slovenian, Catalan, Nynorsk, Tamil, Afrikaans, auto Voice effects configuration.
Supported audio formats:
- Non-streaming:
mp3,wav,flac - Streaming:
mp3
Show child attributes
Show child attributes
Controls whether subtitles are enabled. Default is false. Available for models: speech-2.8-hd, speech-2.8-turbo, speech-2.6-hd, speech-2.6-turbo, speech-02-hd, speech-02-turbo, speech-01-hd, speech-01-turbo.
Subtitle granularity. Default is sentence. Options:
sentence: sentence-level timestampsword: word-level timestampsword_streaming: word-level timestamps optimized for streaming, only valid whenstream=true
In streaming mode (stream=true), subtitles are delivered in data.subtitle along with the audio, and subtitle_file is not returned:
sentence/word: returned once per segment, on the last audio chunk of that segment.sentencedoes not includetimestamped_wordsword_streaming: returned on every audio chunk, carrying the cumulative subtitle of the current segment so far (timestamped_wordsandtime_endkeep growing). Replace the previous result of the same segment with the latest one; a change oftext_beginmeans a new segment has started
In non-streaming mode, subtitles are returned as a file via subtitle_file.
sentence, word, word_streaming Controls the output format. Options: [url, hex]. Default is hex. Only effective in non-streaming scenarios. In streaming, only hex is supported. Returned url is valid for 24 hours.
url, hex Response
The synthesized audio data object. The returned data object may be null, so a null check is required.
Show child attributes
Show child attributes
The session ID, used for troubleshooting and support.
Additional audio information.
Show child attributes
Show child attributes
Status code and details of this request.
Show child attributes
Show child attributes