Alania · TTS
Parameters reference
Every field of the POST /v1/audio/speech body, with its default and range. Send JSON or multipart/form-data.
Request body
inputstringrequired- The Turkish text to speak, up to 5,000 characters. Numbers, dates, times and amounts are read as spoken. Long text is split at sentence boundaries and spoken in order. On Alania-2 a leading
(…)is read as the instruction. modelstringalania-v2oralania-v1. The OpenAI namestts-1,tts-1-hdandgpt-4o-mini-ttsreach the default model. One exception: a request with no model andvoice: "alania-v1"reaches Alania-1.- Default:
alania-v2 voicestring- Alania-2:
alania2-f1(female),alania2-m1(male), or the id of a saved voice. Alania-1:alania. Ignored whenvoice_descriptionorreference_audiois sent. - Default:
alania2-f1 instructionsstringalania-v2- How to speak, in English, at most 200 characters, no parentheses. See Style instructions.
voice_descriptionstringalania-v2- An English description of a new voice, at most 200 characters. Preview. Cannot be combined with
instructionsorreference_audio. See Voice design. reference_audiofile | base64alania-v2- The voice to clone: 3–30 seconds (20–30 is best), at most 10 MB. Base64 or a data URI in JSON, a file in multipart. Needs
consent: true. See Voice cloning. reference_textstringalania-v2- The transcript of
reference_audio, at most 1,000 characters. Optional and not used: the clone is made from the audio alone. consentbooleanalania-v2- Required with
reference_audio: states that the speaker agreed. pronunciationsobject | array- How to say particular words in this request:
{"PatientDesk": "Peyşınt Desk"}or[{"term": …, "say_as": …}], at most 200 entries. Takes precedence over your saved pronunciation dictionary. See Pronunciation dictionary. response_format"wav" | "pcm" | "mp3" | "opus" | "flac" | "aac"wavandpcmare 16-bit mono. Streams accept onlywavandpcm.- Default:
wav seedinteger- 0 to 2147483647. Random when omitted. The same seed, text and settings produce the same audio.
temperaturenumber- 0.1–2.0 on Alania-2, 0–2.0 on Alania-1. Higher values vary the delivery and make the voice less consistent.
- Default:
1.0· 0.30 on Alania-1 cfgnumberalania-v2- 1.0–5.0. Guidance strength: higher values stay closer to the voice, instruction or description. Leave the default unless you are tuning.
- Default:
2.0· 3.0 for design top_pnumberalania-v1- 0.01–1.0. Alania-2 ignores it.
- Default:
0.9
Output
| model | Sample rate | Format |
|---|---|---|
| alania-v2 | 48 kHz | Mono, PCM16 (wav, pcm) or compressed |
| alania-v1 | 24 kHz | Mono, PCM16 (wav, pcm) or compressed |
| response_format | Content-Type |
|---|---|
| wav | audio/wav |
| pcm | audio/L16; rate=48000 (24000) |
| mp3 | audio/mpeg |
| opus | audio/ogg |
| flac | audio/flac |
| aac | audio/aac |
Response headers
x-request-idstring- The request id; include it in support requests.
x-modelstring- The model that produced the audio:
alania-v2oralania-v1. x-disclosure"ai-generated"- On every response. Telling listeners that the audio is AI-generated is your responsibility.
x-charactersinteger- Characters this request billed to your account.
x-audio-secondsnumber- Length of the audio produced. Not sent on streams.
x-queue-wait-msinteger- Time the request waited for a generation slot. The first thing to check when a caller says the API is slow.
Billing
Alania is billed per character, at the same price in every mode (built-in voice, style, design, cloning). What counts is input as you sent it; expanding numbers into words does not add to it. instructions, voice_description and reference_text are not counted; a leading (…) in input is.