Skip to content

Alania · TTS

Parameters reference

Every field of the POST /v1/audio/speech body, with its default and range. Send JSON or multipart/form-data.

POSThttps://voice.patientdesk.ai/v1/audio/speechTry it

Request body

inputstringrequired
The Turkish text to speak, up to 5,000 characters. Numbers, dates, times and amounts are read as spoken. Long text is split at sentence boundaries and spoken in order. On Alania-2 a leading (…) is read as the instruction.
modelstring
alania-v2 or alania-v1. The OpenAI names tts-1, tts-1-hd and gpt-4o-mini-tts reach the default model. One exception: a request with no model and voice: "alania-v1" reaches Alania-1.
Default: alania-v2
voicestring
Alania-2: alania2-f1 (female), alania2-m1 (male), or the id of a saved voice. Alania-1: alania. Ignored when voice_description or reference_audio is sent.
Default: alania2-f1
instructionsstringalania-v2
How to speak, in English, at most 200 characters, no parentheses. See Style instructions.
voice_descriptionstringalania-v2
An English description of a new voice, at most 200 characters. Preview. Cannot be combined with instructions or reference_audio. See Voice design.
reference_audiofile | base64alania-v2
The voice to clone: 3–30 seconds (20–30 is best), at most 10 MB. Base64 or a data URI in JSON, a file in multipart. Needs consent: true. See Voice cloning.
reference_textstringalania-v2
The transcript of reference_audio, at most 1,000 characters. Optional and not used: the clone is made from the audio alone.
consentbooleanalania-v2
Required with reference_audio: states that the speaker agreed.
pronunciationsobject | array
How to say particular words in this request: {"PatientDesk": "Peyşınt Desk"} or [{"term": …, "say_as": …}], at most 200 entries. Takes precedence over your saved pronunciation dictionary. See Pronunciation dictionary.
response_format"wav" | "pcm" | "mp3" | "opus" | "flac" | "aac"
wav and pcm are 16-bit mono. Streams accept only wav and pcm.
Default: wav
streamboolean
When true, audio is sent as it is generated. See Streaming.
Default: false
seedinteger
0 to 2147483647. Random when omitted. The same seed, text and settings produce the same audio.
temperaturenumber
0.1–2.0 on Alania-2, 0–2.0 on Alania-1. Higher values vary the delivery and make the voice less consistent.
Default: 1.0 · 0.30 on Alania-1
cfgnumberalania-v2
1.0–5.0. Guidance strength: higher values stay closer to the voice, instruction or description. Leave the default unless you are tuning.
Default: 2.0 · 3.0 for design
top_pnumberalania-v1
0.01–1.0. Alania-2 ignores it.
Default: 0.9

Output

modelSample rateFormat
alania-v248 kHzMono, PCM16 (wav, pcm) or compressed
alania-v124 kHzMono, PCM16 (wav, pcm) or compressed
response_formatContent-Type
wavaudio/wav
pcmaudio/L16; rate=48000 (24000)
mp3audio/mpeg
opusaudio/ogg
flacaudio/flac
aacaudio/aac

Response headers

x-request-idstring
The request id; include it in support requests.
x-modelstring
The model that produced the audio: alania-v2 or alania-v1.
x-disclosure"ai-generated"
On every response. Telling listeners that the audio is AI-generated is your responsibility.
x-charactersinteger
Characters this request billed to your account.
x-audio-secondsnumber
Length of the audio produced. Not sent on streams.
x-queue-wait-msinteger
Time the request waited for a generation slot. The first thing to check when a caller says the API is slow.

Billing

Alania is billed per character, at the same price in every mode (built-in voice, style, design, cloning). What counts is input as you sent it; expanding numbers into words does not add to it. instructions, voice_description and reference_text are not counted; a leading (…) in input is.