Skip to content

Alania · TTS

Alania: overview

Alania reads Turkish text aloud. Alania-2 is the default model: two built-in voices, short instructions for how to speak, new voices from a description or a recording, and 48 kHz audio.

POSThttps://voice.patientdesk.ai/v1/audio/speechTry it
curl https://voice.patientdesk.ai/v1/audio/speech \
  -H "Authorization: Bearer pd_live_..." \
  -H "Content-Type: application/json" \
  -d '{ "model": "alania-v2", "input": "Merhaba, randevunuz onaylandı.", "voice": "alania2-f1" }' \
  --output merhaba.wav   # audio/wav, 48 kHz mono

The body is the same as OpenAI audio/speech, so your OpenAI client works once you change base_url. Alania-2's extra fields go in the same body. See the parameters reference for every field.

Four ways to choose a voice

WayFieldsDetails
Built-in voicevoiceOne of the two voices below. The same timbre on every request.
Style instructions (preview)voice + instructionsStyle instructions: slowly, calmly, whispering…
Voice design (preview)voice_descriptionVoice design: a new voice described in words.
Voice cloning (beta)reference_audio + consent, or voice set to a saved voice's idVoice cloning: from a 3–30 second recording (20–30 is best) you have permission to use, sent with each request or saved once and used by its id.

The price is the same for all four: per character of input.

Voices

voiceVoiceNote
alania2-f1FemaleAlania-2's default voice.
alania2-m1Male
alaniaMale · Alania-1Only with model: "alania-v1". Its id is alania-v1.
GEThttps://voice.patientdesk.ai/v1/voices
Response
{
  "object": "list",
  "data": [
    {
      "id": "alania-v1", "name": "Alania", "gender": "male", "language": "tr-TR",
      "model": "alania-v1", "sample_rate": 24000, "default": false,
      "f0_band_hz": [125.1, 137.5]
    },
    {
      "id": "alania2-f1", "name": "Alania F1", "gender": "female", "language": "tr-TR",
      "model": "alania-v2", "sample_rate": 48000, "default": true,
      "f0_band_hz": [199.5538, 238.605]
    },
    {
      "id": "alania2-m1", "name": "Alania M1", "gender": "male", "language": "tr-TR",
      "model": "alania-v2", "sample_rate": 48000, "default": false,
      "f0_band_hz": [109.0961, 131.6118]
    }
  ]
}

The list holds the built-in voices and the voices saved to your account (kind: "custom"); see Saved voices. Designed voices and recordings sent with a request are not saved or listed.

Models

GEThttps://voice.patientdesk.ai/v1/alania/models
Response
{
  "object": "list",
  "data": [
    { "id": "alania-v1", "object": "model", "owned_by": "patientdesk-ai",
      "default": false, "legacy": true, "sample_rate": 24000, "status": "ok" },
    { "id": "alania-v2", "object": "model", "owned_by": "patientdesk-ai",
      "default": true, "legacy": false, "sample_rate": 48000, "status": "ok" }
  ]
}

One host serves both products, so the endpoints that exist on both carry /alania on the speech side: /v1/alania/models, /v1/alania/usage, /health/alania. The audio endpoints and /v1/voices stay as they are.

Response

HTTP
HTTP/1.1 200 OK
content-type: audio/wav
x-request-id: 3f9c2b1e7a…
x-model: alania-v2
x-disclosure: ai-generated
x-characters: 45
x-audio-seconds: 3.18
x-queue-wait-ms: 0

<RIFF header, then PCM16 48 kHz mono>

Every response carries x-disclosure: ai-generated. Telling your listeners that the audio is AI-generated is your responsibility. All headers are in the parameters reference.

Measurements

MeasureValueConditions
First audio~190 msStreamed request to first audio chunk, median, measured through the public API with the network included.
Character error rate · alania2-f12.7%Our 292-sentence Turkish test set; the audio was transcribed with Whisper.
Character error rate · alania2-m12.8%Same test set and method.