Skip to content
All posts

· 6 min read · By PatientDesk AI

Alania-2: our new Turkish voice, built from scratch

Today we are making Alania-2 the default voice of the PatientDesk speech API. It is a 488-million-parameter text-to-speech model we trained from scratch for Turkish. It speaks at 48 kHz, comes with two built-in voices, and can follow a style instruction, design a new voice from a description, or clone a voice with the speaker's consent. Alania-1 stays available for everyone already using it.

How we built it

Alania-2 is our own model, not a fine-tune of someone else's weights. We used the open VoxCPM2 architecture: a language model that predicts continuous audio representations instead of discrete audio tokens, decoded to 48 kHz by an audio autoencoder. We started from random weights and trained a 488M-parameter version, about a fifth the size of the largest model we tried.

The hard part was data. Openly licensed Turkish speech is scarce: the largest public sets are a few hundred hours, and our earlier attempts to train from scratch on around 400 hours of real speech did not learn to align text and audio. So we built a corpus. We wrote 1.44 million Turkish sentences of the kind a voice agent actually says (appointments, banking, cargo, numbers, codes, addresses, names), added Turkish web text and Wikipedia, and generated speech for them with the open VoxCPM2 model in thousands of different voices. Every line went through our Turkish normaliser first, and every clip through automatic checks for intelligibility, quality and voice consistency.

The main training run covered about 18,000 hours of Turkish speech. A final fine-tune taught the model style instructions and voice design, from speech paired with plain-language descriptions of how it sounds: pitch, pace, pauses, noise and room.

The two built-in voices were not recorded by actors. We designed candidates, picked one female and one male voice in blind listening, and froze each one's reference and pitch band so every request sounds like the same person.

What changed from Alania-1

Alania-2 doubles the sample rate, roughly halves the time to first audio, and adds three ways to choose a voice that Alania-1 never had.

Alania-2Alania-1
Sample rate48 kHz24 kHz
Built-in voices2 (alania2-f1, alania2-m1)1 (alania)
Style instructionsYes (preview)No
Voice design from a descriptionYes (preview)No
Voice cloningYes, with consentNo
First audio, streamed, public API (median)~190 ms~365 ms
Default modelYesNo, still available

On our 292-sentence Turkish test set, transcribed back with Whisper large-v3, the built-in voices have a character error rate of 2.7 % (female) and 2.8 % (male); with a style instruction it is 3.4 %. Latency was measured on 5 October 2026 from our office through voice.patientdesk.ai, so it includes the network; your number depends on where you are.

Four ways to choose a voice

Every request picks its voice one of four ways, all on the same endpoint, POST /v1/audio/speech. Existing OpenAI clients keep working; the new fields go in the request body.

  1. Built-in voices. alania2-f1 and alania2-m1: the same timbre on every request, held to a fixed pitch band.
  2. Style instructions (preview). Add instructions in English, such as “speaking slowly and calmly”. The voice stays; the delivery changes. In our tests the model did not always follow: a request for “slowly” came out at normal pace, and “whispering” came out voiced. That is why it is labelled preview.
  3. Voice design (preview). Describe a voice in voice_description, such as “a young man with a clear, energetic voice”, and Alania-2 creates it. For long text, the first sentence designs the voice and the rest of the text reuses it, so one request keeps one speaker. Quality varies more than with the built-in voices.
  4. Voice cloning, with consent. Send 3 to 30 seconds of reference_audio, ideally with its reference_text, and Alania-2 speaks in that voice. Every cloning request must carry consent: true, confirming the speaker agreed. Clone only voices you have the right to use.

Every response carries an x-disclosure header marking the audio as AI-generated, whichever way the voice was chosen.

Reading Turkish the way Turks read it

Most mistakes a Turkish listener notices are not in the voice but in the reading: a number, an amount, an abbreviation. Before any text reaches the model, our normaliser writes it out the way it is spoken, and native speakers reviewed every reading rule.

WrittenRead as
12.480,75 TLon iki bin dört yüz seksen lira yetmiş beş kuruş
2,5 mgiki buçuk miligram
%15 indirimyüzde on beş indirim
TBMM, SGK, KPSS, NTVtebememe, segeka, kapesese, entivi
H2Ohaş iki o
HD kaliteheyçdi kalite

Turkish initialisms are read with Turkish letter names (K is “ka”, H is “he”), and English ones such as HD keep their English names. Alania-2 also receives text with its capitals kept, so it can tell proper names and acronyms apart. When native review changed a reading, we replaced the training speech that used the old one, so the model never learned it.

Serving it fast

Alania-2 streams: the first audio leaves the server while the rest of the sentence is still being generated, as 48 kHz PCM inside a WAV stream or over a WebSocket. One engine step produces 0.16 seconds of audio, so playback can start almost as soon as the request arrives.

Every speech server now runs both models on one GPU: Alania-2 in its own worker process, Alania-1 next to it for existing customers. NVIDIA MPS shares the card between them; without it, Alania-1 stalled as soon as Alania-2 got busy. A server only takes traffic once Alania-2 is loaded, so a request without a model never lands on a machine that cannot answer it.

MeasurementResult
First audio, streamed, through the public API (median)~190 ms
First audio at the server, one stream, H200 GPU~50 ms
Sustained throughput, one H200 GPU~12 requests/s
Throughput keeping first audio under 300 ms (p95), one A100 GPU~8 requests/s

The gap between 50 ms at the server and 190 ms for a caller is the network and the load balancer, which is why we quote the larger number on our site.

What we are opening up

Alania-2's weights stay closed, but most of what we built to get there is public, so others can build Turkish speech technology without starting from nothing. Everything below can be used commercially with attribution.

ReleaseWhat it isSizeLicence
Alania Turkish Synthetic Speech48 kHz Turkish speech from 2,752 designed voices, each clip with its text, spoken form and a voice description; some with simulated rooms and microphones. No real person's voice.3,411 h, 2,037,461 clipsCC BY 4.0 (Wikipedia part CC BY-SA 4.0)
Alania Turkish Domain TextSentences a voice agent says: appointments, banking, cargo, support, numbers, codes, addresses, names, with normalised spoken forms.1,439,639 sentencesCC BY 4.0
Alania Turkish Speech Style CaptionsReal Turkish speech segments described in English and Turkish (pitch, pace, pauses, noise, room), with measurements and transcripts; audio fetched from YODAS v3. The first such set for Turkish.835,267 segments, 2,398 hCC BY 4.0 (captions)
Antalia-1Our earlier open Turkish TTS model (24 kHz), its foundation model, evaluation suite and technical report. Development has ended; released in full.~300M parametersSee repository
Antalia voice corpusStudio recordings of one professional voice actor who consented to open release.5.0 hCC BY 4.0
OpenVoiceCS-BenchA benchmark for voice customer-service agents, scored from replayable traces: task resolution, policy compliance, privacy, latency, cost.46 models on its v0.2 leaderboardApache 2.0 code, CC BY 4.0 data

If you train or evaluate a Turkish speech model, the style captions and domain text are the pieces we could not find anywhere else when we started.

What is next, and how to try it

Two features are labelled preview because they are not yet as reliable as the built-in voices: style instructions are followed loosely, and designed voices vary more in quality. Both are where we are working next, along with phone-line presets and more built-in voices.

You can hear Alania-2 now:

  • Type your own Turkish on speech.patientdesk.ai; no account needed.
  • Read the documentation for every parameter, streaming, and the move from Alania-1.
  • Already on Alania-1? Nothing breaks: send model: "alania-v1" to keep the old voice. Requests that name no model now get Alania-2.

Duyu, our speech-to-text model, works with the same key, so one account covers both directions of a conversation.