How we built it
Alania-2 is our own model, not a fine-tune of someone else's weights. We used the open VoxCPM2 architecture: a language model that predicts continuous audio representations instead of discrete audio tokens, decoded to 48 kHz by an audio autoencoder. We started from random weights and trained a 488M-parameter version, about a fifth the size of the largest model we tried.
The hard part was data. Openly licensed Turkish speech is scarce: the largest public sets are a few hundred hours, and our earlier attempts to train from scratch on around 400 hours of real speech did not learn to align text and audio. So we built a corpus. We wrote 1.44 million Turkish sentences of the kind a voice agent actually says (appointments, banking, cargo, numbers, codes, addresses, names), added Turkish web text and Wikipedia, and generated speech for them with the open VoxCPM2 model in thousands of different voices. Every line went through our Turkish normaliser first, and every clip through automatic checks for intelligibility, quality and voice consistency.
The main training run covered about 18,000 hours of Turkish speech. A final fine-tune taught the model style instructions and voice design, from speech paired with plain-language descriptions of how it sounds: pitch, pace, pauses, noise and room.
The two built-in voices were not recorded by actors. We designed candidates, picked one female and one male voice in blind listening, and froze each one's reference and pitch band so every request sounds like the same person.
What changed from Alania-1
Alania-2 doubles the sample rate, roughly halves the time to first audio, and adds three ways to choose a voice that Alania-1 never had.
| Alania-2 | Alania-1 | |
|---|---|---|
| Sample rate | 48 kHz | 24 kHz |
| Built-in voices | 2 (alania2-f1, alania2-m1) | 1 (alania) |
| Style instructions | Yes (preview) | No |
| Voice design from a description | Yes (preview) | No |
| Voice cloning | Yes, with consent | No |
| First audio, streamed, public API (median) | ~190 ms | ~365 ms |
| Default model | Yes | No, still available |
On our 292-sentence Turkish test set, transcribed back with Whisper large-v3, the built-in voices have a character error rate of 2.7 % (female) and 2.8 % (male); with a style instruction it is 3.4 %. Latency was measured on 5 October 2026 from our office through voice.patientdesk.ai, so it includes the network; your number depends on where you are.
Four ways to choose a voice
Every request picks its voice one of four ways, all on the same endpoint, POST /v1/audio/speech. Existing OpenAI clients keep working; the new fields go in the request body.
- Built-in voices.
alania2-f1andalania2-m1: the same timbre on every request, held to a fixed pitch band. - Style instructions (preview). Add
instructionsin English, such as “speaking slowly and calmly”. The voice stays; the delivery changes. In our tests the model did not always follow: a request for “slowly” came out at normal pace, and “whispering” came out voiced. That is why it is labelled preview. - Voice design (preview). Describe a voice in
voice_description, such as “a young man with a clear, energetic voice”, and Alania-2 creates it. For long text, the first sentence designs the voice and the rest of the text reuses it, so one request keeps one speaker. Quality varies more than with the built-in voices. - Voice cloning, with consent. Send 3 to 30 seconds of
reference_audio, ideally with itsreference_text, and Alania-2 speaks in that voice. Every cloning request must carryconsent: true, confirming the speaker agreed. Clone only voices you have the right to use.
Every response carries an x-disclosure header marking the audio as AI-generated, whichever way the voice was chosen.
Reading Turkish the way Turks read it
Most mistakes a Turkish listener notices are not in the voice but in the reading: a number, an amount, an abbreviation. Before any text reaches the model, our normaliser writes it out the way it is spoken, and native speakers reviewed every reading rule.
| Written | Read as |
|---|---|
| 12.480,75 TL | on iki bin dört yüz seksen lira yetmiş beş kuruş |
| 2,5 mg | iki buçuk miligram |
| %15 indirim | yüzde on beş indirim |
| TBMM, SGK, KPSS, NTV | tebememe, segeka, kapesese, entivi |
| H2O | haş iki o |
| HD kalite | heyçdi kalite |
Turkish initialisms are read with Turkish letter names (K is “ka”, H is “he”), and English ones such as HD keep their English names. Alania-2 also receives text with its capitals kept, so it can tell proper names and acronyms apart. When native review changed a reading, we replaced the training speech that used the old one, so the model never learned it.
Serving it fast
Alania-2 streams: the first audio leaves the server while the rest of the sentence is still being generated, as 48 kHz PCM inside a WAV stream or over a WebSocket. One engine step produces 0.16 seconds of audio, so playback can start almost as soon as the request arrives.
Every speech server now runs both models on one GPU: Alania-2 in its own worker process, Alania-1 next to it for existing customers. NVIDIA MPS shares the card between them; without it, Alania-1 stalled as soon as Alania-2 got busy. A server only takes traffic once Alania-2 is loaded, so a request without a model never lands on a machine that cannot answer it.
| Measurement | Result |
|---|---|
| First audio, streamed, through the public API (median) | ~190 ms |
| First audio at the server, one stream, H200 GPU | ~50 ms |
| Sustained throughput, one H200 GPU | ~12 requests/s |
| Throughput keeping first audio under 300 ms (p95), one A100 GPU | ~8 requests/s |
The gap between 50 ms at the server and 190 ms for a caller is the network and the load balancer, which is why we quote the larger number on our site.
What we are opening up
Alania-2's weights stay closed, but most of what we built to get there is public, so others can build Turkish speech technology without starting from nothing. Everything below can be used commercially with attribution.
| Release | What it is | Size | Licence |
|---|---|---|---|
| Alania Turkish Synthetic Speech | 48 kHz Turkish speech from 2,752 designed voices, each clip with its text, spoken form and a voice description; some with simulated rooms and microphones. No real person's voice. | 3,411 h, 2,037,461 clips | CC BY 4.0 (Wikipedia part CC BY-SA 4.0) |
| Alania Turkish Domain Text | Sentences a voice agent says: appointments, banking, cargo, support, numbers, codes, addresses, names, with normalised spoken forms. | 1,439,639 sentences | CC BY 4.0 |
| Alania Turkish Speech Style Captions | Real Turkish speech segments described in English and Turkish (pitch, pace, pauses, noise, room), with measurements and transcripts; audio fetched from YODAS v3. The first such set for Turkish. | 835,267 segments, 2,398 h | CC BY 4.0 (captions) |
| Antalia-1 | Our earlier open Turkish TTS model (24 kHz), its foundation model, evaluation suite and technical report. Development has ended; released in full. | ~300M parameters | See repository |
| Antalia voice corpus | Studio recordings of one professional voice actor who consented to open release. | 5.0 h | CC BY 4.0 |
| OpenVoiceCS-Bench | A benchmark for voice customer-service agents, scored from replayable traces: task resolution, policy compliance, privacy, latency, cost. | 46 models on its v0.2 leaderboard | Apache 2.0 code, CC BY 4.0 data |
If you train or evaluate a Turkish speech model, the style captions and domain text are the pieces we could not find anywhere else when we started.
What is next, and how to try it
Two features are labelled preview because they are not yet as reliable as the built-in voices: style instructions are followed loosely, and designed voices vary more in quality. Both are where we are working next, along with phone-line presets and more built-in voices.
You can hear Alania-2 now:
- Type your own Turkish on speech.patientdesk.ai; no account needed.
- Read the documentation for every parameter, streaming, and the move from Alania-1.
- Already on Alania-1? Nothing breaks: send
model: "alania-v1"to keep the old voice. Requests that name no model now get Alania-2.
Duyu, our speech-to-text model, works with the same key, so one account covers both directions of a conversation.