Benchmark · Freya-TR-Eval
Turkish text to speech, on the same test
We measured Alania against open models and cloud services under the same conditions: 495 everyday Turkish sentences, three takes per system, the same speech recognizer and the same GPU. Every system with public weights we ran again ourselves.
- 495 sentences
- 3 takes per system
- Whisper large-v3 · 8 kHz
- NVIDIA L4
Accuracy
How many words Whisper misses when it turns the audio back into text. Mean of three takes.
Word error rate (WER)
Mean of three takes, log scale · Lower is better
Speed
Generation time against audio length on one NVIDIA L4, one request at a time. Systems with public weights only.
Times faster than real time
1 / RTF, NVIDIA L4, one request, log scale · Higher is better
Naturalness
UTMOS is a model that predicts how natural speech sounds. It was trained on English, so read it as a rough guide.
UTMOS
Predicted naturalness from 1 to 5 · Higher is better
Leaderboard
Every measurement in one table, by word error rate.
| # | System | Provider | Params | WER % | CER % | Clean | UTMOS | × real time | Source |
|---|---|---|---|---|---|---|---|---|---|
| 1 | EMA Lightning (plain)plain PyTorch | EMA | 8.6M | 1.04 ± 0.17 | 0.22 ± 0.07 | 93.7% | 3.30 | 58× | Our run |
| 2 | EMA Lightning (compiled)compiled | EMA | 8.6M | 1.04 ± 0.17 | 0.22 ± 0.07 | 93.7% | 3.30 | 333× | Our run |
| 3 | Trendyol-TTScfg 2.0, 16 steps | Trendyol | 2.38B | 1.16 ± 0.16 | 0.23 ± 0.01 | 93.1% | 3.82 | 1.0× | Our run |
| 4 | Gemini 3.8 Flash TTSvoice Kore | – | 1.23 ± 0.10 | 0.29 ± 0.07 | 92.9% | 3.53 | – | Our run | |
| 5 | Gemini 3.8 Flash-Lite TTSvoice Kore | – | 1.34 ± 0.19 | 0.27 ± 0.05 | 91.9% | 3.63 | – | Our run | |
| 6 | Anka TTSbuilt-in male voice | Anka | 336M | 1.42 ± 0.06 | 0.30 ± 0.03 | 91.6% | 3.52 | 3.1× | Our run |
| 7 | ElevenLabs v4voice George | ElevenLabs | – | 1.43 ± 0.12 | 0.27 ± 0.03 | – | – | – | EMA Lightning card |
| 8 | Alania-2production API, voice F1 | PatientDesk AI | 488M | 1.44 ± 0.15 | 0.28 ± 0.04 | 91.7% | 3.67 | – | Our run |
| 9 | Alania-2 (model)model only, no serving checks | PatientDesk AI | 488M | 1.51 ± 0.19 | 0.30 ± 0.05 | 91.8% | 3.66 | 6.6× | Our run |
| 10 | ElevenLabs v4 Turbovoice George | ElevenLabs | – | 1.54 ± 0.22 | 0.29 ± 0.04 | – | – | – | EMA Lightning card |
| 11 | Chatterbox MultilingualAnka male reference | Resemble AI | 500M | 1.84 ± 0.20 | 0.41 ± 0.08 | – | – | – | Our own run is still being scored; the published number is shown for now. |
| 12 | Pipertr_TR dfki medium, CPU | Rhasspy | 16M | 3.41 ± 0.21 | 0.73 ± 0.05 | 79.4% | 3.71 | 26× | Our run |
| 13 | XTTS-v2Anka male reference | Coqui | 470M | 3.67 ± 0.37 | 1.50 ± 0.16 | 77.1% | 2.97 | 2.2× | Our run |
| 14 | Alania-1production API | PatientDesk AI | 3.3B | 5.99 ± 0.47 | 1.81 ± 0.17 | 67.3% | 3.19 | – | Our run |
| 15 | Antalia-1open weights, release recipe | PatientDesk AI | 305M | 6.28 ± 0.27 | 2.49 ± 0.08 | 66.7% | 2.72 | 2.1× | Our run |
| 16 | MMS-TTStur | Meta | 36M | 6.38 ± 0.19 | 1.48 ± 0.05 | 65.7% | 3.79 | 72× | Our run |
| 17 | Coqui GlowTTStr common-voice, lowercased input | Coqui | 28M | 10.93 ± 0.00 | 2.96 ± 0.00 | 46.1% | 3.08 | 68× | Our run |
| 18 | FreyaTTS-smalldefault voice | Freya | 183M | 12.02 ± 0.00 | 4.95 ± 0.00 | 47.1% | 2.81 | 3.5× | Our run |
| 19 | F5-TTS baseno Turkish training | F5-TTS | 336M | 77.84 ± 0.13 | 32.58 ± 0.54 | – | – | – | Anka TTS card |
Listen
Four sentences from the test, first take of each system. Systems whose licences allow it.
Boğaz ağrın geçmediyse neden hala sıcak bir ıhlamur kaynatmıyorsun?Uzun zaman sonra sesini duymak ne kadar güzel!Bahar ayları gelince polen alerjim yüzünden sürekli hapşırıp duruyorum.Adaylar Kasım ayı başlarında duyurulacak.How we measured
- Sentences
- All 495 sentences of Freya-TR-Eval (CC-BY-4.0): everyday Turkish, 3–13 words. Raw text in: each system uses its own text front end.
- Takes
- Three per system. Open models: seeds 0, 1 and 2. APIs: three independent requests per sentence.
- Scoring
- EMA Lightning's published eval.py, verbatim: audio resampled to 8 kHz and back to 16 kHz, transcribed by faster-whisper large-v3 (Turkish, beam 5), WER and CER after the same Turkish normalisation. A take that fails counts as silence.
- Speed
- Each system ran alone on one NVIDIA L4, one request at a time. Three passes after a warm-up; the middle pass is reported. Piper ran on the CPU, as it does by default.
- Our models
- Alania-2 with voice F1 through our production API (text normaliser and checks included), and the same model with the checks off and one generation per sentence. Alania-1 through our production API. Antalia-1 from its open weights with the released inference recipe.
- Quoted
- ElevenLabs v4 and v4 Turbo from the EMA Lightning card, F5-TTS base from the Anka TTS card. We did not measure these.
What this test does not show
- Word error rate measures whether Whisper understands the audio, not how warm or pleasant the voice is. A listening test is the real measure of that, and we have not run one here.
- The sentences are short and contain no numbers, dates or abbreviations, which is where Turkish text normalisation is hardest.
- Whisper large-v3 makes its own Turkish errors. They count against every system equally, but put a floor under every score.
- Each system speaks in one voice, and some voices are easier to transcribe than others.
Every script, the transcript and UTMOS score of every clip, and the timing of every pass are on Hugging Face:
PatientdeskAI/freya-tr-eval-rerunBenchmark: Freya-TR-Eval by Freya. Protocol and scoring code: the EMA Lightning card and its eval.py. Settings for the voice-cloning systems: the Anka TTS report.