Skip to content

Benchmark · Freya-TR-Eval

Turkish text to speech, on the same test

We measured Alania against open models and cloud services under the same conditions: 495 everyday Turkish sentences, three takes per system, the same speech recognizer and the same GPU. Every system with public weights we ran again ourselves.

  • 495 sentences
  • 3 takes per system
  • Whisper large-v3 · 8 kHz
  • NVIDIA L4

Accuracy

How many words Whisper misses when it turns the audio back into text. Mean of three takes.

Word error rate (WER)

Mean of three takes, log scale · Lower is better

125102050100EMA Lightning (plain) · plain PyTorch: 1.04%1.04EMA Lightning (plain)EMA Lightning (compiled) · compiled: 1.04%1.04EMA Lightning (compiled)Trendyol-TTS · cfg 2.0, 16 steps: 1.16%1.16Trendyol-TTSGemini 3.8 Flash TTS · voice Kore: 1.23%1.23Gemini 3.8 Flash TTSGemini 3.8 Flash-Lite TTS · voice Kore: 1.34%1.34Gemini 3.8 Flash-Lite TTSAnka TTS · built-in male voice: 1.42%1.42Anka TTSElevenLabs v4 · voice George: 1.43% · Quoted from the EMA Lightning card, not measured by us.1.43ElevenLabs v4Alania-2 · production API, voice F1: 1.44%1.44Alania-2Alania-2 (model) · model only, no serving checks: 1.51%1.51Alania-2 (model)ElevenLabs v4 Turbo · voice George: 1.54% · Quoted from the EMA Lightning card, not measured by us.1.54ElevenLabs v4 TurboChatterbox Multilingual · Anka male reference: 1.84% · Our own run is still being scored; the published number is shown for now.1.84Chatterbox MultilingualPiper · tr_TR dfki medium, CPU: 3.41%3.41PiperXTTS-v2 · Anka male reference: 3.67%3.67XTTS-v2Alania-1 · production API: 5.99%5.99Alania-1Antalia-1 · open weights, release recipe: 6.28%6.28Antalia-1MMS-TTS · tur: 6.38%6.38MMS-TTSCoqui GlowTTS · tr common-voice, lowercased input: 10.93%10.93Coqui GlowTTSFreyaTTS-small · default voice: 12.02%12.02FreyaTTS-smallF5-TTS base · no Turkish training: 77.84% · Quoted from the Anka TTS card, not measured by us.77.84F5-TTS base
PatientDesk AIEMATrendyolGoogleAnkaResemble AIRhasspyCoquiMetaFreyaElevenLabsF5-TTSQuoted from another card

Speed

Generation time against audio length on one NVIDIA L4, one request at a time. Systems with public weights only.

Times faster than real time

1 / RTF, NVIDIA L4, one request, log scale · Higher is better

1×10×100×EMA Lightning (compiled) · compiled: 333.3×333EMA Lightning (compiled)MMS-TTS · tur: 72.5×72MMS-TTSCoqui GlowTTS · tr common-voice, lowercased input: 68.5×68Coqui GlowTTSEMA Lightning (plain) · plain PyTorch: 57.8×58EMA Lightning (plain)Piper · tr_TR dfki medium, CPU: 26.5×26PiperAlania-2 (model) · model only, no serving checks: 6.6×6.6Alania-2 (model)FreyaTTS-small · default voice: 3.5×3.5FreyaTTS-smallAnka TTS · built-in male voice: 3.1×3.1Anka TTSXTTS-v2 · Anka male reference: 2.2×2.2XTTS-v2Antalia-1 · open weights, release recipe: 2.1×2.1Antalia-1Trendyol-TTS · cfg 2.0, 16 steps: 1.0×1.0Trendyol-TTS
PatientDesk AIEMATrendyolGoogleAnkaResemble AIRhasspyCoquiMetaFreyaElevenLabsF5-TTSQuoted from another card

Naturalness

UTMOS is a model that predicts how natural speech sounds. It was trained on English, so read it as a rough guide.

UTMOS

Predicted naturalness from 1 to 5 · Higher is better

01234Trendyol-TTS · cfg 2.0, 16 steps: 3.823.82Trendyol-TTSMMS-TTS · tur: 3.793.79MMS-TTSPiper · tr_TR dfki medium, CPU: 3.713.71PiperAlania-2 · production API, voice F1: 3.673.67Alania-2Alania-2 (model) · model only, no serving checks: 3.663.66Alania-2 (model)Gemini 3.8 Flash-Lite TTS · voice Kore: 3.633.63Gemini 3.8 Flash-Lite TTSGemini 3.8 Flash TTS · voice Kore: 3.533.53Gemini 3.8 Flash TTSAnka TTS · built-in male voice: 3.523.52Anka TTSEMA Lightning (plain) · plain PyTorch: 3.303.30EMA Lightning (plain)EMA Lightning (compiled) · compiled: 3.303.30EMA Lightning (compiled)Alania-1 · production API: 3.193.19Alania-1Coqui GlowTTS · tr common-voice, lowercased input: 3.083.08Coqui GlowTTSXTTS-v2 · Anka male reference: 2.972.97XTTS-v2FreyaTTS-small · default voice: 2.812.81FreyaTTS-smallAntalia-1 · open weights, release recipe: 2.722.72Antalia-1
PatientDesk AIEMATrendyolGoogleAnkaResemble AIRhasspyCoquiMetaFreyaElevenLabsF5-TTSQuoted from another card

Leaderboard

Every measurement in one table, by word error rate.

#SystemProviderParamsWER %CER %CleanUTMOS× real timeSource
1EMA Lightning (plain)plain PyTorchEMA8.6M1.04 ± 0.170.22 ± 0.0793.7%3.3058×Our run
2EMA Lightning (compiled)compiledEMA8.6M1.04 ± 0.170.22 ± 0.0793.7%3.30333×Our run
3Trendyol-TTScfg 2.0, 16 stepsTrendyol2.38B1.16 ± 0.160.23 ± 0.0193.1%3.821.0×Our run
4Gemini 3.8 Flash TTSvoice KoreGoogle–1.23 ± 0.100.29 ± 0.0792.9%3.53–Our run
5Gemini 3.8 Flash-Lite TTSvoice KoreGoogle–1.34 ± 0.190.27 ± 0.0591.9%3.63–Our run
6Anka TTSbuilt-in male voiceAnka336M1.42 ± 0.060.30 ± 0.0391.6%3.523.1×Our run
7ElevenLabs v4voice GeorgeElevenLabs–1.43 ± 0.120.27 ± 0.03–––EMA Lightning card
8Alania-2production API, voice F1PatientDesk AI488M1.44 ± 0.150.28 ± 0.0491.7%3.67–Our run
9Alania-2 (model)model only, no serving checksPatientDesk AI488M1.51 ± 0.190.30 ± 0.0591.8%3.666.6×Our run
10ElevenLabs v4 Turbovoice GeorgeElevenLabs–1.54 ± 0.220.29 ± 0.04–––EMA Lightning card
11Chatterbox MultilingualAnka male referenceResemble AI500M1.84 ± 0.200.41 ± 0.08–––Our own run is still being scored; the published number is shown for now.
12Pipertr_TR dfki medium, CPURhasspy16M3.41 ± 0.210.73 ± 0.0579.4%3.7126×Our run
13XTTS-v2Anka male referenceCoqui470M3.67 ± 0.371.50 ± 0.1677.1%2.972.2×Our run
14Alania-1production APIPatientDesk AI3.3B5.99 ± 0.471.81 ± 0.1767.3%3.19–Our run
15Antalia-1open weights, release recipePatientDesk AI305M6.28 ± 0.272.49 ± 0.0866.7%2.722.1×Our run
16MMS-TTSturMeta36M6.38 ± 0.191.48 ± 0.0565.7%3.7972×Our run
17Coqui GlowTTStr common-voice, lowercased inputCoqui28M10.93 ± 0.002.96 ± 0.0046.1%3.0868×Our run
18FreyaTTS-smalldefault voiceFreya183M12.02 ± 0.004.95 ± 0.0047.1%2.813.5×Our run
19F5-TTS baseno Turkish trainingF5-TTS336M77.84 ± 0.1332.58 ± 0.54–––Anka TTS card

Listen

Four sentences from the test, first take of each system. Systems whose licences allow it.

Boğaz ağrın geçmediyse neden hala sıcak bir ıhlamur kaynatmıyorsun?
Uzun zaman sonra sesini duymak ne kadar güzel!
Bahar ayları gelince polen alerjim yüzünden sürekli hapşırıp duruyorum.
Adaylar Kasım ayı başlarında duyurulacak.

How we measured

Sentences
All 495 sentences of Freya-TR-Eval (CC-BY-4.0): everyday Turkish, 3–13 words. Raw text in: each system uses its own text front end.
Takes
Three per system. Open models: seeds 0, 1 and 2. APIs: three independent requests per sentence.
Scoring
EMA Lightning's published eval.py, verbatim: audio resampled to 8 kHz and back to 16 kHz, transcribed by faster-whisper large-v3 (Turkish, beam 5), WER and CER after the same Turkish normalisation. A take that fails counts as silence.
Speed
Each system ran alone on one NVIDIA L4, one request at a time. Three passes after a warm-up; the middle pass is reported. Piper ran on the CPU, as it does by default.
Our models
Alania-2 with voice F1 through our production API (text normaliser and checks included), and the same model with the checks off and one generation per sentence. Alania-1 through our production API. Antalia-1 from its open weights with the released inference recipe.
Quoted
ElevenLabs v4 and v4 Turbo from the EMA Lightning card, F5-TTS base from the Anka TTS card. We did not measure these.

What this test does not show

  • Word error rate measures whether Whisper understands the audio, not how warm or pleasant the voice is. A listening test is the real measure of that, and we have not run one here.
  • The sentences are short and contain no numbers, dates or abbreviations, which is where Turkish text normalisation is hardest.
  • Whisper large-v3 makes its own Turkish errors. They count against every system equally, but put a floor under every score.
  • Each system speaks in one voice, and some voices are easier to transcribe than others.

Every script, the transcript and UTMOS score of every clip, and the timing of every pass are on Hugging Face:

PatientdeskAI/freya-tr-eval-rerun

Benchmark: Freya-TR-Eval by Freya. Protocol and scoring code: the EMA Lightning card and its eval.py. Settings for the voice-cloning systems: the Anka TTS report.