Skip to contentLaunchAlania-1 and Duyu are free for three months.See the offer
DuyuSpeech to text

Spoken Turkish.
Ready to work with.

Bring conversations into your application with Duyu. Transcribe live audio or upload a recording, with timestamps and subtitles when you need them.

Live & recorded audioWord timestampsSRT & VTT

A little inspiration.

Yarın saat 14:30'daki diş randevumu 16:00'ya alabilir miyiz?

Your microphone.

Ready to listen

In your words.

Capabilities

Built around the way
people actually speak.

Duyu
WER · lower is better
Conversation
PatientDesk STT
8.54
whisper-large-v3
10.97
Read speech
PatientDesk STT
4.71
whisper-large-v3
5.04
Call center
PatientDesk STT
9.87
whisper-large-v3
15.18
WER · lower is better

Heard right the first time.

In practice
The transcript is the foundation of every later step. A misheard drug name goes wrong into the LLM, and the conversation turns the wrong way.
Duyu's approach
Lowest WER on three Turkish sets: conversation, read speech and call center. A vocabulary hint (prompt) is supported for drug names, doses and ID numbers.
The details matter

Clear comparisons.
A closer look at performance.

Explore Turkish transcription across three datasets, then see what each timing measure means.

Read the docs

voicedata-TR

235 multi-speaker Turkish clips

  • PatientDesk STT8.54%
  • whisper-large-v310.97%
  • whisper-large-v3-turbo11.19%
  • Qwen3-ASR-1.7B13.65%
  • mihuai/turkish-stt19.89%
  • Qwen3-ASR-0.6B22.30%

Word error rate · Turkish-normalized · Lower is better

Performance

Keep pace
with the conversation.

Receive partial drafts while someone speaks, then a final transcript for each completed segment. Files use the same transcription model.

API reference
Segment end to final text~120ms

Streaming finalization · p50

File transcription real-time factor0.08×

Processing duration ÷ audio duration

Different measures for different parts of the workflow.

API reference

Connect your audio.
Keep your existing client.

Use an OpenAI-compatible transcription endpoint for files, or WebSocket events for live audio.

curl https://voice.patientdesk.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer pd_live_..." \
  -F file=@kayit.m4a \
  -F response_format=json
POST /v1/audio/transcriptions

Choose your output

json
text
text
plain text
verbose_json
segment and word timing
srt
subtitles
vtt
web subtitles

Included with every account

From first word to final transcript.

Partial drafts arrive while someone speaks. A final transcript closes each segment, ready for the next step in your application.

  1. 01

    Listen

    Send microphone audio over a WebSocket connection.

  2. 02

    Follow along

    Use partial events to display the conversation as it happens.

  3. 03

    Act on the final

    Use final events for a complete segment of text.

Launch offer

Three months free.
Both models.

To mark the launch, Alania-1 and Duyu are free for three months. Usage pricing starts after that.

Usage-based

After the three months, keep building at usage prices.

Duyu$0.04/minute
Alania$0.06/1,000 characters
  • Starts after the three free months
  • Pay only for additional usage
  • Same API, more capacity
  • Billed monthly
Start free

Enterprise

For your infrastructure and your team.

Let’s talk
  • Custom quotas and concurrency
  • Deployment on your infrastructure
  • Data-processing terms
  • Priority support
Contact sales

Additional usage is billed monthly. Prices exclude VAT.

Better together

Give the conversation a voice.

One account. Two sides of the conversation.

Explore Alania
Get started

Create your key.
Send a request.

Start free