Streaming finalization · p50
Spoken Turkish.
Ready to work with.
Bring conversations into your application with Duyu. Transcribe live audio or upload a recording, with timestamps and subtitles when you need them.
A little inspiration.
“Yarın saat 14:30'daki diş randevumu 16:00'ya alabilir miyiz?”
Your microphone.
Ready to listen
In your words.
Built around the way
people actually speak.
Heard right the first time.
- In practice
- The transcript is the foundation of every later step. A misheard drug name goes wrong into the LLM, and the conversation turns the wrong way.
- Duyu's approach
- Lowest WER on three Turkish sets: conversation, read speech and call center. A vocabulary hint (
prompt) is supported for drug names, doses and ID numbers.
Clear comparisons.
A closer look at performance.
Explore Turkish transcription across three datasets, then see what each timing measure means.
Read the docsvoicedata-TR
235 multi-speaker Turkish clips
- PatientDesk STT8.54%
- whisper-large-v310.97%
- whisper-large-v3-turbo11.19%
- Qwen3-ASR-1.7B13.65%
- mihuai/turkish-stt19.89%
- Qwen3-ASR-0.6B22.30%
Word error rate · Turkish-normalized · Lower is better
Keep pace
with the conversation.
Receive partial drafts while someone speaks, then a final transcript for each completed segment. Files use the same transcription model.
API referenceProcessing duration ÷ audio duration
Different measures for different parts of the workflow.
Connect your audio.
Keep your existing client.
Use an OpenAI-compatible transcription endpoint for files, or WebSocket events for live audio.
curl https://voice.patientdesk.ai/v1/audio/transcriptions \
-H "Authorization: Bearer pd_live_..." \
-F file=@kayit.m4a \
-F response_format=jsonChoose your output
json- text
text- plain text
verbose_json- segment and word timing
srt- subtitles
vtt- web subtitles
Included with every account
From first word to final transcript.
Partial drafts arrive while someone speaks. A final transcript closes each segment, ready for the next step in your application.
- 01
Listen
Send microphone audio over a WebSocket connection.
- 02
Follow along
Use partial events to display the conversation as it happens.
- 03
Act on the final
Use final events for a complete segment of text.
Three months free.
Both models.
To mark the launch, Alania-1 and Duyu are free for three months. Usage pricing starts after that.
Launch offer
Both models, free for three months.
For the first three months, both models.
- Alania-1 and Duyu, free for three months
- REST and live streaming
- One API key for both models
- No card required
Usage-based
After the three months, keep building at usage prices.
- Starts after the three free months
- Pay only for additional usage
- Same API, more capacity
- Billed monthly
Enterprise
For your infrastructure and your team.
- Custom quotas and concurrency
- Deployment on your infrastructure
- Data-processing terms
- Priority support
Additional usage is billed monthly. Prices exclude VAT.
Give the conversation a voice.
One account. Two sides of the conversation.