Streaming finalization · p50
Spoken Turkish.
Ready to work with.
Bring conversations into your application with Yanki. Transcribe live audio or upload a recording, with timestamps and subtitles when you need them.
A little inspiration.
“Yarın saat 14:30'daki diş randevumu 16:00'ya alabilir miyiz?”
Your microphone.
Ready to listen
Checking availability
As you speak.
Live words appear here while you speak.
In your words.
Your completed transcript will appear here.
Built around the way
people actually speak.
Heard right the first time.
- In practice
- The transcript is the foundation of every later step. A misheard drug name goes wrong into the LLM, and the conversation turns the wrong way.
- Yanki's approach
- Lowest WER on three Turkish sets: conversation, read speech and call center. A vocabulary hint (
prompt) is supported for drug names, doses and ID numbers.
Clear comparisons.
A closer look at performance.
Explore Turkish transcription across three datasets, then see what each timing measure means.
Read the docsvoicedata-TR
235 multi-speaker Turkish clips
- PatientDesk STT8.54%
- whisper-large-v310.97%
- whisper-large-v3-turbo11.19%
- Qwen3-ASR-1.7B13.65%
- mihuai/turkish-stt19.89%
- Qwen3-ASR-0.6B22.30%
Word error rate · Turkish-normalized · Lower is better
Keep pace
with the conversation.
Receive partial drafts while someone speaks, then a final transcript for each completed segment. Files use the same transcription model.
API referenceProcessing duration ÷ audio duration
Different measures for different parts of the workflow.
Connect your audio.
Keep your existing client.
Use an OpenAI-compatible transcription endpoint for files, or WebSocket events for live audio.
curl https://voice.patientdesk.ai/v1/audio/transcriptions \
-H "Authorization: Bearer pd_live_..." \
-F file=@kayit.m4a \
-F response_format=jsonChoose your output
json- text
text- plain text
verbose_json- segment and word timing
srt- subtitles
vtt- web subtitles
Included with every account
From first word to final transcript.
Partial drafts arrive while someone speaks. A final transcript closes each segment, ready for the next step in your application.
- 01
Listen
Send microphone audio over a WebSocket connection.
- 02
Follow along
Use partial events to display the conversation as it happens.
- 03
Act on the final
Use final events for a complete segment of text.
Three months free.
Both models.
To mark the launch, Alania-1 and Yanki are free for three months. Usage pricing starts after that.
Launch offer
Both models, free for three months.
For the first three months, both models.
- Alania-1 and Yanki, free for three months
- REST and live streaming
- One API key for both models
- No card required
Usage-based
After the three months, keep building at usage prices.
- Starts after the three free months
- Pay only for additional usage
- Same API, more capacity
- Billed monthly
Enterprise
For your infrastructure and your team.
- Custom quotas and concurrency
- Deployment on your infrastructure
- Data-processing terms
- Priority support
Additional usage is billed monthly. Prices exclude VAT.
Give the conversation a voice.
One account. Two sides of the conversation.