Skip to content

Getting started

Quickstart

Generate your first audio and transcribe your first recording in a few minutes.

  1. Create an API key in the dashboard. It is shown only once, so store it somewhere safe.
  2. Send Alania a sentence; the response is a WAV file.
  3. Send a recording to Duyu and get the text back.

1 · Speak

POSThttps://voice.patientdesk.ai/v1/audio/speechTry it
curl https://voice.patientdesk.ai/v1/audio/speech \
  -H "Authorization: Bearer pd_live_..." \
  -H "Content-Type: application/json" \
  -d '{ "model": "alania-v2", "input": "Merhaba, randevunuz onaylandı.", "voice": "alania2-f1" }' \
  --output merhaba.wav   # audio/wav, 48 kHz mono

If you leave out model, the request reaches Alania-2 and speaks with the alania2-f1 voice. The response is a 48 kHz mono WAV file. The Python example uses the OpenAI SDK; changing base_url is all it takes.

With Alania-1

The previous model keeps working as alania-v1: 24 kHz, one voice (alania), and no style, design or cloning fields.

curl https://voice.patientdesk.ai/v1/audio/speech \
  -H "Authorization: Bearer pd_live_..." \
  -H "Content-Type: application/json" \
  -d '{ "model": "alania-v1", "input": "Merhaba, randevunuz onaylandı.", "voice": "alania" }' \
  --output merhaba-v1.wav   # audio/wav, 24 kHz mono

2 · Transcribe

POSThttps://voice.patientdesk.ai/v1/audio/transcriptionsTry it
curl https://voice.patientdesk.ai/v1/audio/transcriptions \
  -H "Authorization: Bearer pd_live_..." \
  -F file=@kayit.m4a \
  -F response_format=json
Response
{ "text": "Randevunuz yarın saat 14:05 için oluşturuldu." }

Then