Alania · TTS
Streaming
Start playing while the rest is still being generated: over HTTP with stream: true, or over one WebSocket connection for many sentences.
HTTP streaming
With stream: true the body starts with a 44-byte WAV header, followed by raw PCM16 as it is generated: little-endian, mono, 48 kHz on Alania-2 and 24 kHz on Alania-1. The length of a stream is not known in advance, so the header's length fields are 0xFFFFFFFF.
curl -N https://voice.patientdesk.ai/v1/audio/speech \
-H "Authorization: Bearer pd_live_..." \
-H "Content-Type: application/json" \
-d '{ "model": "alania-v2", "voice": "alania2-f1", "input": "Randevunuz oluşturuldu.", "stream": true }' \
| ffplay -autoexit -nodisp - # 48 kHz PCM16 as it is generated- Streams accept only
wavandpcm.pcmis raw samples without a header; the rate is inContent-Type: audio/L16; rate=48000. - Network chunks may split a sample: carry a leftover byte over to the next chunk.
- Long text is split at sentence boundaries, with 250 ms of silence between parts.
x-audio-secondsis not sent on a stream: it is not known until the audio ends.
48 kHz playback in the browser
The Browser example above schedules each chunk with Web Audio. What matters:
- Create buffers at the right rate:
createBuffer(1, n, 48000). Read the rate from byte 24 of the WAV header, or fromspeech_start.sample_rateon a WebSocket; Alania-1 is 24 kHz. new AudioContext({ sampleRate: 48000 })avoids needless resampling; if the device runs at another rate, the browser converts.- Schedule each chunk at the end of the previous one; a 50 ms margin on the first avoids a stutter.
- An
AudioContextonly plays after a user gesture; start the request from a click.
Latency
On Alania-2, a streamed request gets its first audio in about 190 ms (median, measured through the public API with the network included). What you measure depends on your network and distance to the servers. The x-queue-wait-ms header reports the time the request waited for a generation slot.
WebSocket
wss://voice.patientdesk.ai/v1/audio/speech/streamFor an application that speaks many sentences in a row, such as a voice agent, use one connection. Each speak message returns one audio stream and the connection stays open. Authenticate with ?key= or ?token=; ?model=alania-v1 makes Alania-1 the connection's default model.
const ws = new WebSocket("wss://voice.patientdesk.ai/v1/audio/speech/stream?key=pd_live_...");
ws.binaryType = "arraybuffer";
let rate = 48000;
ws.onmessage = (e) => {
if (e.data instanceof ArrayBuffer) return play(new Int16Array(e.data), rate); // PCM16 mono
const m = JSON.parse(e.data);
if (m.type === "ready") ws.send(JSON.stringify({
type: "speak",
model: "alania-v2",
voice: "alania2-f1",
input: "Randevunuz yarın saat 14:05 için oluşturuldu.",
}));
if (m.type === "speech_start") rate = m.sample_rate; // 48000 (24000 for alania-v1)
if (m.type === "speech_end") ws.send(JSON.stringify({ type: "stop" }));
if (m.type === "error") console.error(m.code, m.message);
};Messages you send
speakJSONinputis required. Every field of the HTTP body works:model,voice,instructions,voice_description,reference_audio(base64),reference_text,consent,cfg,temperature,seed. Messages are handled in order, one at a time.stopJSON- Answered with
stopped, then the connection closes.closedoes the same.
Messages you receive
readyJSONmodel,voice,sample_rate,models,usage. The connection is ready;usageis your remaining allowance.speech_startJSONcharacters,model,voice,sample_rate,encoding(pcm_s16le),disclosure. Arrives before the audio of eachspeak.binaryPCM16- Audio, PCM16 mono. On Alania-2 messages are usually 0.16 seconds (7,680 samples at 48 kHz); the last one may be shorter.
speech_endJSONaudio_s,infer_ms. Thisspeakis done.stoppedJSON- The reply to
stop. errorJSONmessage,code,category,retry_after. After an unknown message type the connection stays open; after any other error it closes.
{"type":"ready","model":"alania-v2","voice":"alania2-f1","sample_rate":48000,"models":["alania-v1","alania-v2"],"usage":{"unit":"characters","used":1204,"quota":100000,"remaining":98796,"free":true,"free_until":null}}
{"type":"speech_start","characters":45,"model":"alania-v2","voice":"alania2-f1","sample_rate":48000,"encoding":"pcm_s16le","disclosure":"ai-generated"}
<binary PCM16 frames, 0.16 s each>
{"type":"speech_end","audio_s":3.2,"infer_ms":…}
{"type":"stopped"}{"type":"error","message":"Cloning a voice needs `consent: true`: …","code":"consent_required","category":"invalid_request_error","retry_after":0}Compressed formats
For non-streaming requests response_format can be mp3, opus (Ogg), flac or aac, which saves bandwidth for telephony and mobile apps. Streams accept only wav and pcm: half a container file is not playable, so anything else is refused with a 400.