KosComputeAPI
Guides

Text to speech

Generate raw audio with model-defined voices, languages, and input limits.

View Markdown

POST /v1/audio/speech returns synthesized audio bytes rather than a JSON wrapper.

curl https://api.koscompute.com/v1/audio/speech \
  -H "Authorization: Bearer $KOSCOMPUTE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "kosmik/tts-kokoro-quality",
    "input": "The deployment completed successfully.",
    "voice": "af_heart",
    "speed": 1.0,
    "response_format": "wav"
  }' \
  --output speech.wav

Parameters

model and input are required. voice defaults to the model's published default. speed is a numeric playback-rate control. Supported output formats are WAV, MP3, FLAC, Opus, and PCM; an unrecognized format currently falls back to WAV, so clients should send a documented value and check Content-Type.

The maximum input length, supported languages, voice IDs, and pricing are model metadata. The current public models use a configured per-request character limit; read tts_max_chars instead of duplicating that number in application code.

Response

On success, read the response as binary audio. Relevant content types include audio/wav, audio/mpeg, audio/flac, and audio/ogg depending on the requested encoding. Save or stream the bytes directly.

Invalid input or voice selection returns 400; oversized text returns 413; capacity returns 429; and no ready TTS route returns 503. Honor Retry-After for temporary failures.

Type to search guides, models, and API reference.