Skip to content
Console
Menu

Queues

Workflows

Getting Started

Authentication

KV Store

Audio

The audio endpoints, their request keys and limits.

Audio has its own endpoints under https://api.sylphx.ai/v1. Each takes the same Sylphx key as chat and is metered on its own surface.

Which models serve audio

An audio endpoint serves the models the catalogue lists for it. The public catalogue lists openai/gpt-audio and openai/gpt-audio-mini. Read the catalogue your key sees (GET https://api.sylphx.ai/v1/models) for the speech and transcription models you can call. A model id the catalogue does not serve answers 404 model_not_found.

#Speech: POST /audio/speech

Speaks text and answers audio bytes; the Content-Type of the answer says what was produced.

Shell
curl https://api.sylphx.ai/v1/audio/speech \
  -H "Authorization: Bearer $SYLPHX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "<a speech model id from your catalogue>", "input": "Hello from Sylphx.", "voice": "alloy", "response_format": "mp3"}' \
  --output hello.mp3
FieldTypeWhat it is
modelstringA speech model id. Required.
inputstringThe text to speak, up to 4,096 characters. Required.
voicestringA voice name of 1 to 64 letters, digits, underscore, dot or dash. The ten OpenAI voice names are mapped to the serving model's own voices.
response_formatstringmp3 (the default), opus, aac, flac, wav or pcm. A model that produces only raw audio refuses opus, aac and flac with a 400 that names pcm, wav and mp3.
speednumber0.25 to 4. Default 1.

#Transcription: POST /audio/transcriptions

Turns audio into text. Send JSON with the audio as base64 in file, or a multipart form. The audio can be up to 25 MiB.

FieldTypeWhat it is
modelstringA transcription model id. Required.
filestringThe audio, base64-encoded. Required in a JSON request.
filenamestringThe name of the uploaded file. Default audio.webm.
languagestringThe spoken language, to improve accuracy.
promptstringContext or vocabulary to bias the transcript.
response_formatstringjson (the default), text, srt, verbose_json or vtt.
temperaturenumber0 to 1. Default 0.

Transcription is billed in whole seconds of audio.

#Translation: POST /audio/translations

The same request as transcription, sent to the translations route.

#Realtime transcription

GET /realtime/transcription upgrades to a WebSocket. The first event, within 10 seconds, is session.start with the model, and optionally language_hints, vocabulary, context, sample_rate (16000 by default, or 24000) and max_audio_seconds (1 to 600, default 300). You then send audio as binary frames and receive transcript.partial and transcript.final events. A session ends when the client closes it, when the audio cap is reached or after 30 seconds without audio.