Menu
Platform
AI
App store purchases
Database
Flags
Jobs and cron
Localization
Monitoring
Notifications
Payments
Queues
Sandboxes
Webhooks
Getting Started
Authentication
KV Store
Deploy & Infrastructure
Reference
Audio
The audio endpoints, their request keys and limits.
Audio has its own endpoints under https://api.sylphx.ai/v1. Each takes the same Sylphx
key as chat and is metered on its own surface.
Which models serve audio
An audio endpoint serves the models the catalogue lists for it. The public
catalogue lists openai/gpt-audio and openai/gpt-audio-mini. Read the
catalogue your key sees (GET https://api.sylphx.ai/v1/models) for the speech and
transcription models you can call. A model id the catalogue does not serve
answers 404 model_not_found.
#Speech: POST /audio/speech
Speaks text and answers audio bytes; the Content-Type of the answer says what
was produced.
curl https://api.sylphx.ai/v1/audio/speech \
-H "Authorization: Bearer $SYLPHX_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "<a speech model id from your catalogue>", "input": "Hello from Sylphx.", "voice": "alloy", "response_format": "mp3"}' \
--output hello.mp3| Field | Type | What it is |
|---|---|---|
model | string | A speech model id. Required. |
input | string | The text to speak, up to 4,096 characters. Required. |
voice | string | A voice name of 1 to 64 letters, digits, underscore, dot or dash. The ten OpenAI voice names are mapped to the serving model's own voices. |
response_format | string | mp3 (the default), opus, aac, flac, wav or pcm. A model that produces only raw audio refuses opus, aac and flac with a 400 that names pcm, wav and mp3. |
speed | number | 0.25 to 4. Default 1. |
#Transcription: POST /audio/transcriptions
Turns audio into text. Send JSON with the audio as base64 in file, or a
multipart form. The audio can be up to 25 MiB.
| Field | Type | What it is |
|---|---|---|
model | string | A transcription model id. Required. |
file | string | The audio, base64-encoded. Required in a JSON request. |
filename | string | The name of the uploaded file. Default audio.webm. |
language | string | The spoken language, to improve accuracy. |
prompt | string | Context or vocabulary to bias the transcript. |
response_format | string | json (the default), text, srt, verbose_json or vtt. |
temperature | number | 0 to 1. Default 0. |
Transcription is billed in whole seconds of audio.
#Translation: POST /audio/translations
The same request as transcription, sent to the translations route.
#Realtime transcription
GET /realtime/transcription upgrades to a WebSocket. The first event, within
10 seconds, is session.start with the model, and optionally
language_hints, vocabulary, context, sample_rate (16000 by default, or
24000) and max_audio_seconds (1 to 600, default 300). You then send audio as
binary frames and receive transcript.partial and transcript.final events.
A session ends when the client closes it, when the audio cap is reached or
after 30 seconds without audio.