---
title: Audio
description: Speech synthesis, transcription and translation on the Sylphx AI API, plus a realtime transcription session.
type: how-to
product: ai
summary: The audio endpoints, their request keys and limits.
updated: 2026-10-01
order: 9
---

Audio has its own endpoints under `https://api.sylphx.ai/v1`. Each takes the same Sylphx
key as chat and is metered on its own surface.

<Callout tone="note" title="Which models serve audio">
An audio endpoint serves the models the catalogue lists for it. The public
catalogue lists `openai/gpt-audio` and `openai/gpt-audio-mini`. Read the
catalogue your key sees (`GET https://api.sylphx.ai/v1/models`) for the speech and
transcription models you can call. A model id the catalogue does not serve
answers `404 model_not_found`.
</Callout>

## Speech: POST /audio/speech

Speaks text and answers audio bytes; the `Content-Type` of the answer says what
was produced.

```bash
curl https://api.sylphx.ai/v1/audio/speech \
  -H "Authorization: Bearer $SYLPHX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "<a speech model id from your catalogue>", "input": "Hello from Sylphx.", "voice": "alloy", "response_format": "mp3"}' \
  --output hello.mp3
```

<PropertyTable
	properties={[
		{
			name: 'model',
			type: 'string',
			description: 'A speech model id. Required.',
		},
		{
			name: 'input',
			type: 'string',
			description: 'The text to speak, up to 4,096 characters. Required.',
		},
		{
			name: 'voice',
			type: 'string',
			description: 'A voice name of 1 to 64 letters, digits, underscore, dot or dash. The ten OpenAI voice names are mapped to the serving model\'s own voices.',
		},
		{
			name: 'response_format',
			type: 'string',
			description: 'mp3 (the default), opus, aac, flac, wav or pcm. A model that produces only raw audio refuses opus, aac and flac with a 400 that names pcm, wav and mp3.',
		},
		{
			name: 'speed',
			type: 'number',
			description: '0.25 to 4. Default 1.',
		},
	]}
/>

## Transcription: POST /audio/transcriptions

Turns audio into text. Send JSON with the audio as base64 in `file`, or a
multipart form. The audio can be up to 25 MiB.

<PropertyTable
	properties={[
		{
			name: 'model',
			type: 'string',
			description: 'A transcription model id. Required.',
		},
		{
			name: 'file',
			type: 'string',
			description: 'The audio, base64-encoded. Required in a JSON request.',
		},
		{
			name: 'filename',
			type: 'string',
			description: 'The name of the uploaded file. Default audio.webm.',
		},
		{
			name: 'language',
			type: 'string',
			description: 'The spoken language, to improve accuracy.',
		},
		{
			name: 'prompt',
			type: 'string',
			description: 'Context or vocabulary to bias the transcript.',
		},
		{
			name: 'response_format',
			type: 'string',
			description: 'json (the default), text, srt, verbose_json or vtt.',
		},
		{
			name: 'temperature',
			type: 'number',
			description: '0 to 1. Default 0.',
		},
	]}
/>

Transcription is billed in whole seconds of audio.

## Translation: POST /audio/translations

The same request as transcription, sent to the translations route.

## Realtime transcription

`GET /realtime/transcription` upgrades to a WebSocket. The first event, within
10 seconds, is `session.start` with the `model`, and optionally
`language_hints`, `vocabulary`, `context`, `sample_rate` (16000 by default, or
24000) and `max_audio_seconds` (1 to 600, default 300). You then send audio as
binary frames and receive `transcript.partial` and `transcript.final` events.
A session ends when the client closes it, when the audio cap is reached or
after 30 seconds without audio.

<RelatedDocs
	links={[
		{
			href: '/docs/ai/catalog',
			label: 'Models, limits and data policy',
			description: 'List the catalogue and read a model before you use it.',
		},
		{
			href: '/docs/ai/errors',
			label: 'Errors and retries',
			description: 'The error envelope and which failures are safe to retry.',
		},
	]}
/>
