---
title: Responses API
description: Create, stream, store, compact and continue responses, upload files, and use the Messages and Chat Completions encodings.
type: reference
product: ai
summary: The Responses endpoints, fields, streaming events, storage and compaction
updated: 2026-09-30
order: 3
---

The Responses API is the primary wire of Sylphx AI: the official OpenAI
Responses document, served at `https://api.sylphx.ai/v1`. Every path below is
relative to that base URL and takes your bearer key.

## Endpoints

| Method | Path | Purpose |
| --- | --- | --- |
| `POST` | `/responses` | Create a response. Add `stream: true` for server-sent events. |
| `POST` | `/responses/compact` | Compact a stored response into an official compaction object. |
| `GET` | `/responses/{response_id}` | Retrieve a stored response while it is retained. |
| `GET` | `/responses/{response_id}/input_items` | List the input items stored for one response. |
| `DELETE` | `/responses/{response_id}` | Forget a stored response. |
| `POST` | `/responses/{response_id}/cancel` | The official cancel route. Responses are not background jobs. |
| `POST` | `/chat/completions` | Chat Completions encoding, normalized into the same Responses document. |
| `POST` | `/messages` | Anthropic Messages encoding, normalized into the same Responses document. |
| `POST` | `/files` | Multipart upload. Mints a `file-*` id for `input_file` and `input_image`. |
| `GET` | `/files` | List the files owned by your key. |
| `GET` | `/files/{file_id}` | File metadata, no bytes. |
| `GET` | `/files/{file_id}/content` | Raw file bytes. |
| `DELETE` | `/files/{file_id}` | Delete a file you own. |

## Create a response

`POST /responses` takes a strict JSON document. The only required field is
`model`; `input` carries what you want answered.

```bash
curl https://api.sylphx.ai/v1/responses \
  -H "Authorization: Bearer $SYLPHX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-5.5",
    "instructions": "Answer as a concise technical writer.",
    "input": "Summarise this incident report in three bullets.",
    "max_output_tokens": 400,
    "store": true
  }'
```

| Field | Type | What it does |
| --- | --- | --- |
| `model` (required) | string | A model id from [the catalogue](/ai/models). There is no router alias: an id we do not sell is a typed `404`, never a silent substitution. |
| `input` | string or array | The user turn, or a full item list. Replayed function calls pair with their outputs by `call_id`. |
| `instructions` | string | System-level guidance for this turn. Only text you write is sent; the platform never appends its own. |
| `stream` | boolean | `false` (the default) returns one JSON response; `true` returns `text/event-stream`. |
| `store` | boolean | Defaults to `true`: the response is retained and can be retrieved or used as `previous_response_id`. `false` skips persistence. |
| `previous_response_id` | string | Continue a stored conversation from an earlier response id. |
| `max_output_tokens` | integer | Upper bound on the tokens this turn may generate. |
| `tools` | array | Function tools your code executes, plus [hosted tools](/docs/ai/tools) the platform executes. |
| `tool_choice` | string or object | `auto` (default), `required`, or name one tool. The model authors the arguments in every case. |
| `max_tool_calls` | integer | Caps tool calls for the whole request. A route that cannot honour the cap fails before any tool runs. |
| `parallel_tool_calls` | boolean | Allow several independent tool calls in one model round. |
| `include` | array | Extra response fields to return, such as the sources a hosted web search used. |
| `context_management` | array | Official compaction configuration, for example `[{"type":"compaction","compact_threshold":100000}]`. |
| `metadata` | object | Your own key-value labels, carried on the response. |

<Callout tone="note" title="Strict by design">
	Requests are strict JSON: duplicate keys, malformed tool definitions,
	non-finite numbers and unknown fields fail before any model runs. The
	OpenAPI contract lists the complete accepted key set.
</Callout>

## The response object

A non-streaming call returns one `application/json` Responses object. Visible
text, refusals, function calls and hosted tool items keep their identity and
order.

```json
{
  "id": "resp_9f2c41d0a8",
  "object": "response",
  "created_at": 1789123456,
  "status": "completed",
  "model": "openai/gpt-5.5",
  "output": [
    {
      "type": "message",
      "role": "assistant",
      "content": [{ "type": "output_text", "text": "…", "annotations": [] }]
    }
  ],
  "usage": { "input_tokens": 128, "output_tokens": 96, "total_tokens": 224, "cached_tokens": 0 }
}
```

- `model` is the id you sent.
- `usage` counts tokens as the serving model counted them; `cached_tokens`
  reports the part served from the provider's prompt cache, when the model
  reports it.
- The official SDKs expose the assistant text as `output_text`; on the raw
  wire, read the text parts inside `output`.

## Streaming

`stream: true` returns server-sent events. Events are complete JSON values with
the official Responses event names.

- The stream ends exactly once, with `response.completed`,
  `response.incomplete`, `response.failed` or `response.cancelled`. EOF,
  `[DONE]` or an HTTP 200 alone is not a successful terminal.
- Hosted tool activity is an item on the way to that terminal: a search never
  ends the stream and never replaces the assistant answer.
- With `stream: false` you always get one JSON object; the content type never
  flips under you.

| Event | Carries |
| --- | --- |
| `response.created`, `response.in_progress` | The response shell, before output. |
| `response.output_item.added`, `.done` | Each output item as it opens and closes. |
| `response.content_part.added`, `.done` | Text, refusal or reasoning parts. |
| `response.output_text.delta`, `.done` | Assistant text as it is produced. |
| `response.function_call_arguments.delta`, `.done` | Arguments for a function call you will execute. |
| `response.web_search_call.in_progress`, `.searching`, `.completed` | Hosted web search progress. |
| `response.tool_search_call.in_progress`, `.completed` | Hosted tool discovery over the tools you declared. |
| `response.completed`, `.incomplete`, `.failed`, `.cancelled` | The single terminal, with its reason. |

Other official events, such as reasoning summaries, pass through under their
own names.

## Stored responses

With the default `store: true`, a completed response stays retrievable. That
is what makes continuation and safe retries possible.

- A completed response is retained for seven days from completion. Retries of
  the completion acknowledgement do not extend the window.
- `GET /responses/{id}` and `GET /responses/{id}/input_items` are
  tenant-isolated: another organization's key sees `404`.
- `DELETE /responses/{id}` returns the official
  `{"id": …, "object": "response", "deleted": true}` object.
- `POST /responses/{id}/cancel` exists as the official route, but background
  execution is not admitted: a stored hit returns `400 response_not_cancellable`
  and an unknown id `404`.
- A `previous_response_id` that points at an unpersisted response (`store:
  false`) is a `400`. Choose one shape per conversation.

## Compaction

`POST /responses/compact` compacts a stored conversation into an official
`response.compaction` object you can continue from, useful when a transcript
approaches the context window.

```bash
curl https://api.sylphx.ai/v1/responses/compact \
  -H "Authorization: Bearer $SYLPHX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "openai/gpt-5.5", "previous_response_id": "resp_9f2c41d0a8"}'
```

- Compact is unary: `stream: true` is rejected on this route.
- Sealed `encrypted_content` inside the result belongs to that model family and
  is not portable to another; continue on the same family.
- You can also let a create compact by sending `context_management`. Where a
  model cannot compact, the request fails with a typed error; there is no
  locally written summary.

## Files

Upload once, then reference the file id from an input item. File ids are
tenant-isolated and expire with the same seven-day window.

```bash
curl https://api.sylphx.ai/v1/files \
  -H "Authorization: Bearer $SYLPHX_API_KEY" \
  -F "file=@report.pdf" \
  -F "purpose=user_data"
```

```json
{
  "model": "openai/gpt-5.5",
  "input": [
    {
      "role": "user",
      "content": [
        { "type": "input_text", "text": "Summarise the attached report." },
        { "type": "input_file", "file_id": "file-3c9…" }
      ]
    }
  ]
}
```

`purpose` is one of `user_data` (the default), `vision`, `assistants`,
`fine-tune` or `evals`. `GET /files` lists your key's files with `first_id`,
`last_id` and `has_more`. A request that names a missing file fails with `400
file_not_found`; the API does not guess which file you meant.

## Messages: a second encoding

`POST /messages` accepts the Anthropic Messages format. It decodes under
Anthropic's rules, normalizes once into the same Responses document, and can
call any model in the catalogue. Hosted tools follow the same ownership: the
model initiates, the platform executes, and you observe server-tool blocks.
Streaming follows Anthropic's stream semantics with one typed terminal.

## Chat Completions: the OpenAI chat encoding

`POST /chat/completions` accepts the OpenAI Chat Completions format and
normalizes once into the same Responses document, so errors and usage are the
Responses ones.

- `messages`, function `tools`, `tool_choice`, `parallel_tool_calls`,
  `response_format` (including strict `json_schema`), `reasoning_effort` and
  `max_completion_tokens` map onto their Responses fields.
- A field the Responses document cannot carry faithfully is refused with `400
  invalid_request` before any model call — for example `n` above 1, a
  non-empty `stop`, `logprobs`, audio output or the deprecated `functions`
  shape. Nothing is dropped silently.
- Streams send `chat.completion.chunk` frames with exactly one `finish_reason`,
  a usage chunk when `stream_options.include_usage` is true, and `data:
  [DONE]`. Hosted tools stay on Responses and Messages.

## Switching models safely

Because `model` is a field, a conversation can move to another model on the
same endpoint. Keep the visible transcript and your function or tool results.
Remove execution observations and provider-sealed content (for example
`encrypted_content` from a compaction) before sending the transcript to a
different model family. If the new model cannot honour the transcript you get
a typed error, never a silently rewritten request. Read the new model's row in
[the catalogue](/ai/models) first: context window, maximum output and data
policy can differ.

<RelatedDocs
	links={[
		{
			href: '/docs/ai/tools',
			label: 'Hosted tools',
			description: 'web_search, web_fetch, tool_search and datetime.',
		},
		{
			href: '/docs/ai/errors',
			label: 'Errors and retries',
			description: 'The envelope, the codes and idempotent retries.',
		},
	]}
/>
