Menu
Platform
AI
App store purchases
Database
Flags
Jobs and cron
Localization
Monitoring
Notifications
Payments
Queues
Sandboxes
Webhooks
Getting Started
Authentication
KV Store
Deploy & Infrastructure
Reference
Responses API
The Responses endpoints, fields, streaming events, storage and compaction
The Responses API is the primary wire of Sylphx AI: the official OpenAI
Responses document, served at https://api.sylphx.ai/v1. Every path below is
relative to that base URL and takes your bearer key.
#Endpoints
| Method | Path | Purpose |
|---|---|---|
POST | /responses | Create a response. Add stream: true for server-sent events. |
POST | /responses/compact | Compact a stored response into an official compaction object. |
GET | /responses/{response_id} | Retrieve a stored response while it is retained. |
GET | /responses/{response_id}/input_items | List the input items stored for one response. |
DELETE | /responses/{response_id} | Forget a stored response. |
POST | /responses/{response_id}/cancel | The official cancel route. Responses are not background jobs. |
POST | /chat/completions | Chat Completions encoding, normalized into the same Responses document. |
POST | /messages | Anthropic Messages encoding, normalized into the same Responses document. |
POST | /files | Multipart upload. Mints a file-* id for input_file and input_image. |
GET | /files | List the files owned by your key. |
GET | /files/{file_id} | File metadata, no bytes. |
GET | /files/{file_id}/content | Raw file bytes. |
DELETE | /files/{file_id} | Delete a file you own. |
#Create a response
POST /responses takes a strict JSON document. The only required field is
model; input carries what you want answered.
curl https://api.sylphx.ai/v1/responses \
-H "Authorization: Bearer $SYLPHX_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "openai/gpt-5.5",
"instructions": "Answer as a concise technical writer.",
"input": "Summarise this incident report in three bullets.",
"max_output_tokens": 400,
"store": true
}'| Field | Type | What it does |
|---|---|---|
model (required) | string | A model id from the catalogue. There is no router alias: an id we do not sell is a typed 404, never a silent substitution. |
input | string or array | The user turn, or a full item list. Replayed function calls pair with their outputs by call_id. |
instructions | string | System-level guidance for this turn. Only text you write is sent; the platform never appends its own. |
stream | boolean | false (the default) returns one JSON response; true returns text/event-stream. |
store | boolean | Defaults to true: the response is retained and can be retrieved or used as previous_response_id. false skips persistence. |
previous_response_id | string | Continue a stored conversation from an earlier response id. |
max_output_tokens | integer | Upper bound on the tokens this turn may generate. |
tools | array | Function tools your code executes, plus hosted tools the platform executes. |
tool_choice | string or object | auto (default), required, or name one tool. The model authors the arguments in every case. |
max_tool_calls | integer | Caps tool calls for the whole request. A route that cannot honour the cap fails before any tool runs. |
parallel_tool_calls | boolean | Allow several independent tool calls in one model round. |
include | array | Extra response fields to return, such as the sources a hosted web search used. |
context_management | array | Official compaction configuration, for example [{"type":"compaction","compact_threshold":100000}]. |
metadata | object | Your own key-value labels, carried on the response. |
Strict by design
Requests are strict JSON: duplicate keys, malformed tool definitions, non-finite numbers and unknown fields fail before any model runs. The OpenAPI contract lists the complete accepted key set.
#The response object
A non-streaming call returns one application/json Responses object. Visible
text, refusals, function calls and hosted tool items keep their identity and
order.
{
"id": "resp_9f2c41d0a8",
"object": "response",
"created_at": 1789123456,
"status": "completed",
"model": "openai/gpt-5.5",
"output": [
{
"type": "message",
"role": "assistant",
"content": [{ "type": "output_text", "text": "…", "annotations": [] }]
}
],
"usage": { "input_tokens": 128, "output_tokens": 96, "total_tokens": 224, "cached_tokens": 0 }
}modelis the id you sent.usagecounts tokens as the serving model counted them;cached_tokensreports the part served from the provider's prompt cache, when the model reports it.- The official SDKs expose the assistant text as
output_text; on the raw wire, read the text parts insideoutput.
#Streaming
stream: true returns server-sent events. Events are complete JSON values with
the official Responses event names.
- The stream ends exactly once, with
response.completed,response.incomplete,response.failedorresponse.cancelled. EOF,[DONE]or an HTTP 200 alone is not a successful terminal. - Hosted tool activity is an item on the way to that terminal: a search never ends the stream and never replaces the assistant answer.
- With
stream: falseyou always get one JSON object; the content type never flips under you.
| Event | Carries |
|---|---|
response.created, response.in_progress | The response shell, before output. |
response.output_item.added, .done | Each output item as it opens and closes. |
response.content_part.added, .done | Text, refusal or reasoning parts. |
response.output_text.delta, .done | Assistant text as it is produced. |
response.function_call_arguments.delta, .done | Arguments for a function call you will execute. |
response.web_search_call.in_progress, .searching, .completed | Hosted web search progress. |
response.tool_search_call.in_progress, .completed | Hosted tool discovery over the tools you declared. |
response.completed, .incomplete, .failed, .cancelled | The single terminal, with its reason. |
Other official events, such as reasoning summaries, pass through under their own names.
#Stored responses
With the default store: true, a completed response stays retrievable. That
is what makes continuation and safe retries possible.
- A completed response is retained for seven days from completion. Retries of the completion acknowledgement do not extend the window.
GET /responses/{id}andGET /responses/{id}/input_itemsare tenant-isolated: another organization's key sees404.DELETE /responses/{id}returns the official{"id": …, "object": "response", "deleted": true}object.POST /responses/{id}/cancelexists as the official route, but background execution is not admitted: a stored hit returns400 response_not_cancellableand an unknown id404.- A
previous_response_idthat points at an unpersisted response (store: false) is a400. Choose one shape per conversation.
#Compaction
POST /responses/compact compacts a stored conversation into an official
response.compaction object you can continue from, useful when a transcript
approaches the context window.
curl https://api.sylphx.ai/v1/responses/compact \
-H "Authorization: Bearer $SYLPHX_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "openai/gpt-5.5", "previous_response_id": "resp_9f2c41d0a8"}'- Compact is unary:
stream: trueis rejected on this route. - Sealed
encrypted_contentinside the result belongs to that model family and is not portable to another; continue on the same family. - You can also let a create compact by sending
context_management. Where a model cannot compact, the request fails with a typed error; there is no locally written summary.
#Files
Upload once, then reference the file id from an input item. File ids are tenant-isolated and expire with the same seven-day window.
curl https://api.sylphx.ai/v1/files \
-H "Authorization: Bearer $SYLPHX_API_KEY" \
-F "file=@report.pdf" \
-F "purpose=user_data"{
"model": "openai/gpt-5.5",
"input": [
{
"role": "user",
"content": [
{ "type": "input_text", "text": "Summarise the attached report." },
{ "type": "input_file", "file_id": "file-3c9…" }
]
}
]
}purpose is one of user_data (the default), vision, assistants,
fine-tune or evals. GET /files lists your key's files with first_id,
last_id and has_more. A request that names a missing file fails with 400 file_not_found; the API does not guess which file you meant.
#Messages: a second encoding
POST /messages accepts the Anthropic Messages format. It decodes under
Anthropic's rules, normalizes once into the same Responses document, and can
call any model in the catalogue. Hosted tools follow the same ownership: the
model initiates, the platform executes, and you observe server-tool blocks.
Streaming follows Anthropic's stream semantics with one typed terminal.
#Chat Completions: the OpenAI chat encoding
POST /chat/completions accepts the OpenAI Chat Completions format and
normalizes once into the same Responses document, so errors and usage are the
Responses ones.
messages, functiontools,tool_choice,parallel_tool_calls,response_format(including strictjson_schema),reasoning_effortandmax_completion_tokensmap onto their Responses fields.- A field the Responses document cannot carry faithfully is refused with
400 invalid_requestbefore any model call — for examplenabove 1, a non-emptystop,logprobs, audio output or the deprecatedfunctionsshape. Nothing is dropped silently. - Streams send
chat.completion.chunkframes with exactly onefinish_reason, a usage chunk whenstream_options.include_usageis true, anddata: [DONE]. Hosted tools stay on Responses and Messages.
#Switching models safely
Because model is a field, a conversation can move to another model on the
same endpoint. Keep the visible transcript and your function or tool results.
Remove execution observations and provider-sealed content (for example
encrypted_content from a compaction) before sending the transcript to a
different model family. If the new model cannot honour the transcript you get
a typed error, never a silently rewritten request. Read the new model's row in
the catalogue first: context window, maximum output and data
policy can differ.