Skip to content
Console
Menu

Queues

Workflows

Getting Started

Authentication

KV Store

Responses API

The Responses endpoints, fields, streaming events, storage and compaction

The Responses API is the primary wire of Sylphx AI: the official OpenAI Responses document, served at https://api.sylphx.ai/v1. Every path below is relative to that base URL and takes your bearer key.

#Endpoints

MethodPathPurpose
POST/responsesCreate a response. Add stream: true for server-sent events.
POST/responses/compactCompact a stored response into an official compaction object.
GET/responses/{response_id}Retrieve a stored response while it is retained.
GET/responses/{response_id}/input_itemsList the input items stored for one response.
DELETE/responses/{response_id}Forget a stored response.
POST/responses/{response_id}/cancelThe official cancel route. Responses are not background jobs.
POST/chat/completionsChat Completions encoding, normalized into the same Responses document.
POST/messagesAnthropic Messages encoding, normalized into the same Responses document.
POST/filesMultipart upload. Mints a file-* id for input_file and input_image.
GET/filesList the files owned by your key.
GET/files/{file_id}File metadata, no bytes.
GET/files/{file_id}/contentRaw file bytes.
DELETE/files/{file_id}Delete a file you own.

#Create a response

POST /responses takes a strict JSON document. The only required field is model; input carries what you want answered.

Shell
curl https://api.sylphx.ai/v1/responses \
  -H "Authorization: Bearer $SYLPHX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "openai/gpt-5.5",
    "instructions": "Answer as a concise technical writer.",
    "input": "Summarise this incident report in three bullets.",
    "max_output_tokens": 400,
    "store": true
  }'
FieldTypeWhat it does
model (required)stringA model id from the catalogue. There is no router alias: an id we do not sell is a typed 404, never a silent substitution.
inputstring or arrayThe user turn, or a full item list. Replayed function calls pair with their outputs by call_id.
instructionsstringSystem-level guidance for this turn. Only text you write is sent; the platform never appends its own.
streambooleanfalse (the default) returns one JSON response; true returns text/event-stream.
storebooleanDefaults to true: the response is retained and can be retrieved or used as previous_response_id. false skips persistence.
previous_response_idstringContinue a stored conversation from an earlier response id.
max_output_tokensintegerUpper bound on the tokens this turn may generate.
toolsarrayFunction tools your code executes, plus hosted tools the platform executes.
tool_choicestring or objectauto (default), required, or name one tool. The model authors the arguments in every case.
max_tool_callsintegerCaps tool calls for the whole request. A route that cannot honour the cap fails before any tool runs.
parallel_tool_callsbooleanAllow several independent tool calls in one model round.
includearrayExtra response fields to return, such as the sources a hosted web search used.
context_managementarrayOfficial compaction configuration, for example [{"type":"compaction","compact_threshold":100000}].
metadataobjectYour own key-value labels, carried on the response.

Strict by design

Requests are strict JSON: duplicate keys, malformed tool definitions, non-finite numbers and unknown fields fail before any model runs. The OpenAPI contract lists the complete accepted key set.

#The response object

A non-streaming call returns one application/json Responses object. Visible text, refusals, function calls and hosted tool items keep their identity and order.

JSON
{
  "id": "resp_9f2c41d0a8",
  "object": "response",
  "created_at": 1789123456,
  "status": "completed",
  "model": "openai/gpt-5.5",
  "output": [
    {
      "type": "message",
      "role": "assistant",
      "content": [{ "type": "output_text", "text": "…", "annotations": [] }]
    }
  ],
  "usage": { "input_tokens": 128, "output_tokens": 96, "total_tokens": 224, "cached_tokens": 0 }
}
  • model is the id you sent.
  • usage counts tokens as the serving model counted them; cached_tokens reports the part served from the provider's prompt cache, when the model reports it.
  • The official SDKs expose the assistant text as output_text; on the raw wire, read the text parts inside output.

#Streaming

stream: true returns server-sent events. Events are complete JSON values with the official Responses event names.

  • The stream ends exactly once, with response.completed, response.incomplete, response.failed or response.cancelled. EOF, [DONE] or an HTTP 200 alone is not a successful terminal.
  • Hosted tool activity is an item on the way to that terminal: a search never ends the stream and never replaces the assistant answer.
  • With stream: false you always get one JSON object; the content type never flips under you.
EventCarries
response.created, response.in_progressThe response shell, before output.
response.output_item.added, .doneEach output item as it opens and closes.
response.content_part.added, .doneText, refusal or reasoning parts.
response.output_text.delta, .doneAssistant text as it is produced.
response.function_call_arguments.delta, .doneArguments for a function call you will execute.
response.web_search_call.in_progress, .searching, .completedHosted web search progress.
response.tool_search_call.in_progress, .completedHosted tool discovery over the tools you declared.
response.completed, .incomplete, .failed, .cancelledThe single terminal, with its reason.

Other official events, such as reasoning summaries, pass through under their own names.

#Stored responses

With the default store: true, a completed response stays retrievable. That is what makes continuation and safe retries possible.

  • A completed response is retained for seven days from completion. Retries of the completion acknowledgement do not extend the window.
  • GET /responses/{id} and GET /responses/{id}/input_items are tenant-isolated: another organization's key sees 404.
  • DELETE /responses/{id} returns the official {"id": …, "object": "response", "deleted": true} object.
  • POST /responses/{id}/cancel exists as the official route, but background execution is not admitted: a stored hit returns 400 response_not_cancellable and an unknown id 404.
  • A previous_response_id that points at an unpersisted response (store: false) is a 400. Choose one shape per conversation.

#Compaction

POST /responses/compact compacts a stored conversation into an official response.compaction object you can continue from, useful when a transcript approaches the context window.

Shell
curl https://api.sylphx.ai/v1/responses/compact \
  -H "Authorization: Bearer $SYLPHX_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "openai/gpt-5.5", "previous_response_id": "resp_9f2c41d0a8"}'
  • Compact is unary: stream: true is rejected on this route.
  • Sealed encrypted_content inside the result belongs to that model family and is not portable to another; continue on the same family.
  • You can also let a create compact by sending context_management. Where a model cannot compact, the request fails with a typed error; there is no locally written summary.

#Files

Upload once, then reference the file id from an input item. File ids are tenant-isolated and expire with the same seven-day window.

Shell
curl https://api.sylphx.ai/v1/files \
  -H "Authorization: Bearer $SYLPHX_API_KEY" \
  -F "file=@report.pdf" \
  -F "purpose=user_data"
JSON
{
  "model": "openai/gpt-5.5",
  "input": [
    {
      "role": "user",
      "content": [
        { "type": "input_text", "text": "Summarise the attached report." },
        { "type": "input_file", "file_id": "file-3c9…" }
      ]
    }
  ]
}

purpose is one of user_data (the default), vision, assistants, fine-tune or evals. GET /files lists your key's files with first_id, last_id and has_more. A request that names a missing file fails with 400 file_not_found; the API does not guess which file you meant.

#Messages: a second encoding

POST /messages accepts the Anthropic Messages format. It decodes under Anthropic's rules, normalizes once into the same Responses document, and can call any model in the catalogue. Hosted tools follow the same ownership: the model initiates, the platform executes, and you observe server-tool blocks. Streaming follows Anthropic's stream semantics with one typed terminal.

#Chat Completions: the OpenAI chat encoding

POST /chat/completions accepts the OpenAI Chat Completions format and normalizes once into the same Responses document, so errors and usage are the Responses ones.

  • messages, function tools, tool_choice, parallel_tool_calls, response_format (including strict json_schema), reasoning_effort and max_completion_tokens map onto their Responses fields.
  • A field the Responses document cannot carry faithfully is refused with 400 invalid_request before any model call — for example n above 1, a non-empty stop, logprobs, audio output or the deprecated functions shape. Nothing is dropped silently.
  • Streams send chat.completion.chunk frames with exactly one finish_reason, a usage chunk when stream_options.include_usage is true, and data: [DONE]. Hosted tools stay on Responses and Messages.

#Switching models safely

Because model is a field, a conversation can move to another model on the same endpoint. Keep the visible transcript and your function or tool results. Remove execution observations and provider-sealed content (for example encrypted_content from a compaction) before sending the transcript to a different model family. If the new model cannot honour the transcript you get a typed error, never a silently rewritten request. Read the new model's row in the catalogue first: context window, maximum output and data policy can differ.