Menu
Platform
AI
App store purchases
Database
Flags
Jobs and cron
Localization
Monitoring
Notifications
Payments
Queues
Sandboxes
Webhooks
Getting Started
Authentication
KV Store
Deploy & Infrastructure
Reference
Indexing documents
The three calls that put a document in an index, take it back and remove it.
A document lives in one index, at one address, and the two ids in the URL are what make that address unique.
| Method | Path | What it does |
|---|---|---|
PUT | /v1/documents/{index_id}/{document_id} | Writes a document to a search index, replacing the current version. |
GET | /v1/documents/{index_id}/{document_id} | Reads a document. |
DELETE | /v1/documents/{index_id}/{document_id} | Deletes a document. |
All three are served at https://api.data.sylphx.com. index_id is the id
segment of the index's name in the environment, and document_id is the
document's own id; there is no CLI command for a document, so these are API
calls. PUT and DELETE need data:write; GET needs data:read.
#The document id is yours
Nothing assigns a document id. You write one, and writing the same id again replaces what is at that address, so the id is the part to get right: it is the address the index retrieves the document by, and the address a delete names. An id you choose is also the only key you need — a document's contents can change under a stable id without anything else knowing.
#Writing one
curl -X PUT "https://api.data.sylphx.com/v1/documents/index-id/document-id" \
-H "Authorization: Bearer $SYLPHX_API_KEY" \
-H "Content-Type: application/json" \
-d '{}'Three fields carry the work:
| Field | Type | What it is |
|---|---|---|
document_json | bytes | The document as JSON, base64; at most 1 MiB. |
vector | double[] | The document’s embedding, at most 4096 numbers; the index’s vector dimensions when it has them. |
expected_version | uint64 | Write only when the current version equals this one (0: any). |
The write is a replace, not a merge: what the address held is gone, and
version moves on to the next one. That is why expected_version earns its
place when two writers can reach the same document — a write that names the
version it read is refused instead of quietly overwriting a version someone
else wrote since.
The answer is the stored document:
| Field | Type | What it is |
|---|---|---|
document_json | bytes | The document as JSON, base64. |
vector | double[] | The document’s embedding; its length is the index’s vector dimensions. |
sha256 | string | `sha256:<hex>` of the document. |
version | uint64 | The version, increasing per document. |
updated_at_unix_ms | int64 | When this version was written, in Unix milliseconds. |
#Embeddings are yours to supply
Data never invents embeddings. In an index with dimensions, a document carries the embedding you wrote with it, and a vector query matches against that vector rather than against anything the index derived. So the model that produces the embeddings is a decision you keep: an embedding written by one model and a query embedded by another are two different spaces, and the nearness between them means nothing.
#Reading one back
curl "https://api.data.sylphx.com/v1/documents/index-id/document-id" \
-H "Authorization: Bearer $SYLPHX_API_KEY"The answer is the whole stored record, version included. A read is what a
conditional write starts from: take version from here, and pass it as
expected_version there.
#Deleting one
curl -X DELETE "https://api.data.sylphx.com/v1/documents/index-id/document-id" \
-H "Authorization: Bearer $SYLPHX_API_KEY"The answer's deleted says whether the document existed, and the query takes
expected_version for the same reason the write does. A document delete
removes that document and nothing else — the index keeps its name, its
settings and its other documents, and deleting the index is a separate call.