Skip to content
Console
Menu

Queues

Workflows

Getting Started

Authentication

KV Store

Health checks

The probe a Service configures, and the one line that answers it well.

Today this runs on the management API

Probe settings are part of the service, and you change them with PATCH /v1/projects/{id}/services/{name} on the management API. The sylphx hosting services update command below shows the resource-model spelling of the same setting. See the Hosting quickstart.

A health check is the platform asking an instance whether it is ready. It is also the fastest way to tell a bad deploy from a slow one, so it is worth making the answer honest.

#What the Service configures

FieldTypeWhat it is
health.protocolHealthProtocolhttp or tcp. Required.
health.pathstringThe HTTP path to probe when the protocol is HTTP.
health.startup_timeoutdurationHow long a new instance may take to pass its first probe.
Shell
sylphx hosting services update orgs/acme/projects/shop/envs/production/services/shop \
  --spec.health.protocol http \
  --spec.health.path /healthz \
  --spec.health.startup-timeout 30s

A tcp probe asks only whether something is listening — right for a worker with no HTTP surface, and it cannot tell a booting process from a serving one. An http probe waits for the path to answer successfully, which is what makes a rollout safe: an instance that never passes its probe does not take traffic, so a wave does not advance on an instance that is not ready.

startup_timeout is the tolerance for a slow first boot: a JIT warming up, a connection pool filling, a migration check. Set it to something real — the time the slowest legitimate start takes, not a round number.

Raise the timeout for the right reason

A timeout that keeps needing to be raised is usually an instance that is failing rather than slow. The probe's failures are evidence the rollout reads; a longer timeout only delays the answer.

#Serve a healthz worth reading

@sylphx/apps/health is the package that answers /healthz for you. Mount one middleware and the endpoint exists:

Shell
bun add @sylphx/apps
import { Hono } from 'hono'
import { withHealth } from '@sylphx/apps/health'

const app = new Hono()

// One line — every service gets a multi-signal /healthz
app.use('*', withHealth.hono())

app.get('/', (c) => c.text('hello world'))

export default app

The path defaults to /healthz and is configurable, so a Service's health.path and the middleware's path are one setting, not two.

#What it answers

JSON
{
  "score": 0.92,
  "signals": {
    "eventLoopLagMs": 12,
    "database": "ok",
    "redis": "ok",
    "errorRate": 0.001,
    "memoryPressure": 0.45
  },
  "signalFactors": {
    "eventLoopLagMs": 1,
    "database": 1,
    "redis": 1,
    "errorRate": 0.99,
    "memoryPressure": 0.7
  },
  "lastTickAt": "2026-05-20T12:34:56.789Z"
}

The score is one number in [0, 1] composed across the signals, so a probe answer is a decision rather than a boolean guess. The HTTP status follows the same number: 200 while the score is at or above the liveness floor, 503 below it — which is exactly what an http probe needs to read.

What the score means downstream:

ScoreDirect probeSidecarWhat happens
above 0.8200200Normal traffic
0.5 to 0.8200503Drains new traffic, does not kill the instance
at or below 0.5503503Kill after the threshold

A single number that moves through those bands is what stops flapping: an instance that is merely unhappy drains, and only one that is failing is replaced.

#Compose the signals it should know about

With no configuration the endpoint reports event-loop lag. A service that knows more should say more:

TypeScript
import { sylphxHealth } from '@sylphx/apps/health'
import {
  eventLoopLagSignal,
  databaseSignal,
  redisSignal,
  errorRateSignal,
  memoryPressureSignal,
} from '@sylphx/apps/health'

const errors = errorRateSignal({ window: '5s', degradedRate: 0.05 })

const health = sylphxHealth({
  signals: [
    eventLoopLagSignal({ degradedMs: 5000, deadMs: 30000 }),
    databaseSignal({ ping: () => pool.query('SELECT 1') }),
    redisSignal({ ping: () => redis.ping() }),
    errors,
    memoryPressureSignal({ degradedRatio: 0.85 }),
  ],
})

// Track requests for the error-rate signal
app.use(async (c, next) => {
  try {
    await next()
    errors.recordSuccess()
  } catch (err) {
    errors.recordError()
    throw err
  }
})

The seven built-in signals are event-loop lag, queue depth, error rate, memory pressure, a generic ping, database and cache. A database that is down is then visible in the probe — which is the point: the probe answers for the things the service needs, not just for its process.

A dependency has its own health

A databaseSignal failing while the event loop is fine says the fault is downstream. Raising startup_timeout, or replacing the instance, does not fix it — read the signals before treating the probe as noise.

#Where the scores go

Every evaluation emits OpenTelemetry metrics — health.score, health.signal.factor, health.signal.value (and health.signal.value_state for string-valued signals) — sets service.health.score on the active span, and records a health.evaluated event. Point an OTel collector at the service and the scores are in the same pipeline as the traces, with no scraping endpoint to configure.

The same evaluation can be recorded into a tamper-evident history: a chain in which each entry carries the hash of the one before it, so a history can be exported and verified afterwards rather than trusted. createHistoryRecorder records it, and verifyChain checks an exported chain end to end and reports the first break if there is one.

#Before you ship

  • Give the Service an http probe on a path that answers quickly and does not do work: a probe that queries three services is a probe that fails three ways.
  • Set startup_timeout from a real measurement of the slowest honest boot.
  • Compose the signals for the dependencies that matter, and let the probe answer for them.
  • Keep the endpoint unauthenticated inside the service and unreachable from the public edge — the platform probes it, callers do not.