---
title: Health checks
description: How an instance earns traffic, and how to serve a /healthz that says something useful.
type: how-to
product: hosting
summary: The probe a Service configures, and the one line that answers it well.
updated: 2026-10-01
order: 2
---

<Callout tone="note" title="Today this runs on the management API">
Probe settings are part of the service, and you change them with `PATCH /v1/projects/{id}/services/{name}` on the management API. The `sylphx hosting services update` command below shows the resource-model spelling of the same setting. See the [Hosting quickstart](/docs/hosting/quickstart).
</Callout>

A health check is the platform asking an instance whether it is ready. It is
also the fastest way to tell a bad deploy from a slow one, so it is worth
making the answer honest.

## What the Service configures

<PropertyTable
	properties={[
		{
			name: 'health.protocol',
			type: 'HealthProtocol',
			description: 'http or tcp. Required.',
		},
		{
			name: 'health.path',
			type: 'string',
			description: 'The HTTP path to probe when the protocol is HTTP.',
		},
		{
			name: 'health.startup_timeout',
			type: 'duration',
			description: 'How long a new instance may take to pass its first probe.',
		},
	]}
/>

```bash
sylphx hosting services update orgs/acme/projects/shop/envs/production/services/shop \
  --spec.health.protocol http \
  --spec.health.path /healthz \
  --spec.health.startup-timeout 30s
```

A `tcp` probe asks only whether something is listening — right for a worker
with no HTTP surface, and it cannot tell a booting process from a serving one.
An `http` probe waits for the path to answer successfully, which is what makes
a rollout safe: an instance that never passes its probe does not take traffic,
so a wave does not advance on an instance that is not ready.

`startup_timeout` is the tolerance for a slow first boot: a JIT warming up, a
connection pool filling, a migration check. Set it to something real — the
time the slowest legitimate start takes, not a round number.

<Callout tone="warning" title="Raise the timeout for the right reason">
A timeout that keeps needing to be raised is usually an instance that is
failing rather than slow. The probe's failures are evidence the rollout reads;
a longer timeout only delays the answer.
</Callout>

## Serve a healthz worth reading

`@sylphx/apps/health` is the package that answers `/healthz` for you. Mount
one middleware and the endpoint exists:

```bash
bun add @sylphx/apps
```

<CodeTabs>
	<CodeTab
		label="Hono"
		language="ts"
		code={`import { Hono } from 'hono'
import { withHealth } from '@sylphx/apps/health'

const app = new Hono()

// One line — every service gets a multi-signal /healthz
app.use('*', withHealth.hono())

app.get('/', (c) => c.text('hello world'))

export default app`}
	/>
	<CodeTab
		label="Express"
		language="ts"
		code={`import express from 'express'
import { withHealth } from '@sylphx/apps/health'

const app = express()
app.use(withHealth.express())

app.get('/', (req, res) => res.send('hello world'))

app.listen(3000)`}
	/>
	<CodeTab
		label="Fastify"
		language="ts"
		code={`import Fastify from 'fastify'
import { withHealth } from '@sylphx/apps/health'

const app = Fastify()
await app.register(withHealth.fastify())

app.get('/', async () => 'hello world')

await app.listen({ port: 3000 })`}
	/>
	<CodeTab
		label="Next.js"
		language="ts"
		code={`// app/healthz/route.ts
import { withHealth } from '@sylphx/apps/health'

const handler = withHealth.fetch()

export function GET(req: Request): Promise<Response> {
  return handler(req)
}`}
	/>
</CodeTabs>

The path defaults to `/healthz` and is configurable, so a Service's
`health.path` and the middleware's path are one setting, not two.

## What it answers

```json
{
  "score": 0.92,
  "signals": {
    "eventLoopLagMs": 12,
    "database": "ok",
    "redis": "ok",
    "errorRate": 0.001,
    "memoryPressure": 0.45
  },
  "signalFactors": {
    "eventLoopLagMs": 1,
    "database": 1,
    "redis": 1,
    "errorRate": 0.99,
    "memoryPressure": 0.7
  },
  "lastTickAt": "2026-05-20T12:34:56.789Z"
}
```

The score is one number in `[0, 1]` composed across the signals, so a probe
answer is a decision rather than a boolean guess. The HTTP status follows the
same number: `200` while the score is at or above the liveness floor, `503`
below it — which is exactly what an `http` probe needs to read.

What the score means downstream:

| Score | Direct probe | Sidecar | What happens |
| --- | --- | --- | --- |
| above 0.8 | 200 | 200 | Normal traffic |
| 0.5 to 0.8 | 200 | 503 | Drains new traffic, does not kill the instance |
| at or below 0.5 | 503 | 503 | Kill after the threshold |

A single number that moves through those bands is what stops flapping: an
instance that is merely unhappy drains, and only one that is failing is
replaced.

## Compose the signals it should know about

With no configuration the endpoint reports event-loop lag. A service that
knows more should say more:

```ts
import { sylphxHealth } from '@sylphx/apps/health'
import {
  eventLoopLagSignal,
  databaseSignal,
  redisSignal,
  errorRateSignal,
  memoryPressureSignal,
} from '@sylphx/apps/health'

const errors = errorRateSignal({ window: '5s', degradedRate: 0.05 })

const health = sylphxHealth({
  signals: [
    eventLoopLagSignal({ degradedMs: 5000, deadMs: 30000 }),
    databaseSignal({ ping: () => pool.query('SELECT 1') }),
    redisSignal({ ping: () => redis.ping() }),
    errors,
    memoryPressureSignal({ degradedRatio: 0.85 }),
  ],
})

// Track requests for the error-rate signal
app.use(async (c, next) => {
  try {
    await next()
    errors.recordSuccess()
  } catch (err) {
    errors.recordError()
    throw err
  }
})
```

The seven built-in signals are event-loop lag, queue depth, error rate,
memory pressure, a generic ping, database and cache. A database that is down
is then visible in the probe — which is the point: the probe answers for the
things the service needs, not just for its process.

<Callout tone="note" title="A dependency has its own health">
A `databaseSignal` failing while the event loop is fine says the fault is
downstream. Raising `startup_timeout`, or replacing the instance, does not fix
it — read the signals before treating the probe as noise.
</Callout>

## Where the scores go

Every evaluation emits OpenTelemetry metrics — `health.score`,
`health.signal.factor`, `health.signal.value` (and `health.signal.value_state`
for string-valued signals) — sets `service.health.score` on the active span,
and records a `health.evaluated` event. Point an OTel collector at the
service and the scores are in the same pipeline as the traces, with no
scraping endpoint to configure.

The same evaluation can be recorded into a tamper-evident history: a chain in
which each entry carries the hash of the one before it, so a history can be
exported and verified afterwards rather than trusted. `createHistoryRecorder`
records it, and `verifyChain` checks an exported chain end to end and reports
the first break if there is one.

## Before you ship

- Give the Service an `http` probe on a path that answers quickly and does not
  do work: a probe that queries three services is a probe that fails three
  ways.
- Set `startup_timeout` from a real measurement of the slowest honest boot.
- Compose the signals for the dependencies that matter, and let the probe
  answer for them.
- Keep the endpoint unauthenticated inside the service and unreachable from the
  public edge — the platform probes it, callers do not.

<RelatedDocs
	links={[
		{
			href: '/docs/hosting/configuration',
			label: 'Configuration',
			description: 'The probe is one field of the Service spec.',
		},
		{
			href: '/docs/hosting/deploy',
			label: 'Deploys',
			description: 'How probe evidence gates each wave.',
		},
		{
			href: '/docs/api/services/update',
			label: 'services.update',
			description: 'Every field, with examples.',
		},
	]}
/>
