What every deployment needs
The three things an orchestrator — or a person with a shell — needs from nanocached before any of the shapes in the deployment guides: probes to route around instances that aren't ready, metrics to scale on, and a SIGTERM-driven drain so scale-in never loses data.
The operations endpoint
Each binary takes --metrics-port <port> and serves
three paths on it (plain HTTP, on --host):
| Path | What it returns |
|---|---|
GET /metrics | Prometheus text format. |
GET /healthz | 200 while the process is serving — a liveness probe. |
GET /readyz | 200 only when the instance can do useful work — a readiness probe. A node is ready once it has adopted cluster membership; a proxy once it has a node roster; a discovery server once its startup grace period is over. |
The endpoint is unauthenticated by design (probes and scrapers don't do protocol handshakes) — bind it to an internal interface and don't expose it outside the task network. Omitting the flag disables the endpoint entirely.
Metrics to scale on
Nodes scale on the memory watermark. Sum
nanocached_node_memory_used_bytes over the fleet, divide by
the summed nanocached_node_memory_max_bytes, and target a
utilization (say 70%) with your autoscaler. Entry counts, hit/miss/set/
delete/eviction/expiration counters, connection counts
(nanocached_node_connections against
…_connections_max, the --max-connections bound
— raise --max-connections-per-ip too when clients share a
source IP, as behind NAT or on Kubernetes), and per-namespace
breakdowns (nanocached_node_namespace_used_bytes,
…_namespace_entries) are exported alongside for dashboards
and alerting.
Proxies scale on connections:
nanocached_proxy_client_connections against
…_client_connections_max (plus CPU).
…_requests_total and …_upstream_failures_total
give request and error rates.
Discovery exports roster sizes
(nanocached_discovery_members, …_proxies,
…_waiting_nodes, …_joining_nodes) and join
counters — not scaling signals, but the first thing to graph when a
scale event misbehaves.
Graceful scale-in
Scale-in terminates instances that are not dead, and orchestrators signal that with SIGTERM followed by SIGKILL after a grace period. Both server binaries treat SIGTERM as "drain, then exit" — so ordinary deployments, autoscaler scale-in, and spot/Spot reclaim all get the graceful path with no extra hooks.
Node: decommission
On SIGTERM a clustered node runs a planned leave within
--drain-timeout seconds (default 25):
- It hands every key it owns to the node that becomes the new owner once it's gone — so even at R=1 the only copy moves rather than dies.
- It leaves the cluster membership immediately (no liveness timeout to wait out), so clients refresh onto the surviving nodes.
- It keeps serving and forwards writes it no longer owns for the rest of the budget while clients catch up, then exits 0.
--drain-timeout 0 skips the handoff (the old
crash-equivalent shutdown) — acceptable for fast restarts at
R ≥ 2, where replicas cover the gap.
Scale in one node at a time. A draining node plans its entire handoff from one roster snapshot taken as the drain begins. Two nodes draining at once each plan against a roster that still lists the other, so a key whose owners included both can come out of the double-leave one replica short until a later membership change, a rewrite, or its TTL turns it over — and at R=1, a handoff that lands on the other leaver can be dropped outright (see Architecture on the self-healing consistency model). Scale-in policies that step the desired count down one instance per action avoid the window entirely; a larger single step (or a spot reclaim of several instances at once) trades that safety for speed and should be reserved for R ≥ 2.
Proxy: drain
On SIGTERM a proxy deregisters from discovery's proxy roster
immediately (new clients stop being handed its address), finishes
in-flight requests within --drain-timeout (default 25),
and exits 0. Clients connected to it reconnect to a surviving proxy
transparently — at most one reconnect per connection. Where the
proxy tier runs and how many to run is its own page.
Sizing the stop grace
The orchestrator's grace period must outlast the drain:
set ECS stopTimeout / Kubernetes
terminationGracePeriodSeconds to at least
--drain-timeout plus a margin (twice it is a
comfortable rule). A node's handoff time scales with the data it holds
— a 20 s budget drains hundreds of MB comfortably on an
in-network transfer, but measure with your value sizes; keys that miss
the budget are logged and left to the surviving replicas. The default
25 s fits inside a spot interruption's 2-minute warning with room
to spare.
Verifying a deployment
The repository ships runnable scale-in scenarios under
tests/e2e:
scalein.sh (a node leaves under load at R=2 — zero loss,
zero corruption, membership follows) and proxydrain.sh (a
proxy drains under via-proxy load — immediate roster removal, exit 0,
clients reconnect). They document the exact guarantees the drain makes
and are the fastest way to see the choreography before wiring up an
autoscaler.