What every deployment needs

The three things an orchestrator — or a person with a shell — needs from nanocached before any of the shapes in the deployment guides: probes to route around instances that aren't ready, metrics to scale on, and a SIGTERM-driven drain so scale-in never loses data.

The operations endpoint

Each binary takes --metrics-port <port> and serves three paths on it (plain HTTP, on --host):

PathWhat it returns
GET /metricsPrometheus text format.
GET /healthz200 while the process is serving — a liveness probe.
GET /readyz200 only when the instance can do useful work — a readiness probe. A node is ready once it has adopted cluster membership; a proxy once it has a node roster; a discovery server once its startup grace period is over.

The endpoint is unauthenticated by design (probes and scrapers don't do protocol handshakes) — bind it to an internal interface and don't expose it outside the task network. Omitting the flag disables the endpoint entirely.

Metrics to scale on

Nodes scale on the memory watermark. Sum nanocached_node_memory_used_bytes over the fleet, divide by the summed nanocached_node_memory_max_bytes, and target a utilization (say 70%) with your autoscaler. Entry counts, hit/miss/set/ delete/eviction/expiration counters, connection counts (nanocached_node_connections against …_connections_max, the --max-connections bound — raise --max-connections-per-ip too when clients share a source IP, as behind NAT or on Kubernetes), and per-namespace breakdowns (nanocached_node_namespace_used_bytes, …_namespace_entries) are exported alongside for dashboards and alerting.

Proxies scale on connections: nanocached_proxy_client_connections against …_client_connections_max (plus CPU). …_requests_total and …_upstream_failures_total give request and error rates.

Discovery exports roster sizes (nanocached_discovery_members, …_proxies, …_waiting_nodes, …_joining_nodes) and join counters — not scaling signals, but the first thing to graph when a scale event misbehaves.

Graceful scale-in

Scale-in terminates instances that are not dead, and orchestrators signal that with SIGTERM followed by SIGKILL after a grace period. Both server binaries treat SIGTERM as "drain, then exit" — so ordinary deployments, autoscaler scale-in, and spot/Spot reclaim all get the graceful path with no extra hooks.

Node: decommission

On SIGTERM a clustered node runs a planned leave within --drain-timeout seconds (default 25):

  1. It hands every key it owns to the node that becomes the new owner once it's gone — so even at R=1 the only copy moves rather than dies.
  2. It leaves the cluster membership immediately (no liveness timeout to wait out), so clients refresh onto the surviving nodes.
  3. It keeps serving and forwards writes it no longer owns for the rest of the budget while clients catch up, then exits 0.

--drain-timeout 0 skips the handoff (the old crash-equivalent shutdown) — acceptable for fast restarts at R ≥ 2, where replicas cover the gap.

Scale in one node at a time. A draining node plans its entire handoff from one roster snapshot taken as the drain begins. Two nodes draining at once each plan against a roster that still lists the other, so a key whose owners included both can come out of the double-leave one replica short until a later membership change, a rewrite, or its TTL turns it over — and at R=1, a handoff that lands on the other leaver can be dropped outright (see Architecture on the self-healing consistency model). Scale-in policies that step the desired count down one instance per action avoid the window entirely; a larger single step (or a spot reclaim of several instances at once) trades that safety for speed and should be reserved for R ≥ 2.

Proxy: drain

On SIGTERM a proxy deregisters from discovery's proxy roster immediately (new clients stop being handed its address), finishes in-flight requests within --drain-timeout (default 25), and exits 0. Clients connected to it reconnect to a surviving proxy transparently — at most one reconnect per connection. Where the proxy tier runs and how many to run is its own page.

Sizing the stop grace

The orchestrator's grace period must outlast the drain: set ECS stopTimeout / Kubernetes terminationGracePeriodSeconds to at least --drain-timeout plus a margin (twice it is a comfortable rule). A node's handoff time scales with the data it holds — a 20 s budget drains hundreds of MB comfortably on an in-network transfer, but measure with your value sizes; keys that miss the budget are logged and left to the surviving replicas. The default 25 s fits inside a spot interruption's 2-minute warning with room to spare.

Verifying a deployment

The repository ships runnable scale-in scenarios under tests/e2e: scalein.sh (a node leaves under load at R=2 — zero loss, zero corruption, membership follows) and proxydrain.sh (a proxy drains under via-proxy load — immediate roster removal, exit 0, clients reconnect). They document the exact guarantees the drain makes and are the fastest way to see the choreography before wiring up an autoscaler.