Cluster on EKS (multi-AZ)

The same cluster as the ECS guide — one discovery and a node Deployment across two zones — expressed in Kubernetes objects, on an EKS cluster that eksctl builds from one config file (VPC, subnets in two zones, security groups, IAM roles, control plane, a managed node group). Discovery's stable address is a ClusterIP Service; the node Deployment grows and shrinks by its replica count, and the drain on the way down is the SIGTERM Kubernetes already sends. This page builds the cluster and grows it by one pod; a HorizontalPodAutoscaler on top of the same Deployment is how you would tie the count to a metric, and is not covered here. Every command was run in this order against a real account and torn down with the commands at the end.

The cluster

eks-cluster.yaml:

apiVersion: eksctl.io/v1alpha5
kind: ClusterConfig
metadata:
  name: nanocached
  region: us-east-1
  version: "1.33"
  tags: {Project: nanocached}
availabilityZones: [us-east-1a, us-east-1b]
vpc:
  nat:
    gateway: Disable          # nodes sit in public subnets; nothing needs a NAT
managedNodeGroups:
  - name: ng-1
    instanceType: t3.medium
    desiredCapacity: 2
    minSize: 2
    maxSize: 2
    volumeSize: 20
    privateNetworking: false
    ssh: {allow: false}
    tags: {Project: nanocached}
eksctl create cluster -f eks-cluster.yaml     # about 15 minutes
kubectl get nodes -L topology.kubernetes.io/zone
# ip-192-168-27-95.ec2.internal    Ready   us-east-1a
# ip-192-168-38-195.ec2.internal   Ready   us-east-1b

Two worker nodes, one per zone, is the minimum for the zone spread below to mean anything. The node group is pinned at two because growth here is in pods: four 1 GiB node pods fit on two t3.mediums with room for discovery and the system pods. Growing past that means either bigger instances or a cluster autoscaler for the node group — a separate decision. The cluster's own security groups already allow all traffic between pods and nothing from outside: the worker nodes have public IPs (for image pulls and the control plane) but no inbound rule reaches them, and neither 8356 nor 9356 answers from the internet.

eksctl installs the metrics-server EKS add-on by default; nothing on this page uses it, but if you add an HPA later it is what the HPA reads — don't apply the upstream components.yaml on top of it, which overwrites the add-on's Service selector and takes the Metrics API down.

The workload

nanocached.yaml — everything in one namespace:

apiVersion: v1
kind: Namespace
metadata:
  name: nanocached
---
apiVersion: v1
kind: Secret
metadata:
  name: nanocached-auth
  namespace: nanocached
type: Opaque
stringData:
  NANOCACHED_AUTH_SECRET: change-me
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: discovery
  namespace: nanocached
spec:
  replicas: 1
  selector:
    matchLabels: {app: nanocached-discovery}
  template:
    metadata:
      labels: {app: nanocached-discovery}
    spec:
      containers:
        - name: discovery
          image: ghcr.io/nanocached/nanocached-discovery@sha256:f3368f0d35fdac03264915bdfe6818b458f14c25068f069aa79bcb2c22805288
          args: ["--host", "0.0.0.0", "--port", "8357", "--replication-factor", "2",
                 "--metrics-port", "9357"]
          envFrom:
            - secretRef: {name: nanocached-auth}
          ports:
            - {containerPort: 8357, name: discovery}
            - {containerPort: 9357, name: ops}
          readinessProbe:
            httpGet: {path: /readyz, port: ops}
            periodSeconds: 5
          livenessProbe:
            httpGet: {path: /healthz, port: ops}
            periodSeconds: 10
          resources:
            requests: {cpu: 100m, memory: 128Mi}
            limits: {memory: 256Mi}
---
apiVersion: v1
kind: Service
metadata:
  name: disc
  namespace: nanocached
spec:
  selector: {app: nanocached-discovery}
  ports:
    - {name: discovery, port: 8357, targetPort: discovery}
    - {name: ops, port: 9357, targetPort: ops}
---
apiVersion: apps/v1
kind: Deployment
metadata:
  name: node
  namespace: nanocached
spec:
  replicas: 2
  selector:
    matchLabels: {app: nanocached-node}
  template:
    metadata:
      labels: {app: nanocached-node}
    spec:
      terminationGracePeriodSeconds: 45
      topologySpreadConstraints:
        - maxSkew: 1
          topologyKey: topology.kubernetes.io/zone
          whenUnsatisfiable: DoNotSchedule
          labelSelector:
            matchLabels: {app: nanocached-node}
      containers:
        - name: node
          image: ghcr.io/nanocached/nanocached-node@sha256:7ca62b477976620b9a81d8d903683187fa901a4ef09d74f23ab8371b0b147623
          args: ["--host", "0.0.0.0", "--port", "8356", "--discovery", "disc:8357",
                 "--drain-timeout", "20", "--metrics-port", "9356", "--max-memory", "536870912"]
          envFrom:
            - secretRef: {name: nanocached-auth}
          ports:
            - {containerPort: 8356, name: data}
            - {containerPort: 9356, name: ops}
          readinessProbe:
            httpGet: {path: /readyz, port: ops}
            periodSeconds: 5
            failureThreshold: 36        # 3 minutes of not-ready before it counts
          livenessProbe:
            httpGet: {path: /healthz, port: ops}
            periodSeconds: 10
          resources:
            requests: {cpu: 500m, memory: 1Gi}
            limits: {memory: 1Gi}

The choices that matter, beyond those the ECS guide already explains (/healthz for discovery's liveness and /readyz for the node's readiness, a memory budget that leaves room under the limit, images pinned by digest):

kubectl apply -f nanocached.yaml
kubectl -n nanocached rollout status deploy/discovery
kubectl -n nanocached rollout status deploy/node
kubectl -n nanocached get pods -o wide
# discovery-…   1/1  Running   ip-192-168-38-195.ec2.internal
# node-…-9ddbd  1/1  Running   ip-192-168-27-95.ec2.internal
# node-…-dgwcp  1/1  Running   ip-192-168-38-195.ec2.internal

Both rollouts were complete 31 s after the apply; the two node pods landed one per zone (the spread constraint) and each logged joined the cluster via discovery at disc:8357 on its first attempt.

Discovery's roster metric, from inside the cluster:

kubectl -n nanocached run curl --rm -it --image=curlimages/curl --restart=Never -- \
  -s http://disc:9357/metrics | grep ^nanocached_discovery_members
# nanocached_discovery_members 2

Verifying

From the application's point of view the cluster is the Service name: disc (or disc.nanocached.svc from another namespace), port 8357, and the secret — the SDK fetches the roster from discovery and talks to the node pods directly (see the SDKs). This cluster was exercised that way from pods in the same namespace with all six SDKs at 0.4.1: set / get / delete succeed through the Service, a wrong secret is rejected, and 20,000 keys written before the next step were read back complete after it. Logs are the ordinary kubectl logs; discovery's narrates every membership change and the nodes' every handoff:

kubectl -n nanocached logs deploy/discovery
kubectl -n nanocached logs -l app=nanocached-node --prefix | grep -E 'joined|migration|decommission'

Growing the cluster

One more node is one more replica:

kubectl -n nanocached scale deploy/node --replicas=3
kubectl -n nanocached rollout status deploy/node

Observed: the third pod registered with discovery one second after it was created and was promoted 15 s later, once each of the two members had handed it its share of the ring — migration completed … sent 13391 keys in both members' logs, join promoted … members now 3 in discovery's — and all 20,000 keys read back. The rollout reported complete 19 s after the scale. Going the other way is --replicas=2: Kubernetes sends SIGTERM, the node runs its planned leave inside --drain-timeout, and terminationGracePeriodSeconds gives it the room (graceful scale-in). Capture a leaver's log with kubectl logs -f before scaling down — a deleted pod's log goes with it.

Note what the drain needs: discovery. Deleting the whole namespace (as the teardown does) stops discovery and the nodes together, and every node then logs decommission: fetching the roster failed (Connection refused); leaving without a handoff — correct, since there is nobody left to hand off to, but not what a scale-down looks like. Under a scale-down discovery stays up, the leaver hands its entries to the survivors and discovery logs node left the cluster.

Teardown

kubectl delete namespace nanocached
eksctl delete cluster --name nanocached --region us-east-1 --wait     # about 12 minutes

eksctl removes its two CloudFormation stacks and with them the VPC, security groups, IAM roles and the OIDC provider. The caller needs ec2:DeleteRoute for that to finish: without it the cluster stack ends DELETE_FAILED on the public subnets' route and the VPC has to be unpicked by hand (this guide's first run). Confirm with aws eks list-clusters (empty), aws cloudformation list-stacks --stack-status-filter CREATE_COMPLETE DELETE_FAILED (no eksctl-nanocached-*), aws ec2 describe-vpcs --filters Name=isDefault,Values=false (empty) and aws iam list-roles --query 'Roles[?starts_with(RoleName,`eksctl-nanocached`)]' (empty). If the create ever fails part-way, the cluster stack is left with termination protection on — aws cloudformation update-termination-protection --no-enable-termination-protection --stack-name eksctl-nanocached-cluster before deleting it.