Cluster on EKS (multi-AZ)
The same cluster as the ECS
guide — one discovery and a node Deployment across two zones —
expressed in Kubernetes objects, on an EKS cluster that
eksctl builds from one config file (VPC, subnets in two
zones, security groups, IAM roles, control plane, a managed node
group). Discovery's stable address is a ClusterIP Service;
the node Deployment grows and shrinks by its replica count, and the
drain on the way down is the SIGTERM Kubernetes already sends. This
page builds the cluster and grows it by one pod; a
HorizontalPodAutoscaler on top of the same Deployment is
how you would tie the count to a metric, and is not covered here.
Every command was run in this order against a real account and torn
down with the commands at the end.
The cluster
eks-cluster.yaml:
apiVersion: eksctl.io/v1alpha5
kind: ClusterConfig
metadata:
name: nanocached
region: us-east-1
version: "1.33"
tags: {Project: nanocached}
availabilityZones: [us-east-1a, us-east-1b]
vpc:
nat:
gateway: Disable # nodes sit in public subnets; nothing needs a NAT
managedNodeGroups:
- name: ng-1
instanceType: t3.medium
desiredCapacity: 2
minSize: 2
maxSize: 2
volumeSize: 20
privateNetworking: false
ssh: {allow: false}
tags: {Project: nanocached}
eksctl create cluster -f eks-cluster.yaml # about 15 minutes
kubectl get nodes -L topology.kubernetes.io/zone
# ip-192-168-27-95.ec2.internal Ready us-east-1a
# ip-192-168-38-195.ec2.internal Ready us-east-1b
Two worker nodes, one per zone, is the minimum for the zone spread
below to mean anything. The node group is pinned at two because
growth here is in pods: four 1 GiB node pods fit on two
t3.mediums with room for discovery and the system pods.
Growing past that means either bigger instances or a cluster
autoscaler for the node group — a separate decision. The
cluster's own security groups already allow all traffic between pods
and nothing from outside: the worker nodes have public IPs (for image
pulls and the control plane) but no inbound rule reaches them, and
neither 8356 nor 9356 answers from the internet.
eksctl installs the metrics-server EKS
add-on by default; nothing on this page uses it, but if you add an HPA
later it is what the HPA reads — don't apply the upstream
components.yaml on top of it, which overwrites the
add-on's Service selector and takes the Metrics API down.
The workload
nanocached.yaml — everything in one namespace:
apiVersion: v1
kind: Namespace
metadata:
name: nanocached
---
apiVersion: v1
kind: Secret
metadata:
name: nanocached-auth
namespace: nanocached
type: Opaque
stringData:
NANOCACHED_AUTH_SECRET: change-me
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: discovery
namespace: nanocached
spec:
replicas: 1
selector:
matchLabels: {app: nanocached-discovery}
template:
metadata:
labels: {app: nanocached-discovery}
spec:
containers:
- name: discovery
image: ghcr.io/nanocached/nanocached-discovery@sha256:f3368f0d35fdac03264915bdfe6818b458f14c25068f069aa79bcb2c22805288
args: ["--host", "0.0.0.0", "--port", "8357", "--replication-factor", "2",
"--metrics-port", "9357"]
envFrom:
- secretRef: {name: nanocached-auth}
ports:
- {containerPort: 8357, name: discovery}
- {containerPort: 9357, name: ops}
readinessProbe:
httpGet: {path: /readyz, port: ops}
periodSeconds: 5
livenessProbe:
httpGet: {path: /healthz, port: ops}
periodSeconds: 10
resources:
requests: {cpu: 100m, memory: 128Mi}
limits: {memory: 256Mi}
---
apiVersion: v1
kind: Service
metadata:
name: disc
namespace: nanocached
spec:
selector: {app: nanocached-discovery}
ports:
- {name: discovery, port: 8357, targetPort: discovery}
- {name: ops, port: 9357, targetPort: ops}
---
apiVersion: apps/v1
kind: Deployment
metadata:
name: node
namespace: nanocached
spec:
replicas: 2
selector:
matchLabels: {app: nanocached-node}
template:
metadata:
labels: {app: nanocached-node}
spec:
terminationGracePeriodSeconds: 45
topologySpreadConstraints:
- maxSkew: 1
topologyKey: topology.kubernetes.io/zone
whenUnsatisfiable: DoNotSchedule
labelSelector:
matchLabels: {app: nanocached-node}
containers:
- name: node
image: ghcr.io/nanocached/nanocached-node@sha256:7ca62b477976620b9a81d8d903683187fa901a4ef09d74f23ab8371b0b147623
args: ["--host", "0.0.0.0", "--port", "8356", "--discovery", "disc:8357",
"--drain-timeout", "20", "--metrics-port", "9356", "--max-memory", "536870912"]
envFrom:
- secretRef: {name: nanocached-auth}
ports:
- {containerPort: 8356, name: data}
- {containerPort: 9356, name: ops}
readinessProbe:
httpGet: {path: /readyz, port: ops}
periodSeconds: 5
failureThreshold: 36 # 3 minutes of not-ready before it counts
livenessProbe:
httpGet: {path: /healthz, port: ops}
periodSeconds: 10
resources:
requests: {cpu: 500m, memory: 1Gi}
limits: {memory: 1Gi}
The choices that matter, beyond those the ECS guide already
explains (/healthz for discovery's liveness and
/readyz for the node's readiness, a memory budget that
leaves room under the limit, images pinned by digest):
disc:8357is the Service name, resolvable the moment it is created. Node pods that start before discovery is ready see Connection refused (the Service has no endpoints until discovery's readiness probe passes), retry every 5 s and join as soon as it does — no ordering step is needed, unlike Cloud Map on ECS.topologySpreadConstraintson zone withDoNotSchedule: with R=2, both copies of a key in one zone is the failure mode this prevents. Two, three and four pods were placed 1/1, 2/1 and 2/2 across the zones in verification.terminationGracePeriodSeconds: 45against--drain-timeout 20, as on ECS. NopreStophook: the SIGTERM is the drain.- Readiness
failureThreshold: 36(three minutes) plays the role ECS'sstartPerioddoes — a joiner waiting its turn in a staged join is not-ready but not broken. Liveness stays at the default threshold; a process that stops answering/healthzshould be restarted. resources— the 1 GiB memory limit is--max-memory(512 MiB) plus the process's own headroom, and the CPU request is what the scheduler packs by (and what a later HPA would compute its percentage of).- The Secret is
envFrom, so the auth secret becomesNANOCACHED_AUTH_SECRETwithout appearing in the pod spec. Rotating it iskubectl applyof the new value and arollout restartof both Deployments.
kubectl apply -f nanocached.yaml
kubectl -n nanocached rollout status deploy/discovery
kubectl -n nanocached rollout status deploy/node
kubectl -n nanocached get pods -o wide
# discovery-… 1/1 Running ip-192-168-38-195.ec2.internal
# node-…-9ddbd 1/1 Running ip-192-168-27-95.ec2.internal
# node-…-dgwcp 1/1 Running ip-192-168-38-195.ec2.internal
Both rollouts were complete 31 s after the apply; the two
node pods landed one per zone (the spread constraint) and each logged
joined the cluster via discovery at disc:8357 on its
first attempt.
Discovery's roster metric, from inside the cluster:
kubectl -n nanocached run curl --rm -it --image=curlimages/curl --restart=Never -- \
-s http://disc:9357/metrics | grep ^nanocached_discovery_members
# nanocached_discovery_members 2
Verifying
From the application's point of view the cluster is the Service
name: disc (or disc.nanocached.svc from
another namespace), port 8357, and the secret — the SDK fetches the
roster from discovery and talks to the node pods directly (see the SDKs). This cluster was exercised that way
from pods in the same namespace with all six SDKs at 0.4.1: set / get /
delete succeed through the Service, a wrong secret is rejected, and
20,000 keys written before the next step were read back complete
after it. Logs are the ordinary kubectl logs;
discovery's narrates every membership change and the nodes' every
handoff:
kubectl -n nanocached logs deploy/discovery
kubectl -n nanocached logs -l app=nanocached-node --prefix | grep -E 'joined|migration|decommission'
Growing the cluster
One more node is one more replica:
kubectl -n nanocached scale deploy/node --replicas=3
kubectl -n nanocached rollout status deploy/node
Observed: the third pod registered with discovery one second after
it was created and was promoted 15 s later, once each of the two
members had handed it its share of the ring — migration
completed … sent 13391 keys in both members' logs, join
promoted … members now 3 in discovery's — and all 20,000 keys
read back. The rollout reported complete 19 s after the scale.
Going the other way is --replicas=2: Kubernetes sends
SIGTERM, the node runs its planned leave inside
--drain-timeout, and terminationGracePeriodSeconds
gives it the room (graceful
scale-in). Capture a leaver's log with kubectl logs -f
before scaling down — a deleted pod's log goes with it.
Note what the drain needs: discovery. Deleting the whole namespace
(as the teardown does) stops discovery and the nodes together, and
every node then logs decommission: fetching the roster failed
(Connection refused); leaving without a handoff — correct, since
there is nobody left to hand off to, but not what a scale-down looks
like. Under a scale-down discovery stays up, the leaver hands its
entries to the survivors and discovery logs node left the
cluster.
Teardown
kubectl delete namespace nanocached
eksctl delete cluster --name nanocached --region us-east-1 --wait # about 12 minutes
eksctl removes its two CloudFormation stacks and with
them the VPC, security groups, IAM roles and the OIDC provider. The
caller needs ec2:DeleteRoute for that to finish: without
it the cluster stack ends DELETE_FAILED on the public
subnets' route and the VPC has to be unpicked by hand (this guide's
first run).
Confirm with aws eks list-clusters (empty), aws
cloudformation list-stacks --stack-status-filter CREATE_COMPLETE
DELETE_FAILED (no eksctl-nanocached-*),
aws ec2 describe-vpcs --filters Name=isDefault,Values=false
(empty) and aws iam list-roles --query
'Roles[?starts_with(RoleName,`eksctl-nanocached`)]' (empty). If
the create ever fails part-way, the cluster stack is left with
termination protection on — aws cloudformation
update-termination-protection --no-enable-termination-protection
--stack-name eksctl-nanocached-cluster before deleting it.