Cluster on EC2

A replicated cache on plain EC2 instances, without an orchestrator: one instance runs the discovery server and a node at a fixed private address, every other instance runs an identical node, and the cluster grows by launching more instances from an image of one of those nodes. Built from an empty account with the AWS CLI, verified two nodes up, a third added from the image, one node removed, and torn down again. Read what every deployment needs first for the operations endpoint, the metrics and the SIGTERM drain every shape relies on; this page builds on the single node on EC2.

The shape

Two kinds of instance, one systemd unit for the node on both:

Adding a node is run-instances from that image with no user-data; the node joins with a staged data handoff, so the keys it takes ownership of move to it before it serves. Removing one is systemctl stop, which the node turns into a planned leave (graceful scale-in), and then terminate-instances. With discovery's default replication factor of 2, every key lives on two nodes, so a single instance loss — planned or not — loses nothing.

What this shape does not do: it has one discovery server. Nodes and clients that are already connected keep working if it is down (only membership updates stop), but a new client cannot bootstrap and a restarted node cannot re-register until it is back; the address is fixed precisely so the seed can be rebuilt in place. Nor is there an Auto Scaling group — growth is a command you run. Both are deliberate: this is the shape for a handful of instances that people operate; the ECS and EKS shapes are the ones that scale on their own.

export AWS_REGION=us-east-1
NAME=nanocached
TAG="{Key=Project,Value=$NAME}"     # every resource carries it — see teardown

Network

The same VPC as the single-node shape: two public subnets, one for the cache tier and one for the application tier, on one route table through an internet gateway. No NAT gateway — outbound (image pulls, package installs, Systems Manager) goes straight out, and inbound is whatever the security groups admit. The subnet CIDR matters more here than it did with one node: 10.0.0.11 below has to be a free address inside the node subnet.

VPC_ID=$(aws ec2 create-vpc --cidr-block 10.0.0.0/16 \
  --tag-specifications "ResourceType=vpc,Tags=[$TAG]" --query Vpc.VpcId --output text)
aws ec2 modify-vpc-attribute --vpc-id $VPC_ID --enable-dns-support
aws ec2 modify-vpc-attribute --vpc-id $VPC_ID --enable-dns-hostnames

NODE_SUBNET_ID=$(aws ec2 create-subnet --vpc-id $VPC_ID --cidr-block 10.0.0.0/24 \
  --availability-zone ${AWS_REGION}a \
  --tag-specifications "ResourceType=subnet,Tags=[$TAG,{Key=Name,Value=$NAME-node}]" \
  --query Subnet.SubnetId --output text)
APP_SUBNET_ID=$(aws ec2 create-subnet --vpc-id $VPC_ID --cidr-block 10.0.1.0/24 \
  --availability-zone ${AWS_REGION}a \
  --tag-specifications "ResourceType=subnet,Tags=[$TAG,{Key=Name,Value=$NAME-app}]" \
  --query Subnet.SubnetId --output text)
aws ec2 modify-subnet-attribute --subnet-id $NODE_SUBNET_ID --map-public-ip-on-launch
aws ec2 modify-subnet-attribute --subnet-id $APP_SUBNET_ID --map-public-ip-on-launch

IGW_ID=$(aws ec2 create-internet-gateway \
  --tag-specifications "ResourceType=internet-gateway,Tags=[$TAG]" \
  --query InternetGateway.InternetGatewayId --output text)
aws ec2 attach-internet-gateway --internet-gateway-id $IGW_ID --vpc-id $VPC_ID

RTB_ID=$(aws ec2 create-route-table --vpc-id $VPC_ID \
  --tag-specifications "ResourceType=route-table,Tags=[$TAG]" --query RouteTable.RouteTableId --output text)
aws ec2 create-route --route-table-id $RTB_ID --destination-cidr-block 0.0.0.0/0 --gateway-id $IGW_ID
aws ec2 associate-route-table --route-table-id $RTB_ID --subnet-id $NODE_SUBNET_ID
aws ec2 associate-route-table --route-table-id $RTB_ID --subnet-id $APP_SUBNET_ID

Name the availability zone explicitly (a subnet left to AWS can land in a zone without t3). Already have a VPC? Set VPC_ID, NODE_SUBNET_ID and APP_SUBNET_ID to existing public subnets, pick a free address in the node subnet for the seed, and substitute your VPC's CIDR for 10.0.0.0/16 in the operations-port rule.

Security groups

Two groups, as before — nanocached-app is the label the application servers wear, nanocached-node refers to it — with two additions for a cluster: the discovery port (8357) opens to the app group alongside the data port, and the node group refers to itself, because nodes hand keys to each other on 8356 and register with discovery on 8357. The operations range now covers both the node's 9356 and discovery's 9357.

APP_SG=$(aws ec2 create-security-group --group-name $NAME-app \
  --description "nanocached application tier" --vpc-id $VPC_ID \
  --tag-specifications "ResourceType=security-group,Tags=[$TAG]" --query GroupId --output text)
NODE_SG=$(aws ec2 create-security-group --group-name $NAME-node \
  --description "nanocached node" --vpc-id $VPC_ID \
  --tag-specifications "ResourceType=security-group,Tags=[$TAG]" --query GroupId --output text)

aws ec2 authorize-security-group-ingress --group-id $NODE_SG --ip-permissions \
  "IpProtocol=tcp,FromPort=8356,ToPort=8357,UserIdGroupPairs=[{GroupId=$APP_SG,Description=data and discovery ports from app tier}]" \
  "IpProtocol=tcp,FromPort=8356,ToPort=8357,UserIdGroupPairs=[{GroupId=$NODE_SG,Description=node-to-node handoff and discovery registration}]" \
  "IpProtocol=tcp,FromPort=9356,ToPort=9357,IpRanges=[{CidrIp=10.0.0.0/16,Description=operations endpoints (node 9356 and discovery 9357) from inside the VPC}]"

Port 22 stays absent; the shell is Session Manager. The operations endpoints are unauthenticated, hence a CIDR inside the VPC and never 0.0.0.0/0.

IAM

One role with AmazonSSMManagedInstanceCore, shared by every instance:

aws iam create-role --role-name $NAME-ec2 --tags Key=Project,Value=$NAME \
  --assume-role-policy-document '{"Version":"2012-10-17","Statement":[{"Effect":"Allow",
    "Principal":{"Service":"ec2.amazonaws.com"},"Action":"sts:AssumeRole"}]}'
aws iam attach-role-policy --role-name $NAME-ec2 \
  --policy-arn arn:aws:iam::aws:policy/AmazonSSMManagedInstanceCore
aws iam create-instance-profile --instance-profile-name $NAME-ec2
aws iam add-role-to-instance-profile --instance-profile-name $NAME-ec2 --role-name $NAME-ec2
sleep 15   # a brand-new instance profile takes a few seconds to become launchable

The seed

Save this as seed-user-data.sh. It is the single-node user-data with a second unit for discovery and two flags on the node: --discovery naming the seed's own fixed address, and --drain-timeout, the budget a stopping node gets to hand its keys away (20 s here, inside the unit's 30 s docker stop grace). Discovery reads the same secret file; its --replication-factor 2 is the cluster's R.

#!/bin/bash
set -euo pipefail
dnf install -y docker
systemctl enable --now docker

install -d -m 700 /etc/nanocached
cat > /etc/nanocached/env <<'ENV'
NANOCACHED_AUTH_SECRET=change-me
ENV
chmod 600 /etc/nanocached/env

cat > /etc/systemd/system/nanocached-discovery.service <<'UNIT'
[Unit]
Description=nanocached discovery
After=docker.service
Requires=docker.service

[Service]
ExecStartPre=-/usr/bin/docker rm -f nanocached-discovery
ExecStartPre=/usr/bin/docker pull ghcr.io/nanocached/nanocached-discovery@sha256:f3368f0d35fdac03264915bdfe6818b458f14c25068f069aa79bcb2c22805288
ExecStart=/usr/bin/docker run --rm --name nanocached-discovery \
  --network host \
  --env-file /etc/nanocached/env \
  ghcr.io/nanocached/nanocached-discovery@sha256:f3368f0d35fdac03264915bdfe6818b458f14c25068f069aa79bcb2c22805288 \
  --host 0.0.0.0 --port 8357 --metrics-port 9357 \
  --replication-factor 2
ExecStop=/usr/bin/docker stop -t 30 nanocached-discovery
Restart=always
RestartSec=2

[Install]
WantedBy=multi-user.target
UNIT

cat > /etc/systemd/system/nanocached-node.service <<'UNIT'
[Unit]
Description=nanocached node
After=docker.service nanocached-discovery.service
Requires=docker.service

[Service]
ExecStartPre=-/usr/bin/docker rm -f nanocached-node
ExecStartPre=/usr/bin/docker pull ghcr.io/nanocached/nanocached-node@sha256:7ca62b477976620b9a81d8d903683187fa901a4ef09d74f23ab8371b0b147623
ExecStart=/usr/bin/docker run --rm --name nanocached-node \
  --network host \
  --env-file /etc/nanocached/env \
  ghcr.io/nanocached/nanocached-node@sha256:7ca62b477976620b9a81d8d903683187fa901a4ef09d74f23ab8371b0b147623 \
  --host 0.0.0.0 --port 8356 --metrics-port 9356 \
  --discovery 10.0.0.11:8357 --drain-timeout 20 \
  --max-memory 1073741824
ExecStop=/usr/bin/docker stop -t 30 nanocached-node
Restart=always
RestartSec=2

[Install]
WantedBy=multi-user.target
UNIT

systemctl daemon-reload
systemctl enable --now nanocached-discovery nanocached-node

The node unit names nanocached-discovery.service in its After= so that on the seed the registry is up before the node registers; on a plain node, where that unit does not exist, systemd ignores the reference — which is what lets the unit be byte-for-byte the same on both. Everything the single-node guide says about NANOCACHED_AUTH_SECRET, --max-memory and the pinned image digests applies unchanged.

A node

Save this as node-user-data.sh: the same preamble and the same node unit, nothing else.

#!/bin/bash
set -euo pipefail
dnf install -y docker
systemctl enable --now docker

install -d -m 700 /etc/nanocached
cat > /etc/nanocached/env <<'ENV'
NANOCACHED_AUTH_SECRET=change-me
ENV
chmod 600 /etc/nanocached/env

cat > /etc/systemd/system/nanocached-node.service <<'UNIT'
[Unit]
Description=nanocached node
After=docker.service nanocached-discovery.service
Requires=docker.service

[Service]
ExecStartPre=-/usr/bin/docker rm -f nanocached-node
ExecStartPre=/usr/bin/docker pull ghcr.io/nanocached/nanocached-node@sha256:7ca62b477976620b9a81d8d903683187fa901a4ef09d74f23ab8371b0b147623
ExecStart=/usr/bin/docker run --rm --name nanocached-node \
  --network host \
  --env-file /etc/nanocached/env \
  ghcr.io/nanocached/nanocached-node@sha256:7ca62b477976620b9a81d8d903683187fa901a4ef09d74f23ab8371b0b147623 \
  --host 0.0.0.0 --port 8356 --metrics-port 9356 \
  --discovery 10.0.0.11:8357 --drain-timeout 20 \
  --max-memory 1073741824
ExecStop=/usr/bin/docker stop -t 30 nanocached-node
Restart=always
RestartSec=2

[Install]
WantedBy=multi-user.target
UNIT

systemctl daemon-reload
systemctl enable --now nanocached-node

Launch

The seed goes first, at its fixed address; then a node. Both from the current Amazon Linux 2023 AMI:

AMI_ID=$(aws ssm get-parameter \
  --name /aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-x86_64 \
  --query Parameter.Value --output text)

# A: the seed — discovery plus a node, at the fixed address every other
# node and every client will name.
SEED_ID=$(aws ec2 run-instances --image-id $AMI_ID --instance-type t3.small \
  --subnet-id $NODE_SUBNET_ID --security-group-ids $NODE_SG --private-ip-address 10.0.0.11 \
  --iam-instance-profile Name=$NAME-ec2 --user-data file://seed-user-data.sh \
  --tag-specifications "ResourceType=instance,Tags=[$TAG,{Key=Name,Value=$NAME-seed}]" \
  --query 'Instances[0].InstanceId' --output text)

# B: a plain node.
NODE_ID=$(aws ec2 run-instances --image-id $AMI_ID --instance-type t3.small \
  --subnet-id $NODE_SUBNET_ID --security-group-ids $NODE_SG \
  --iam-instance-profile Name=$NAME-ec2 --user-data file://node-user-data.sh \
  --tag-specifications "ResourceType=instance,Tags=[$TAG,{Key=Name,Value=$NAME-node}]" \
  --query 'Instances[0].InstanceId' --output text)

aws ec2 wait instance-running --instance-ids $SEED_ID $NODE_ID
NODE_IP=$(aws ec2 describe-instances --instance-ids $NODE_ID \
  --query 'Reservations[0].Instances[0].PrivateIpAddress' --output text)

Timing, from launch: the seed's discovery and node both answered /readyz 200 and discovery reported two members at 86 s after boot; the plain node registered at about 80 s and was promoted to a member twelve seconds later, once the seed's own node had joined ahead of it. The order in the discovery log is the staged join at work — a node waits until the members before it have finished handing over.

Verifying

On the seed, both services and the membership count. Discovery serves /healthz, /readyz and /metrics on its own operations port, 9357:

aws ssm send-command --instance-ids $SEED_ID --document-name AWS-RunShellScript \
  --parameters 'commands=["systemctl is-active nanocached-discovery nanocached-node",
    "curl -s -o /dev/null -w discovery-readyz=%{http_code}\\\\n http://127.0.0.1:9357/readyz",
    "curl -s -o /dev/null -w node-readyz=%{http_code}\\\\n http://127.0.0.1:9356/readyz",
    "curl -s http://127.0.0.1:9357/metrics | grep ^nanocached_discovery_members"]'
# active / active / discovery-readyz=200 / node-readyz=200 / nanocached_discovery_members 2

journalctl -u nanocached-discovery narrates every membership change (node registered, join started, join promoted, node left the cluster) and journalctl -u nanocached-node the handoffs (migration started/completed, re-replication after membership change) — the two logs quoted in the rest of this page.

From the application tier, the SDK's configuration is discovery's address — 10.0.0.11, port 8357 — and the secret; the SDK fetches the node list and routes each key to its owners itself (see the SDKs). This cluster was exercised that way from an instance in the app subnet with all six SDKs at 0.4.1: set / get / delete succeed through discovery, a wrong secret is rejected, and 2,000 keys written before the steps below were read back complete after each of them.

Growing the cluster

Take an image of a node. --no-reboot keeps it running: a node keeps nothing on disk, so there is nothing to quiesce, and a reboot would only cost the cluster a leave-and-rejoin. The image carries the unit, the secret file and the already-pulled container image:

# Bake the plain node into an AMI without stopping it: the node keeps
# nothing on disk, so there is nothing to quiesce.
IMAGE_ID=$(aws ec2 create-image --instance-id $NODE_ID --name $NAME-node-$(date +%Y%m%d-%H%M) \
  --no-reboot --tag-specifications "ResourceType=image,Tags=[$TAG]" "ResourceType=snapshot,Tags=[$TAG]" \
  --query ImageId --output text)
aws ec2 wait image-available --image-ids $IMAGE_ID

Then add a node from it — no user-data, no configuration:

# Add a node from the image: no user-data — the unit, the secret file and
# the pulled image are already on the disk, and the node picks a fresh
# identity on start.
NODE2_ID=$(aws ec2 run-instances --image-id $IMAGE_ID --instance-type t3.small \
  --subnet-id $NODE_SUBNET_ID --security-group-ids $NODE_SG \
  --iam-instance-profile Name=$NAME-ec2 \
  --tag-specifications "ResourceType=instance,Tags=[$TAG,{Key=Name,Value=$NAME-node}]" \
  --query 'Instances[0].InstanceId' --output text)
aws ec2 wait instance-running --instance-ids $NODE2_ID

Observed: the image was available three minutes after create-image; the new instance registered with discovery 37 s after launch and was promoted the same second, the seed's node handing it 1,328 of the 2,000 test keys (the share the ring now assigns to it) before the join completed; every key still read back from the application tier. Repeat the block for each node you want — the image is good until you change the unit, the secret or the pinned digest, at which point you bake a new one from a node running the new configuration.

Removing a node

Stop the service first, then terminate the instance. systemctl stop sends SIGTERM through docker stop; the node deregisters and hands its keys to their next owners within --drain-timeout:

aws ssm send-command --instance-ids $NODE_ID --document-name AWS-RunShellScript \
  --parameters 'commands=["systemctl stop nanocached-node"]'
# then, once the command reports Success:
aws ec2 terminate-instances --instance-ids $NODE_ID

Observed with three members: discovery logged node left the cluster in the same second the stop was issued, the two remaining nodes re-replicated 672 entries between them so every key was again on two nodes, and systemctl stop returned after 20 s — the drain budget. The 2,000 keys read back complete throughout. Terminating an instance without the stop is what an unplanned loss looks like: discovery drops the node after its liveness timeout instead of immediately, and the keys are still on their other owner.

Teardown

Instances first, then the image and its snapshot (deregistering an AMI does not delete the snapshot behind it), then the rest in reverse order. Anything you launched in the app group goes first too.

aws ec2 terminate-instances --instance-ids $SEED_ID $NODE_ID $NODE2_ID
aws ec2 wait instance-terminated --instance-ids $SEED_ID $NODE_ID $NODE2_ID

SNAPSHOT_ID=$(aws ec2 describe-images --image-ids $IMAGE_ID \
  --query 'Images[0].BlockDeviceMappings[0].Ebs.SnapshotId' --output text)
aws ec2 deregister-image --image-id $IMAGE_ID
aws ec2 delete-snapshot --snapshot-id $SNAPSHOT_ID

aws iam remove-role-from-instance-profile --instance-profile-name $NAME-ec2 --role-name $NAME-ec2
aws iam delete-instance-profile --instance-profile-name $NAME-ec2
aws iam detach-role-policy --role-name $NAME-ec2 \
  --policy-arn arn:aws:iam::aws:policy/AmazonSSMManagedInstanceCore
aws iam delete-role --role-name $NAME-ec2

aws ec2 delete-security-group --group-id $NODE_SG   # first: it references $APP_SG
aws ec2 delete-security-group --group-id $APP_SG

for ASSOC in $(aws ec2 describe-route-tables --route-table-ids $RTB_ID \
  --query 'RouteTables[0].Associations[?!Main].RouteTableAssociationId' --output text); do
  aws ec2 disassociate-route-table --association-id $ASSOC
done
aws ec2 delete-route-table --route-table-id $RTB_ID
aws ec2 detach-internet-gateway --internet-gateway-id $IGW_ID --vpc-id $VPC_ID
aws ec2 delete-internet-gateway --internet-gateway-id $IGW_ID
aws ec2 delete-subnet --subnet-id $NODE_SUBNET_ID
aws ec2 delete-subnet --subnet-id $APP_SUBNET_ID
aws ec2 delete-vpc --vpc-id $VPC_ID

Confirm nothing is left: aws ec2 describe-vpcs --filters Name=isDefault,Values=false, aws ec2 describe-instances --filters Name=instance-state-name,Values=running, aws ec2 describe-images --owners self and aws ec2 describe-snapshots --owner-ids self should all come back empty, and aws iam get-role --role-name nanocached-ec2 should say NoSuchEntity.