Cluster on EC2
A replicated cache on plain EC2 instances, without an orchestrator: one instance runs the discovery server and a node at a fixed private address, every other instance runs an identical node, and the cluster grows by launching more instances from an image of one of those nodes. Built from an empty account with the AWS CLI, verified two nodes up, a third added from the image, one node removed, and torn down again. Read what every deployment needs first for the operations endpoint, the metrics and the SIGTERM drain every shape relies on; this page builds on the single node on EC2.
The shape
Two kinds of instance, one systemd unit for the node on both:
- The seed —
nanocached-discoveryand ananocached-nodeon one instance, launched with a fixed private IP (10.0.0.11below). That address is the one thing every other node and every client is configured with: nodes register and heartbeat there (--discovery 10.0.0.11:8357), clients fetch the node list from it and then talk to nodes directly. Discovery is never on the data path. - Nodes — every other instance. Same unit file as the seed's node, same command line; a node has no identity of its own on disk (it picks a random name each start, and its address is learned by discovery from the heartbeat connection), so nodes are interchangeable and an image of one is an image of all of them.
Adding a node is run-instances from that image with no
user-data; the node joins with a staged data handoff, so the keys it
takes ownership of move to it before it serves. Removing one is
systemctl stop, which the node turns into a planned leave
(graceful scale-in),
and then terminate-instances. With discovery's default
replication factor of 2, every key lives on two nodes, so a single
instance loss — planned or not — loses nothing.
What this shape does not do: it has one discovery server. Nodes and clients that are already connected keep working if it is down (only membership updates stop), but a new client cannot bootstrap and a restarted node cannot re-register until it is back; the address is fixed precisely so the seed can be rebuilt in place. Nor is there an Auto Scaling group — growth is a command you run. Both are deliberate: this is the shape for a handful of instances that people operate; the ECS and EKS shapes are the ones that scale on their own.
export AWS_REGION=us-east-1
NAME=nanocached
TAG="{Key=Project,Value=$NAME}" # every resource carries it — see teardown
Network
The same VPC as the single-node shape: two public subnets, one for
the cache tier and one for the application tier, on one route table
through an internet gateway. No NAT gateway — outbound (image pulls,
package installs, Systems Manager) goes straight out, and inbound is
whatever the security groups admit. The subnet CIDR matters more here
than it did with one node: 10.0.0.11 below has to be a
free address inside the node subnet.
VPC_ID=$(aws ec2 create-vpc --cidr-block 10.0.0.0/16 \
--tag-specifications "ResourceType=vpc,Tags=[$TAG]" --query Vpc.VpcId --output text)
aws ec2 modify-vpc-attribute --vpc-id $VPC_ID --enable-dns-support
aws ec2 modify-vpc-attribute --vpc-id $VPC_ID --enable-dns-hostnames
NODE_SUBNET_ID=$(aws ec2 create-subnet --vpc-id $VPC_ID --cidr-block 10.0.0.0/24 \
--availability-zone ${AWS_REGION}a \
--tag-specifications "ResourceType=subnet,Tags=[$TAG,{Key=Name,Value=$NAME-node}]" \
--query Subnet.SubnetId --output text)
APP_SUBNET_ID=$(aws ec2 create-subnet --vpc-id $VPC_ID --cidr-block 10.0.1.0/24 \
--availability-zone ${AWS_REGION}a \
--tag-specifications "ResourceType=subnet,Tags=[$TAG,{Key=Name,Value=$NAME-app}]" \
--query Subnet.SubnetId --output text)
aws ec2 modify-subnet-attribute --subnet-id $NODE_SUBNET_ID --map-public-ip-on-launch
aws ec2 modify-subnet-attribute --subnet-id $APP_SUBNET_ID --map-public-ip-on-launch
IGW_ID=$(aws ec2 create-internet-gateway \
--tag-specifications "ResourceType=internet-gateway,Tags=[$TAG]" \
--query InternetGateway.InternetGatewayId --output text)
aws ec2 attach-internet-gateway --internet-gateway-id $IGW_ID --vpc-id $VPC_ID
RTB_ID=$(aws ec2 create-route-table --vpc-id $VPC_ID \
--tag-specifications "ResourceType=route-table,Tags=[$TAG]" --query RouteTable.RouteTableId --output text)
aws ec2 create-route --route-table-id $RTB_ID --destination-cidr-block 0.0.0.0/0 --gateway-id $IGW_ID
aws ec2 associate-route-table --route-table-id $RTB_ID --subnet-id $NODE_SUBNET_ID
aws ec2 associate-route-table --route-table-id $RTB_ID --subnet-id $APP_SUBNET_ID
Name the availability zone explicitly (a subnet left to AWS can land
in a zone without t3). Already have a VPC? Set
VPC_ID, NODE_SUBNET_ID and
APP_SUBNET_ID to existing public subnets, pick a free
address in the node subnet for the seed, and substitute your VPC's
CIDR for 10.0.0.0/16 in the operations-port rule.
Security groups
Two groups, as before — nanocached-app is the label the
application servers wear, nanocached-node refers to it —
with two additions for a cluster: the discovery port (8357) opens to
the app group alongside the data port, and the node group refers to
itself, because nodes hand keys to each other on 8356 and
register with discovery on 8357. The operations range now covers both
the node's 9356 and discovery's 9357.
APP_SG=$(aws ec2 create-security-group --group-name $NAME-app \
--description "nanocached application tier" --vpc-id $VPC_ID \
--tag-specifications "ResourceType=security-group,Tags=[$TAG]" --query GroupId --output text)
NODE_SG=$(aws ec2 create-security-group --group-name $NAME-node \
--description "nanocached node" --vpc-id $VPC_ID \
--tag-specifications "ResourceType=security-group,Tags=[$TAG]" --query GroupId --output text)
aws ec2 authorize-security-group-ingress --group-id $NODE_SG --ip-permissions \
"IpProtocol=tcp,FromPort=8356,ToPort=8357,UserIdGroupPairs=[{GroupId=$APP_SG,Description=data and discovery ports from app tier}]" \
"IpProtocol=tcp,FromPort=8356,ToPort=8357,UserIdGroupPairs=[{GroupId=$NODE_SG,Description=node-to-node handoff and discovery registration}]" \
"IpProtocol=tcp,FromPort=9356,ToPort=9357,IpRanges=[{CidrIp=10.0.0.0/16,Description=operations endpoints (node 9356 and discovery 9357) from inside the VPC}]"
Port 22 stays absent; the shell is Session Manager. The operations
endpoints are unauthenticated, hence a CIDR inside the VPC and never
0.0.0.0/0.
IAM
One role with AmazonSSMManagedInstanceCore, shared by
every instance:
aws iam create-role --role-name $NAME-ec2 --tags Key=Project,Value=$NAME \
--assume-role-policy-document '{"Version":"2012-10-17","Statement":[{"Effect":"Allow",
"Principal":{"Service":"ec2.amazonaws.com"},"Action":"sts:AssumeRole"}]}'
aws iam attach-role-policy --role-name $NAME-ec2 \
--policy-arn arn:aws:iam::aws:policy/AmazonSSMManagedInstanceCore
aws iam create-instance-profile --instance-profile-name $NAME-ec2
aws iam add-role-to-instance-profile --instance-profile-name $NAME-ec2 --role-name $NAME-ec2
sleep 15 # a brand-new instance profile takes a few seconds to become launchable
The seed
Save this as seed-user-data.sh. It is the single-node
user-data with a second unit for discovery and two flags on the node:
--discovery naming the seed's own fixed address, and
--drain-timeout, the budget a stopping node gets to hand
its keys away (20 s here, inside the unit's 30 s
docker stop grace). Discovery reads the same secret file;
its --replication-factor 2 is the cluster's R.
#!/bin/bash
set -euo pipefail
dnf install -y docker
systemctl enable --now docker
install -d -m 700 /etc/nanocached
cat > /etc/nanocached/env <<'ENV'
NANOCACHED_AUTH_SECRET=change-me
ENV
chmod 600 /etc/nanocached/env
cat > /etc/systemd/system/nanocached-discovery.service <<'UNIT'
[Unit]
Description=nanocached discovery
After=docker.service
Requires=docker.service
[Service]
ExecStartPre=-/usr/bin/docker rm -f nanocached-discovery
ExecStartPre=/usr/bin/docker pull ghcr.io/nanocached/nanocached-discovery@sha256:f3368f0d35fdac03264915bdfe6818b458f14c25068f069aa79bcb2c22805288
ExecStart=/usr/bin/docker run --rm --name nanocached-discovery \
--network host \
--env-file /etc/nanocached/env \
ghcr.io/nanocached/nanocached-discovery@sha256:f3368f0d35fdac03264915bdfe6818b458f14c25068f069aa79bcb2c22805288 \
--host 0.0.0.0 --port 8357 --metrics-port 9357 \
--replication-factor 2
ExecStop=/usr/bin/docker stop -t 30 nanocached-discovery
Restart=always
RestartSec=2
[Install]
WantedBy=multi-user.target
UNIT
cat > /etc/systemd/system/nanocached-node.service <<'UNIT'
[Unit]
Description=nanocached node
After=docker.service nanocached-discovery.service
Requires=docker.service
[Service]
ExecStartPre=-/usr/bin/docker rm -f nanocached-node
ExecStartPre=/usr/bin/docker pull ghcr.io/nanocached/nanocached-node@sha256:7ca62b477976620b9a81d8d903683187fa901a4ef09d74f23ab8371b0b147623
ExecStart=/usr/bin/docker run --rm --name nanocached-node \
--network host \
--env-file /etc/nanocached/env \
ghcr.io/nanocached/nanocached-node@sha256:7ca62b477976620b9a81d8d903683187fa901a4ef09d74f23ab8371b0b147623 \
--host 0.0.0.0 --port 8356 --metrics-port 9356 \
--discovery 10.0.0.11:8357 --drain-timeout 20 \
--max-memory 1073741824
ExecStop=/usr/bin/docker stop -t 30 nanocached-node
Restart=always
RestartSec=2
[Install]
WantedBy=multi-user.target
UNIT
systemctl daemon-reload
systemctl enable --now nanocached-discovery nanocached-node
The node unit names nanocached-discovery.service in its
After= so that on the seed the registry is up before the
node registers; on a plain node, where that unit does not exist,
systemd ignores the reference — which is what lets the unit be
byte-for-byte the same on both. Everything the single-node guide says
about NANOCACHED_AUTH_SECRET, --max-memory
and the pinned image digests applies unchanged.
A node
Save this as node-user-data.sh: the same preamble and
the same node unit, nothing else.
#!/bin/bash
set -euo pipefail
dnf install -y docker
systemctl enable --now docker
install -d -m 700 /etc/nanocached
cat > /etc/nanocached/env <<'ENV'
NANOCACHED_AUTH_SECRET=change-me
ENV
chmod 600 /etc/nanocached/env
cat > /etc/systemd/system/nanocached-node.service <<'UNIT'
[Unit]
Description=nanocached node
After=docker.service nanocached-discovery.service
Requires=docker.service
[Service]
ExecStartPre=-/usr/bin/docker rm -f nanocached-node
ExecStartPre=/usr/bin/docker pull ghcr.io/nanocached/nanocached-node@sha256:7ca62b477976620b9a81d8d903683187fa901a4ef09d74f23ab8371b0b147623
ExecStart=/usr/bin/docker run --rm --name nanocached-node \
--network host \
--env-file /etc/nanocached/env \
ghcr.io/nanocached/nanocached-node@sha256:7ca62b477976620b9a81d8d903683187fa901a4ef09d74f23ab8371b0b147623 \
--host 0.0.0.0 --port 8356 --metrics-port 9356 \
--discovery 10.0.0.11:8357 --drain-timeout 20 \
--max-memory 1073741824
ExecStop=/usr/bin/docker stop -t 30 nanocached-node
Restart=always
RestartSec=2
[Install]
WantedBy=multi-user.target
UNIT
systemctl daemon-reload
systemctl enable --now nanocached-node
Launch
The seed goes first, at its fixed address; then a node. Both from the current Amazon Linux 2023 AMI:
AMI_ID=$(aws ssm get-parameter \
--name /aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-x86_64 \
--query Parameter.Value --output text)
# A: the seed — discovery plus a node, at the fixed address every other
# node and every client will name.
SEED_ID=$(aws ec2 run-instances --image-id $AMI_ID --instance-type t3.small \
--subnet-id $NODE_SUBNET_ID --security-group-ids $NODE_SG --private-ip-address 10.0.0.11 \
--iam-instance-profile Name=$NAME-ec2 --user-data file://seed-user-data.sh \
--tag-specifications "ResourceType=instance,Tags=[$TAG,{Key=Name,Value=$NAME-seed}]" \
--query 'Instances[0].InstanceId' --output text)
# B: a plain node.
NODE_ID=$(aws ec2 run-instances --image-id $AMI_ID --instance-type t3.small \
--subnet-id $NODE_SUBNET_ID --security-group-ids $NODE_SG \
--iam-instance-profile Name=$NAME-ec2 --user-data file://node-user-data.sh \
--tag-specifications "ResourceType=instance,Tags=[$TAG,{Key=Name,Value=$NAME-node}]" \
--query 'Instances[0].InstanceId' --output text)
aws ec2 wait instance-running --instance-ids $SEED_ID $NODE_ID
NODE_IP=$(aws ec2 describe-instances --instance-ids $NODE_ID \
--query 'Reservations[0].Instances[0].PrivateIpAddress' --output text)
Timing, from launch: the seed's discovery and node both answered
/readyz 200 and discovery reported two members at
86 s after boot; the plain node registered at about 80 s and
was promoted to a member twelve seconds later, once the seed's own
node had joined ahead of it. The order in the discovery log is the
staged join at work — a node waits until the members before it have
finished handing over.
Verifying
On the seed, both services and the membership count. Discovery
serves /healthz, /readyz and
/metrics on its own operations port, 9357:
aws ssm send-command --instance-ids $SEED_ID --document-name AWS-RunShellScript \
--parameters 'commands=["systemctl is-active nanocached-discovery nanocached-node",
"curl -s -o /dev/null -w discovery-readyz=%{http_code}\\\\n http://127.0.0.1:9357/readyz",
"curl -s -o /dev/null -w node-readyz=%{http_code}\\\\n http://127.0.0.1:9356/readyz",
"curl -s http://127.0.0.1:9357/metrics | grep ^nanocached_discovery_members"]'
# active / active / discovery-readyz=200 / node-readyz=200 / nanocached_discovery_members 2
journalctl -u nanocached-discovery narrates every
membership change (node registered, join
started, join promoted, node left the
cluster) and journalctl -u nanocached-node the
handoffs (migration started/completed, re-replication
after membership change) — the two logs quoted in the rest of
this page.
From the application tier, the SDK's configuration is discovery's
address — 10.0.0.11, port 8357 — and the secret; the SDK
fetches the node list and routes each key to its owners itself (see
the SDKs). This cluster was exercised that way
from an instance in the app subnet with all six SDKs at 0.4.1:
set / get / delete succeed through discovery, a wrong secret is
rejected, and 2,000 keys written before the steps below were read back
complete after each of them.
Growing the cluster
Take an image of a node. --no-reboot keeps it running:
a node keeps nothing on disk, so there is nothing to quiesce, and a
reboot would only cost the cluster a leave-and-rejoin. The image
carries the unit, the secret file and the already-pulled container
image:
# Bake the plain node into an AMI without stopping it: the node keeps
# nothing on disk, so there is nothing to quiesce.
IMAGE_ID=$(aws ec2 create-image --instance-id $NODE_ID --name $NAME-node-$(date +%Y%m%d-%H%M) \
--no-reboot --tag-specifications "ResourceType=image,Tags=[$TAG]" "ResourceType=snapshot,Tags=[$TAG]" \
--query ImageId --output text)
aws ec2 wait image-available --image-ids $IMAGE_ID
Then add a node from it — no user-data, no configuration:
# Add a node from the image: no user-data — the unit, the secret file and
# the pulled image are already on the disk, and the node picks a fresh
# identity on start.
NODE2_ID=$(aws ec2 run-instances --image-id $IMAGE_ID --instance-type t3.small \
--subnet-id $NODE_SUBNET_ID --security-group-ids $NODE_SG \
--iam-instance-profile Name=$NAME-ec2 \
--tag-specifications "ResourceType=instance,Tags=[$TAG,{Key=Name,Value=$NAME-node}]" \
--query 'Instances[0].InstanceId' --output text)
aws ec2 wait instance-running --instance-ids $NODE2_ID
Observed: the image was available three minutes after
create-image; the new instance registered with discovery
37 s after launch and was promoted the same second, the seed's
node handing it 1,328 of the 2,000 test keys (the share the ring now
assigns to it) before the join completed; every key still read back
from the application tier. Repeat the block for each node you want —
the image is good until you change the unit, the secret or the pinned
digest, at which point you bake a new one from a node running the new
configuration.
Removing a node
Stop the service first, then terminate the instance.
systemctl stop sends SIGTERM through docker
stop; the node deregisters and hands its keys to their next
owners within --drain-timeout:
aws ssm send-command --instance-ids $NODE_ID --document-name AWS-RunShellScript \
--parameters 'commands=["systemctl stop nanocached-node"]'
# then, once the command reports Success:
aws ec2 terminate-instances --instance-ids $NODE_ID
Observed with three members: discovery logged node left the
cluster in the same second the stop was issued, the two
remaining nodes re-replicated 672 entries between them so every key
was again on two nodes, and systemctl stop returned after
20 s — the drain budget. The 2,000 keys read back complete
throughout. Terminating an instance without the stop is what an
unplanned loss looks like: discovery drops the node after its
liveness timeout instead of immediately, and the keys are still on
their other owner.
Teardown
Instances first, then the image and its snapshot (deregistering an AMI does not delete the snapshot behind it), then the rest in reverse order. Anything you launched in the app group goes first too.
aws ec2 terminate-instances --instance-ids $SEED_ID $NODE_ID $NODE2_ID
aws ec2 wait instance-terminated --instance-ids $SEED_ID $NODE_ID $NODE2_ID
SNAPSHOT_ID=$(aws ec2 describe-images --image-ids $IMAGE_ID \
--query 'Images[0].BlockDeviceMappings[0].Ebs.SnapshotId' --output text)
aws ec2 deregister-image --image-id $IMAGE_ID
aws ec2 delete-snapshot --snapshot-id $SNAPSHOT_ID
aws iam remove-role-from-instance-profile --instance-profile-name $NAME-ec2 --role-name $NAME-ec2
aws iam delete-instance-profile --instance-profile-name $NAME-ec2
aws iam detach-role-policy --role-name $NAME-ec2 \
--policy-arn arn:aws:iam::aws:policy/AmazonSSMManagedInstanceCore
aws iam delete-role --role-name $NAME-ec2
aws ec2 delete-security-group --group-id $NODE_SG # first: it references $APP_SG
aws ec2 delete-security-group --group-id $APP_SG
for ASSOC in $(aws ec2 describe-route-tables --route-table-ids $RTB_ID \
--query 'RouteTables[0].Associations[?!Main].RouteTableAssociationId' --output text); do
aws ec2 disassociate-route-table --association-id $ASSOC
done
aws ec2 delete-route-table --route-table-id $RTB_ID
aws ec2 detach-internet-gateway --internet-gateway-id $IGW_ID --vpc-id $VPC_ID
aws ec2 delete-internet-gateway --internet-gateway-id $IGW_ID
aws ec2 delete-subnet --subnet-id $NODE_SUBNET_ID
aws ec2 delete-subnet --subnet-id $APP_SUBNET_ID
aws ec2 delete-vpc --vpc-id $VPC_ID
Confirm nothing is left: aws ec2 describe-vpcs --filters
Name=isDefault,Values=false, aws ec2 describe-instances
--filters Name=instance-state-name,Values=running, aws ec2
describe-images --owners self and aws ec2
describe-snapshots --owner-ids self should all come back empty,
and aws iam get-role --role-name nanocached-ec2 should say
NoSuchEntity.