Single node on EC2
The simplest production shape: one nanocached-node on one
instance, standalone (no discovery server, no replication — a cache, so
losing the instance loses its contents and nothing else). This section
builds everything around it from an empty account with the AWS CLI:
a VPC with two public subnets (one for the node, one for the
application tier), two security groups, an SSM-only instance
role, and an Amazon Linux 2023 instance that installs Docker and runs
the node under systemd from user-data. Every command was run against a
real account in this order and then torn down with the commands at the
end.
The shape of the network: the node has a public IP only so it can
pull its image and reach Systems Manager — nothing inbound ever uses
it. Its security group admits the data port only from the
application's security group, the operations port only from
inside the VPC, and no SSH at all; you get a shell through SSM Session
Manager. The application servers that use the cache — an API or
backend tier, whatever holds the SDK — run in their own subnet of the
same VPC and carry the nanocached-app security group; the
node's rule names that group, not a subnet or an address range, so
where the application tier lives is a routing decision and not a
firewall one. No NAT gateway: both subnets
are public, and outbound (image pull, package installs, Systems
Manager) goes straight out through the internet gateway.
export AWS_REGION=us-east-1
NAME=nanocached
TAG="{Key=Project,Value=$NAME}" # every resource carries it — see teardown
Network
VPC_ID=$(aws ec2 create-vpc --cidr-block 10.0.0.0/16 \
--tag-specifications "ResourceType=vpc,Tags=[$TAG]" --query Vpc.VpcId --output text)
aws ec2 modify-vpc-attribute --vpc-id $VPC_ID --enable-dns-support
aws ec2 modify-vpc-attribute --vpc-id $VPC_ID --enable-dns-hostnames
NODE_SUBNET_ID=$(aws ec2 create-subnet --vpc-id $VPC_ID --cidr-block 10.0.0.0/24 \
--availability-zone ${AWS_REGION}a \
--tag-specifications "ResourceType=subnet,Tags=[$TAG,{Key=Name,Value=$NAME-node}]" \
--query Subnet.SubnetId --output text)
APP_SUBNET_ID=$(aws ec2 create-subnet --vpc-id $VPC_ID --cidr-block 10.0.1.0/24 \
--availability-zone ${AWS_REGION}a \
--tag-specifications "ResourceType=subnet,Tags=[$TAG,{Key=Name,Value=$NAME-app}]" \
--query Subnet.SubnetId --output text)
aws ec2 modify-subnet-attribute --subnet-id $NODE_SUBNET_ID --map-public-ip-on-launch
aws ec2 modify-subnet-attribute --subnet-id $APP_SUBNET_ID --map-public-ip-on-launch
IGW_ID=$(aws ec2 create-internet-gateway \
--tag-specifications "ResourceType=internet-gateway,Tags=[$TAG]" \
--query InternetGateway.InternetGatewayId --output text)
aws ec2 attach-internet-gateway --internet-gateway-id $IGW_ID --vpc-id $VPC_ID
RTB_ID=$(aws ec2 create-route-table --vpc-id $VPC_ID \
--tag-specifications "ResourceType=route-table,Tags=[$TAG]" --query RouteTable.RouteTableId --output text)
aws ec2 create-route --route-table-id $RTB_ID --destination-cidr-block 0.0.0.0/0 --gateway-id $IGW_ID
aws ec2 associate-route-table --route-table-id $RTB_ID --subnet-id $NODE_SUBNET_ID
aws ec2 associate-route-table --route-table-id $RTB_ID --subnet-id $APP_SUBNET_ID
Name the availability zone explicitly. Left to AWS, the subnet can
land in a zone that doesn't offer the instance type you're about to
launch (us-east-1e has no t3, and
run-instances fails with Unsupported only
after the network is built).
Already have a VPC? Set VPC_ID,
NODE_SUBNET_ID and APP_SUBNET_ID to existing
public subnets (the same one twice is fine for a trial) and skip this
block; substitute your VPC's CIDR for 10.0.0.0/16 in the
operations-port rule below.
Security groups
Two groups: nanocached-app is a label the application
servers wear (no rules of its own); nanocached-node refers
to it, so adding an application server is a matter of launching it in
the app group, never of editing an IP allow-list.
APP_SG=$(aws ec2 create-security-group --group-name $NAME-app \
--description "nanocached application tier" --vpc-id $VPC_ID \
--tag-specifications "ResourceType=security-group,Tags=[$TAG]" --query GroupId --output text)
NODE_SG=$(aws ec2 create-security-group --group-name $NAME-node \
--description "nanocached node" --vpc-id $VPC_ID \
--tag-specifications "ResourceType=security-group,Tags=[$TAG]" --query GroupId --output text)
aws ec2 authorize-security-group-ingress --group-id $NODE_SG --ip-permissions \
"IpProtocol=tcp,FromPort=8356,ToPort=8356,UserIdGroupPairs=[{GroupId=$APP_SG,Description=data port from app tier}]" \
"IpProtocol=tcp,FromPort=9356,ToPort=9356,IpRanges=[{CidrIp=10.0.0.0/16,Description=operations endpoint from inside the VPC}]"
Port 22 is deliberately absent. The operations endpoint is
unauthenticated (see the operations endpoint), which is why its rule is a
CIDR inside the VPC rather than 0.0.0.0/0; the default
outbound rule (allow all) stays, since the instance needs to pull its
image and talk to SSM.
IAM
One role, one managed policy: AmazonSSMManagedInstanceCore
is what lets Session Manager and Run Command reach the instance without
a key pair or an open port.
aws iam create-role --role-name $NAME-ec2 --tags Key=Project,Value=$NAME \
--assume-role-policy-document '{"Version":"2012-10-17","Statement":[{"Effect":"Allow",
"Principal":{"Service":"ec2.amazonaws.com"},"Action":"sts:AssumeRole"}]}'
aws iam attach-role-policy --role-name $NAME-ec2 \
--policy-arn arn:aws:iam::aws:policy/AmazonSSMManagedInstanceCore
aws iam create-instance-profile --instance-profile-name $NAME-ec2
aws iam add-role-to-instance-profile --instance-profile-name $NAME-ec2 --role-name $NAME-ec2
sleep 15 # a brand-new instance profile takes a few seconds to become launchable
The instance
Save this as node-user-data.sh. It installs Docker,
writes the auth secret to a root-only environment file, and registers
a systemd unit that runs the published image with host networking
(no Docker NAT on the data path) and restarts it if it ever exits.
#!/bin/bash
set -euo pipefail
dnf install -y docker
systemctl enable --now docker
install -d -m 700 /etc/nanocached
cat > /etc/nanocached/env <<'ENV'
NANOCACHED_AUTH_SECRET=change-me
ENV
chmod 600 /etc/nanocached/env
cat > /etc/systemd/system/nanocached-node.service <<'UNIT'
[Unit]
Description=nanocached node
After=docker.service
Requires=docker.service
[Service]
ExecStartPre=-/usr/bin/docker rm -f nanocached-node
ExecStartPre=/usr/bin/docker pull ghcr.io/nanocached/nanocached-node@sha256:7ca62b477976620b9a81d8d903683187fa901a4ef09d74f23ab8371b0b147623
ExecStart=/usr/bin/docker run --rm --name nanocached-node \
--network host \
--env-file /etc/nanocached/env \
ghcr.io/nanocached/nanocached-node@sha256:7ca62b477976620b9a81d8d903683187fa901a4ef09d74f23ab8371b0b147623 \
--host 0.0.0.0 --port 8356 --metrics-port 9356 \
--max-memory 1073741824
ExecStop=/usr/bin/docker stop -t 30 nanocached-node
Restart=always
RestartSec=2
[Install]
WantedBy=multi-user.target
UNIT
systemctl daemon-reload
systemctl enable --now nanocached-node
Three things to change for your own deployment:
NANOCACHED_AUTH_SECRET— pick a real one. It is an environment variable rather than a flag so it never shows up inps; the file is0600for the same reason. The application's SDK configuration carries the same value (see authentication).--max-memory— the cache's budget in bytes. 1 GiB is right for a 2 GiBt3.small: leave the kernel, Docker and the per-connection buffers the rest. Scale it with the instance; the capacity planner turns a hit-rate target into a number.- The image — pinned by digest so an instance replacement
can't silently pick up a different build. The digest above is the
0.4.4release (0.4.0was the first tag with--metrics-portand--drain-timeout;0.3.0refuses this command line).docker buildx imagetools inspect ghcr.io/nanocached/nanocached-node:0.4.4prints it, and the same command against a newer tag prints the digest to move to.
Launch it:
AMI_ID=$(aws ssm get-parameter \
--name /aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-x86_64 \
--query Parameter.Value --output text)
NODE_ID=$(aws ec2 run-instances --image-id $AMI_ID --instance-type t3.small \
--subnet-id $NODE_SUBNET_ID --security-group-ids $NODE_SG \
--iam-instance-profile Name=$NAME-ec2 --user-data file://node-user-data.sh \
--tag-specifications "ResourceType=instance,Tags=[$TAG,{Key=Name,Value=$NAME-node}]" \
--query 'Instances[0].InstanceId' --output text)
aws ec2 wait instance-running --instance-ids $NODE_ID
NODE_IP=$(aws ec2 describe-instances --instance-ids $NODE_ID \
--query 'Reservations[0].Instances[0].PrivateIpAddress' --output text)
From launch to a ready node took about 70 s (Docker install plus the image pull), and the instance was registered with Systems Manager by then, so Run Command could reach it as soon as it was useful.
Verifying
On the node itself (aws ssm start-session --target $NODE_ID
for a shell, or aws ssm send-command as below):
aws ssm send-command --instance-ids $NODE_ID --document-name AWS-RunShellScript \
--parameters 'commands=["systemctl is-active nanocached-node",
"curl -s -o /dev/null -w healthz=%{http_code}\\\\n http://127.0.0.1:9356/healthz",
"curl -s -o /dev/null -w readyz=%{http_code}\\\\n http://127.0.0.1:9356/readyz"]'
# active / healthz=200 / readyz=200
From the application tier, the cache is reachable as soon as an
instance carries the app group: the SDK's configuration is the node's
private IP ($NODE_IP), port 8356 and the secret, and
nothing else — see the SDKs for each
language and the framework adapters. This node was exercised that way
from an instance in the app subnet with all six SDKs at 0.4.1 (Python,
TypeScript, Go, Java, .NET and Rust): set / get / delete succeed with
the secret, and a wrong secret is rejected with each SDK's
authentication error.
The operations endpoint answers from anywhere inside the VPC:
curl -s http://$NODE_IP:9356/metrics | grep -v '^#' | head -3
# nanocached_node_memory_used_bytes 0
# nanocached_node_memory_max_bytes 1073741824
# nanocached_node_entries 0
One more negative check is worth running once, because it is what the security groups are for: from anywhere outside the app group — your laptop, say — both ports refuse:
NODE_PUBLIC_IP=$(aws ec2 describe-instances --instance-ids $NODE_ID \
--query 'Reservations[0].Instances[0].PublicIpAddress' --output text)
nc -z -w 5 $NODE_PUBLIC_IP 8356 || echo "8356 blocked (expected)"
nc -z -w 5 $NODE_PUBLIC_IP 9356 || echo "9356 blocked (expected)"
(On macOS, nc's -w is an idle timeout, not
a connect timeout, so a filtered port keeps it waiting for the kernel's
75 s; use -G 5 there instead and each line comes back
in five seconds.)
Finally, systemctl restart nanocached-node on the node:
/readyz answered 200 again about half a second after the
restart returned (the image is already local, so the container only has
to start). Logs are journalctl -u nanocached-node —
the container's stderr goes straight to the journal. Note that a
standalone node has nobody to hand its entries to, so a restart is a
cold cache — that is the trade-off this shape makes, and the reason to
graduate to the clustered ECS / EKS shapes below once an empty cache after a restart is
no longer acceptable.
Teardown
The reverse order — the instance first (a security group can't be deleted while an ENI references it, and the VPC can't go until everything inside it has; anything you launched in the app group goes first too):
aws ec2 terminate-instances --instance-ids $NODE_ID
aws ec2 wait instance-terminated --instance-ids $NODE_ID
aws iam remove-role-from-instance-profile --instance-profile-name $NAME-ec2 --role-name $NAME-ec2
aws iam delete-instance-profile --instance-profile-name $NAME-ec2
aws iam detach-role-policy --role-name $NAME-ec2 \
--policy-arn arn:aws:iam::aws:policy/AmazonSSMManagedInstanceCore
aws iam delete-role --role-name $NAME-ec2
aws ec2 delete-security-group --group-id $NODE_SG # first: it references $APP_SG
aws ec2 delete-security-group --group-id $APP_SG
for ASSOC in $(aws ec2 describe-route-tables --route-table-ids $RTB_ID \
--query 'RouteTables[0].Associations[?!Main].RouteTableAssociationId' --output text); do
aws ec2 disassociate-route-table --association-id $ASSOC
done
aws ec2 delete-route-table --route-table-id $RTB_ID
aws ec2 detach-internet-gateway --internet-gateway-id $IGW_ID --vpc-id $VPC_ID
aws ec2 delete-internet-gateway --internet-gateway-id $IGW_ID
aws ec2 delete-subnet --subnet-id $NODE_SUBNET_ID
aws ec2 delete-subnet --subnet-id $APP_SUBNET_ID
aws ec2 delete-vpc --vpc-id $VPC_ID
Confirm nothing is left: aws ec2 describe-vpcs --filters
Name=isDefault,Values=false and aws ec2 describe-instances
--filters Name=instance-state-name,Values=running should both
come back empty, and aws iam get-role --role-name
nanocached-ec2 should say NoSuchEntity. (Terminated
instances stay visible to the tagging API for about an hour; that is
display, not billing.)