Single node on EC2

The simplest production shape: one nanocached-node on one instance, standalone (no discovery server, no replication — a cache, so losing the instance loses its contents and nothing else). This section builds everything around it from an empty account with the AWS CLI: a VPC with two public subnets (one for the node, one for the application tier), two security groups, an SSM-only instance role, and an Amazon Linux 2023 instance that installs Docker and runs the node under systemd from user-data. Every command was run against a real account in this order and then torn down with the commands at the end.

The shape of the network: the node has a public IP only so it can pull its image and reach Systems Manager — nothing inbound ever uses it. Its security group admits the data port only from the application's security group, the operations port only from inside the VPC, and no SSH at all; you get a shell through SSM Session Manager. The application servers that use the cache — an API or backend tier, whatever holds the SDK — run in their own subnet of the same VPC and carry the nanocached-app security group; the node's rule names that group, not a subnet or an address range, so where the application tier lives is a routing decision and not a firewall one. No NAT gateway: both subnets are public, and outbound (image pull, package installs, Systems Manager) goes straight out through the internet gateway.

export AWS_REGION=us-east-1
NAME=nanocached
TAG="{Key=Project,Value=$NAME}"     # every resource carries it — see teardown

Network

VPC_ID=$(aws ec2 create-vpc --cidr-block 10.0.0.0/16 \
  --tag-specifications "ResourceType=vpc,Tags=[$TAG]" --query Vpc.VpcId --output text)
aws ec2 modify-vpc-attribute --vpc-id $VPC_ID --enable-dns-support
aws ec2 modify-vpc-attribute --vpc-id $VPC_ID --enable-dns-hostnames

NODE_SUBNET_ID=$(aws ec2 create-subnet --vpc-id $VPC_ID --cidr-block 10.0.0.0/24 \
  --availability-zone ${AWS_REGION}a \
  --tag-specifications "ResourceType=subnet,Tags=[$TAG,{Key=Name,Value=$NAME-node}]" \
  --query Subnet.SubnetId --output text)
APP_SUBNET_ID=$(aws ec2 create-subnet --vpc-id $VPC_ID --cidr-block 10.0.1.0/24 \
  --availability-zone ${AWS_REGION}a \
  --tag-specifications "ResourceType=subnet,Tags=[$TAG,{Key=Name,Value=$NAME-app}]" \
  --query Subnet.SubnetId --output text)
aws ec2 modify-subnet-attribute --subnet-id $NODE_SUBNET_ID --map-public-ip-on-launch
aws ec2 modify-subnet-attribute --subnet-id $APP_SUBNET_ID --map-public-ip-on-launch

IGW_ID=$(aws ec2 create-internet-gateway \
  --tag-specifications "ResourceType=internet-gateway,Tags=[$TAG]" \
  --query InternetGateway.InternetGatewayId --output text)
aws ec2 attach-internet-gateway --internet-gateway-id $IGW_ID --vpc-id $VPC_ID

RTB_ID=$(aws ec2 create-route-table --vpc-id $VPC_ID \
  --tag-specifications "ResourceType=route-table,Tags=[$TAG]" --query RouteTable.RouteTableId --output text)
aws ec2 create-route --route-table-id $RTB_ID --destination-cidr-block 0.0.0.0/0 --gateway-id $IGW_ID
aws ec2 associate-route-table --route-table-id $RTB_ID --subnet-id $NODE_SUBNET_ID
aws ec2 associate-route-table --route-table-id $RTB_ID --subnet-id $APP_SUBNET_ID

Name the availability zone explicitly. Left to AWS, the subnet can land in a zone that doesn't offer the instance type you're about to launch (us-east-1e has no t3, and run-instances fails with Unsupported only after the network is built).

Already have a VPC? Set VPC_ID, NODE_SUBNET_ID and APP_SUBNET_ID to existing public subnets (the same one twice is fine for a trial) and skip this block; substitute your VPC's CIDR for 10.0.0.0/16 in the operations-port rule below.

Security groups

Two groups: nanocached-app is a label the application servers wear (no rules of its own); nanocached-node refers to it, so adding an application server is a matter of launching it in the app group, never of editing an IP allow-list.

APP_SG=$(aws ec2 create-security-group --group-name $NAME-app \
  --description "nanocached application tier" --vpc-id $VPC_ID \
  --tag-specifications "ResourceType=security-group,Tags=[$TAG]" --query GroupId --output text)
NODE_SG=$(aws ec2 create-security-group --group-name $NAME-node \
  --description "nanocached node" --vpc-id $VPC_ID \
  --tag-specifications "ResourceType=security-group,Tags=[$TAG]" --query GroupId --output text)

aws ec2 authorize-security-group-ingress --group-id $NODE_SG --ip-permissions \
  "IpProtocol=tcp,FromPort=8356,ToPort=8356,UserIdGroupPairs=[{GroupId=$APP_SG,Description=data port from app tier}]" \
  "IpProtocol=tcp,FromPort=9356,ToPort=9356,IpRanges=[{CidrIp=10.0.0.0/16,Description=operations endpoint from inside the VPC}]"

Port 22 is deliberately absent. The operations endpoint is unauthenticated (see the operations endpoint), which is why its rule is a CIDR inside the VPC rather than 0.0.0.0/0; the default outbound rule (allow all) stays, since the instance needs to pull its image and talk to SSM.

IAM

One role, one managed policy: AmazonSSMManagedInstanceCore is what lets Session Manager and Run Command reach the instance without a key pair or an open port.

aws iam create-role --role-name $NAME-ec2 --tags Key=Project,Value=$NAME \
  --assume-role-policy-document '{"Version":"2012-10-17","Statement":[{"Effect":"Allow",
    "Principal":{"Service":"ec2.amazonaws.com"},"Action":"sts:AssumeRole"}]}'
aws iam attach-role-policy --role-name $NAME-ec2 \
  --policy-arn arn:aws:iam::aws:policy/AmazonSSMManagedInstanceCore
aws iam create-instance-profile --instance-profile-name $NAME-ec2
aws iam add-role-to-instance-profile --instance-profile-name $NAME-ec2 --role-name $NAME-ec2
sleep 15   # a brand-new instance profile takes a few seconds to become launchable

The instance

Save this as node-user-data.sh. It installs Docker, writes the auth secret to a root-only environment file, and registers a systemd unit that runs the published image with host networking (no Docker NAT on the data path) and restarts it if it ever exits.

#!/bin/bash
set -euo pipefail
dnf install -y docker
systemctl enable --now docker

install -d -m 700 /etc/nanocached
cat > /etc/nanocached/env <<'ENV'
NANOCACHED_AUTH_SECRET=change-me
ENV
chmod 600 /etc/nanocached/env

cat > /etc/systemd/system/nanocached-node.service <<'UNIT'
[Unit]
Description=nanocached node
After=docker.service
Requires=docker.service

[Service]
ExecStartPre=-/usr/bin/docker rm -f nanocached-node
ExecStartPre=/usr/bin/docker pull ghcr.io/nanocached/nanocached-node@sha256:7ca62b477976620b9a81d8d903683187fa901a4ef09d74f23ab8371b0b147623
ExecStart=/usr/bin/docker run --rm --name nanocached-node \
  --network host \
  --env-file /etc/nanocached/env \
  ghcr.io/nanocached/nanocached-node@sha256:7ca62b477976620b9a81d8d903683187fa901a4ef09d74f23ab8371b0b147623 \
  --host 0.0.0.0 --port 8356 --metrics-port 9356 \
  --max-memory 1073741824
ExecStop=/usr/bin/docker stop -t 30 nanocached-node
Restart=always
RestartSec=2

[Install]
WantedBy=multi-user.target
UNIT

systemctl daemon-reload
systemctl enable --now nanocached-node

Three things to change for your own deployment:

Launch it:

AMI_ID=$(aws ssm get-parameter \
  --name /aws/service/ami-amazon-linux-latest/al2023-ami-kernel-default-x86_64 \
  --query Parameter.Value --output text)

NODE_ID=$(aws ec2 run-instances --image-id $AMI_ID --instance-type t3.small \
  --subnet-id $NODE_SUBNET_ID --security-group-ids $NODE_SG \
  --iam-instance-profile Name=$NAME-ec2 --user-data file://node-user-data.sh \
  --tag-specifications "ResourceType=instance,Tags=[$TAG,{Key=Name,Value=$NAME-node}]" \
  --query 'Instances[0].InstanceId' --output text)

aws ec2 wait instance-running --instance-ids $NODE_ID
NODE_IP=$(aws ec2 describe-instances --instance-ids $NODE_ID \
  --query 'Reservations[0].Instances[0].PrivateIpAddress' --output text)

From launch to a ready node took about 70 s (Docker install plus the image pull), and the instance was registered with Systems Manager by then, so Run Command could reach it as soon as it was useful.

Verifying

On the node itself (aws ssm start-session --target $NODE_ID for a shell, or aws ssm send-command as below):

aws ssm send-command --instance-ids $NODE_ID --document-name AWS-RunShellScript \
  --parameters 'commands=["systemctl is-active nanocached-node",
    "curl -s -o /dev/null -w healthz=%{http_code}\\\\n http://127.0.0.1:9356/healthz",
    "curl -s -o /dev/null -w readyz=%{http_code}\\\\n http://127.0.0.1:9356/readyz"]'
# active / healthz=200 / readyz=200

From the application tier, the cache is reachable as soon as an instance carries the app group: the SDK's configuration is the node's private IP ($NODE_IP), port 8356 and the secret, and nothing else — see the SDKs for each language and the framework adapters. This node was exercised that way from an instance in the app subnet with all six SDKs at 0.4.1 (Python, TypeScript, Go, Java, .NET and Rust): set / get / delete succeed with the secret, and a wrong secret is rejected with each SDK's authentication error.

The operations endpoint answers from anywhere inside the VPC:

curl -s http://$NODE_IP:9356/metrics | grep -v '^#' | head -3
# nanocached_node_memory_used_bytes 0
# nanocached_node_memory_max_bytes 1073741824
# nanocached_node_entries 0

One more negative check is worth running once, because it is what the security groups are for: from anywhere outside the app group — your laptop, say — both ports refuse:

NODE_PUBLIC_IP=$(aws ec2 describe-instances --instance-ids $NODE_ID \
  --query 'Reservations[0].Instances[0].PublicIpAddress' --output text)
nc -z -w 5 $NODE_PUBLIC_IP 8356 || echo "8356 blocked (expected)"
nc -z -w 5 $NODE_PUBLIC_IP 9356 || echo "9356 blocked (expected)"

(On macOS, nc's -w is an idle timeout, not a connect timeout, so a filtered port keeps it waiting for the kernel's 75 s; use -G 5 there instead and each line comes back in five seconds.)

Finally, systemctl restart nanocached-node on the node: /readyz answered 200 again about half a second after the restart returned (the image is already local, so the container only has to start). Logs are journalctl -u nanocached-node — the container's stderr goes straight to the journal. Note that a standalone node has nobody to hand its entries to, so a restart is a cold cache — that is the trade-off this shape makes, and the reason to graduate to the clustered ECS / EKS shapes below once an empty cache after a restart is no longer acceptable.

Teardown

The reverse order — the instance first (a security group can't be deleted while an ENI references it, and the VPC can't go until everything inside it has; anything you launched in the app group goes first too):

aws ec2 terminate-instances --instance-ids $NODE_ID
aws ec2 wait instance-terminated --instance-ids $NODE_ID

aws iam remove-role-from-instance-profile --instance-profile-name $NAME-ec2 --role-name $NAME-ec2
aws iam delete-instance-profile --instance-profile-name $NAME-ec2
aws iam detach-role-policy --role-name $NAME-ec2 \
  --policy-arn arn:aws:iam::aws:policy/AmazonSSMManagedInstanceCore
aws iam delete-role --role-name $NAME-ec2

aws ec2 delete-security-group --group-id $NODE_SG   # first: it references $APP_SG
aws ec2 delete-security-group --group-id $APP_SG

for ASSOC in $(aws ec2 describe-route-tables --route-table-ids $RTB_ID \
  --query 'RouteTables[0].Associations[?!Main].RouteTableAssociationId' --output text); do
  aws ec2 disassociate-route-table --association-id $ASSOC
done
aws ec2 delete-route-table --route-table-id $RTB_ID
aws ec2 detach-internet-gateway --internet-gateway-id $IGW_ID --vpc-id $VPC_ID
aws ec2 delete-internet-gateway --internet-gateway-id $IGW_ID
aws ec2 delete-subnet --subnet-id $NODE_SUBNET_ID
aws ec2 delete-subnet --subnet-id $APP_SUBNET_ID
aws ec2 delete-vpc --vpc-id $VPC_ID

Confirm nothing is left: aws ec2 describe-vpcs --filters Name=isDefault,Values=false and aws ec2 describe-instances --filters Name=instance-state-name,Values=running should both come back empty, and aws iam get-role --role-name nanocached-ec2 should say NoSuchEntity. (Terminated instances stay visible to the tagging API for about an hour; that is display, not billing.)