Cluster on ECS (Fargate, multi-AZ)

A replicated cluster on Fargate: one discovery task and a node service, spread across two availability zones. Discovery gets a stable name through Cloud Map private DNS; nodes find it by that name, join the ring with a staged handoff, and hand their entries off on the way out when ECS stops them. The node service grows and shrinks by its desired-count — this page builds the cluster and grows it by one; wiring the count to a metric (the memory watermark, see metrics to scale on) is Application Auto Scaling configuration on top of exactly this and is not covered here. Every block below was run in this order against a real account and torn down with the commands at the end.

What the pieces are for: two cache subnets in two zones so a zone loss costs at most half the ring (R=2 means the surviving copy of every key is elsewhere) and a third subnet for the application tier; one security group for the cluster's own traffic and the same nanocached-app label as the EC2 guides for the application servers; the auth secret in Parameter Store, pulled into the tasks by the execution role so it appears in no task definition; a health check on /readyz so the scheduler only routes to nodes that have joined; and stopTimeout sized to the drain.

export AWS_REGION=us-east-1
ACCOUNT=$(aws sts get-caller-identity --query Account --output text)
NAME=nanocached
TAG="{Key=Project,Value=$NAME}"

Network

VPC_ID=$(aws ec2 create-vpc --cidr-block 10.1.0.0/16 \
  --tag-specifications "ResourceType=vpc,Tags=[$TAG]" --query Vpc.VpcId --output text)
aws ec2 modify-vpc-attribute --vpc-id $VPC_ID --enable-dns-support
aws ec2 modify-vpc-attribute --vpc-id $VPC_ID --enable-dns-hostnames   # Cloud Map needs both

# Two cache subnets in two zones for the tasks, one for the application tier.
SUBNET_A=$(aws ec2 create-subnet --vpc-id $VPC_ID --cidr-block 10.1.0.0/24 \
  --availability-zone ${AWS_REGION}a \
  --tag-specifications "ResourceType=subnet,Tags=[$TAG,{Key=Name,Value=$NAME-cache-a}]" \
  --query Subnet.SubnetId --output text)
SUBNET_B=$(aws ec2 create-subnet --vpc-id $VPC_ID --cidr-block 10.1.1.0/24 \
  --availability-zone ${AWS_REGION}b \
  --tag-specifications "ResourceType=subnet,Tags=[$TAG,{Key=Name,Value=$NAME-cache-b}]" \
  --query Subnet.SubnetId --output text)
APP_SUBNET_ID=$(aws ec2 create-subnet --vpc-id $VPC_ID --cidr-block 10.1.2.0/24 \
  --availability-zone ${AWS_REGION}a \
  --tag-specifications "ResourceType=subnet,Tags=[$TAG,{Key=Name,Value=$NAME-app}]" \
  --query Subnet.SubnetId --output text)
for s in $SUBNET_A $SUBNET_B $APP_SUBNET_ID; do
  aws ec2 modify-subnet-attribute --subnet-id $s --map-public-ip-on-launch
done

IGW_ID=$(aws ec2 create-internet-gateway \
  --tag-specifications "ResourceType=internet-gateway,Tags=[$TAG]" \
  --query InternetGateway.InternetGatewayId --output text)
aws ec2 attach-internet-gateway --internet-gateway-id $IGW_ID --vpc-id $VPC_ID
RTB_ID=$(aws ec2 create-route-table --vpc-id $VPC_ID \
  --tag-specifications "ResourceType=route-table,Tags=[$TAG]" --query RouteTable.RouteTableId --output text)
aws ec2 create-route --route-table-id $RTB_ID --destination-cidr-block 0.0.0.0/0 --gateway-id $IGW_ID
for s in $SUBNET_A $SUBNET_B $APP_SUBNET_ID; do
  aws ec2 associate-route-table --route-table-id $RTB_ID --subnet-id $s
done

Public subnets with assignPublicIp=ENABLED on the services is how the tasks pull their image and reach Parameter Store and CloudWatch without a NAT gateway; nothing inbound uses those addresses (the security group sees to that). The application tier has its own subnet for the same reason it does in the EC2 guides — where it lives is a routing decision; what may reach the cluster is the security group's.

Security groups

APP_SG=$(aws ec2 create-security-group --group-name $NAME-app \
  --description "nanocached application tier" --vpc-id $VPC_ID \
  --tag-specifications "ResourceType=security-group,Tags=[$TAG]" --query GroupId --output text)
CLUSTER_SG=$(aws ec2 create-security-group --group-name $NAME-cluster \
  --description "nanocached cluster tasks" --vpc-id $VPC_ID \
  --tag-specifications "ResourceType=security-group,Tags=[$TAG]" --query GroupId --output text)

aws ec2 wait security-group-exists --group-ids $CLUSTER_SG $APP_SG   # the groups are not always visible to the next call yet
aws ec2 authorize-security-group-ingress --group-id $CLUSTER_SG --ip-permissions \
  "IpProtocol=tcp,FromPort=8356,ToPort=8357,UserIdGroupPairs=[{GroupId=$CLUSTER_SG,Description=node/discovery traffic within the cluster}]" \
  "IpProtocol=tcp,FromPort=8356,ToPort=8357,UserIdGroupPairs=[{GroupId=$APP_SG,Description=data + discovery from app tier}]" \
  "IpProtocol=tcp,FromPort=9356,ToPort=9357,IpRanges=[{CidrIp=10.1.0.0/16,Description=operations endpoints from inside the VPC}]"

The cluster group refers to itself: nodes heartbeat to discovery on 8357 and hand entries to each other on 8356. The application tier needs both ports too — its SDK fetches the roster from discovery and then talks to nodes directly. The wait is there because in one run the authorize call, issued straight after the two create-security-group calls, failed and succeeded on re-run: a new group is not always visible to the next API call.

The auth secret

aws ssm put-parameter --name /$NAME/auth-secret --type SecureString \
  --value change-me --tags Key=Project,Value=$NAME

Both task definitions reference this parameter in secrets; ECS resolves it at task start and injects it as NANOCACHED_AUTH_SECRET. To rotate it, put a new value and force a new deployment of both services (aws ecs update-service --force-new-deployment) — nodes and discovery must agree.

IAM

The execution role is what ECS itself uses to pull the image, write logs and read the secret — the tasks have no role of their own, since nanocached calls no AWS API. The instance role is for the throwaway client below, as in the EC2 guide.

aws iam create-role --role-name $NAME-ecs-execution --tags Key=Project,Value=$NAME \
  --assume-role-policy-document '{"Version":"2012-10-17","Statement":[{"Effect":"Allow",
    "Principal":{"Service":"ecs-tasks.amazonaws.com"},"Action":"sts:AssumeRole"}]}'
aws iam attach-role-policy --role-name $NAME-ecs-execution \
  --policy-arn arn:aws:iam::aws:policy/service-role/AmazonECSTaskExecutionRolePolicy
aws iam put-role-policy --role-name $NAME-ecs-execution --policy-name read-auth-secret \
  --policy-document "{\"Version\":\"2012-10-17\",\"Statement\":[{\"Effect\":\"Allow\",
    \"Action\":\"ssm:GetParameters\",
    \"Resource\":\"arn:aws:ssm:$AWS_REGION:$ACCOUNT:parameter/$NAME/auth-secret\"}]}"

aws iam create-role --role-name $NAME-ec2 --tags Key=Project,Value=$NAME \
  --assume-role-policy-document '{"Version":"2012-10-17","Statement":[{"Effect":"Allow",
    "Principal":{"Service":"ec2.amazonaws.com"},"Action":"sts:AssumeRole"}]}'
aws iam attach-role-policy --role-name $NAME-ec2 \
  --policy-arn arn:aws:iam::aws:policy/AmazonSSMManagedInstanceCore
aws iam create-instance-profile --instance-profile-name $NAME-ec2
aws iam add-role-to-instance-profile --instance-profile-name $NAME-ec2 --role-name $NAME-ec2
sleep 15

Cloud Map, cluster, logs

NS_OP=$(aws servicediscovery create-private-dns-namespace --name $NAME.local --vpc $VPC_ID \
  --tag Key=Project,Value=$NAME --query OperationId --output text)
until [ "$(aws servicediscovery get-operation --operation-id $NS_OP \
  --query Operation.Status --output text)" = SUCCESS ]; do sleep 5; done
NS_ID=$(aws servicediscovery get-operation --operation-id $NS_OP \
  --query 'Operation.Targets.NAMESPACE' --output text)
DISC_SD_ARN=$(aws servicediscovery create-service --name disc --namespace-id $NS_ID \
  --dns-config "RoutingPolicy=MULTIVALUE,DnsRecords=[{Type=A,TTL=5}]" \
  --health-check-custom-config FailureThreshold=1 \
  --tags Key=Project,Value=$NAME --query Service.Arn --output text)

aws ecs create-cluster --cluster-name $NAME --tags key=Project,value=$NAME
aws logs create-log-group --log-group-name /ecs/$NAME --tags Project=$NAME
aws logs put-retention-policy --log-group-name /ecs/$NAME --retention-in-days 7

Discovery is therefore disc.nanocached.local:8357 from anywhere in the VPC. A 5-second TTL keeps a discovery restart short for the nodes' heartbeats; the custom health check lets ECS withdraw the record the moment the task stops. The namespace is a Route 53 private hosted zone; if this block is interrupted after create-private-dns-namespace, re-running it fails with ConflictingDomainExists — take NS_ID from aws servicediscovery list-namespaces instead (the new namespace can take a minute to appear there).

Task definitions

td-discovery.json:

[{
  "name": "nanocached-discovery",
  "image": "ghcr.io/nanocached/nanocached-discovery@sha256:f3368f0d35fdac03264915bdfe6818b458f14c25068f069aa79bcb2c22805288",
  "essential": true,
  "command": ["--host", "0.0.0.0", "--port", "8357", "--replication-factor", "2",
              "--metrics-port", "9357"],
  "secrets": [{"name": "NANOCACHED_AUTH_SECRET",
               "valueFrom": "arn:aws:ssm:us-east-1:123456789012:parameter/nanocached/auth-secret"}],
  "portMappings": [{"containerPort": 8357}, {"containerPort": 9357}],
  "healthCheck": {"command": ["CMD-SHELL", "wget -qO- http://127.0.0.1:9357/healthz || exit 1"],
                  "interval": 10, "timeout": 5, "retries": 3, "startPeriod": 30},
  "logConfiguration": {"logDriver": "awslogs", "options": {"awslogs-group": "/ecs/nanocached",
                       "awslogs-region": "us-east-1", "awslogs-stream-prefix": "discovery"}}
}]

td-node.json:

[{
  "name": "nanocached-node",
  "image": "ghcr.io/nanocached/nanocached-node@sha256:7ca62b477976620b9a81d8d903683187fa901a4ef09d74f23ab8371b0b147623",
  "essential": true,
  "command": ["--host", "0.0.0.0", "--port", "8356", "--discovery", "disc.nanocached.local:8357",
              "--drain-timeout", "20", "--metrics-port", "9356", "--max-memory", "536870912"],
  "secrets": [{"name": "NANOCACHED_AUTH_SECRET",
               "valueFrom": "arn:aws:ssm:us-east-1:123456789012:parameter/nanocached/auth-secret"}],
  "portMappings": [{"containerPort": 8356}, {"containerPort": 9356}],
  "stopTimeout": 45,
  "healthCheck": {"command": ["CMD-SHELL", "wget -qO- http://127.0.0.1:9356/readyz || exit 1"],
                  "interval": 10, "timeout": 5, "retries": 3, "startPeriod": 120},
  "logConfiguration": {"logDriver": "awslogs", "options": {"awslogs-group": "/ecs/nanocached",
                       "awslogs-region": "us-east-1", "awslogs-stream-prefix": "node"}}
}]

Substitute your account ID in the two valueFrom ARNs. The choices that matter:

EXEC_ARN=arn:aws:iam::$ACCOUNT:role/$NAME-ecs-execution
aws ecs register-task-definition --family $NAME-discovery --network-mode awsvpc \
  --requires-compatibilities FARGATE --cpu 256 --memory 512 --execution-role-arn $EXEC_ARN \
  --tags key=Project,value=$NAME --container-definitions file://td-discovery.json
aws ecs register-task-definition --family $NAME-node --network-mode awsvpc \
  --requires-compatibilities FARGATE --cpu 512 --memory 1024 --execution-role-arn $EXEC_ARN \
  --tags key=Project,value=$NAME --container-definitions file://td-node.json

Services

NETCFG="awsvpcConfiguration={subnets=[$SUBNET_A,$SUBNET_B],securityGroups=[$CLUSTER_SG],assignPublicIp=ENABLED}"

aws ecs create-service --cluster $NAME --service-name discovery \
  --task-definition $NAME-discovery --desired-count 1 --launch-type FARGATE \
  --network-configuration "$NETCFG" --service-registries registryArn=$DISC_SD_ARN \
  --tags key=Project,value=$NAME
aws ecs wait services-stable --cluster $NAME --services discovery

# services-stable returns before the Cloud Map record exists — wait for it
until [ "$(aws servicediscovery discover-instances --namespace-name $NAME.local \
  --service-name disc --health-status HEALTHY --query 'length(Instances)' --output text)" -ge 1 ]
do sleep 5; done

aws ecs create-service --cluster $NAME --service-name node \
  --task-definition $NAME-node --desired-count 2 --launch-type FARGATE \
  --network-configuration "$NETCFG" --tags key=Project,value=$NAME
aws ecs wait services-stable --cluster $NAME --services node

The wait on discover-instances is not decoration. In verification, services-stable came back 40 s before the DNS record did, and without the wait one of the two first node tasks started, looked up disc.nanocached.local, got Name does not resolve, and kept getting it every 5 s for over three minutes — while the other task, and every task launched later, resolved it at once. The health check caught it (failed container health checks in the service events) and ECS replaced the task, which joined normally; but that is a self-healing detour, not a start. With the wait, both tasks resolved the name on their first try and were healthy inside a minute. A service launched across two subnets is placed across both zones by default; describe-tasks shows availabilityZone per task if you want to see it.

The stuck task is not the node holding on to a stale answer. Against a local DNS server that answered NXDOMAIN and then published the record, both the Alpine image and a glibc build re-query the name on every 5 s retry and join on the first retry after the record appears (issue #174). What keeps a Fargate task stuck for minutes is most likely a negative-cache entry in the VPC resolver its ENI talks to — the other zone's task used a different resolver, and the replacement task started after the entry had expired. Such a task recovers by itself once the cache expires; a longer node startPeriod stops ECS from replacing it in the meantime, at the cost of slower replacement of a task that is genuinely broken.

Verifying

Discovery's name and its metrics answer the first question — how many nodes are in the ring. From anywhere in the VPC:

DISC_IP=$(getent hosts disc.nanocached.local | awk '{print $1}')
curl -s http://$DISC_IP:9357/metrics | grep ^nanocached_discovery_members
# nanocached_discovery_members 2

Each task's log is in the one log group, prefixed by service. describe-tasks shows the zone each task landed in:

aws logs tail /ecs/$NAME --since 10m --format short
aws logs filter-log-events --log-group-name /ecs/$NAME \
  --log-stream-name-prefix node/ --filter-pattern '"joined the cluster"'
aws ecs describe-tasks --cluster $NAME --query 'tasks[].[availabilityZone,healthStatus]' --output text \
  --tasks $(aws ecs list-tasks --cluster $NAME --service-name node --query taskArns --output text)

Observed: the node service reached a steady state 2 min 40 s after the first create-service (of which the Cloud Map wait was about 35 s); the two tasks were placed one per zone, both HEALTHY, and both logged joined the cluster via discovery at disc.nanocached.local:8357 on their first attempt — zero Name does not resolve lines in the log group.

From the application tier, the SDK's configuration is disc.nanocached.local, port 8357 and the secret; the SDK sees from the handshake that it is talking to a discovery server, fetches the roster and connects to the nodes itself (see the SDKs). This cluster was exercised that way from an instance in the app subnet with all six SDKs at 0.4.1: set / get / delete succeed through discovery, a wrong secret is rejected, and 20,000 keys written before the next step were read back complete after it.

Growing the cluster

One more node is one more task:

aws ecs update-service --cluster $NAME --service node --desired-count 3
aws ecs wait services-stable --cluster $NAME --services node

Observed: the third task was placed in the zone with one task, registered with discovery 24 s after ECS created it, and was promoted the same second; each of the two members handed it 13,367 keys — its share of the ring under R=2 — before the join completed (migration completed … sent 13367 keys in both members' logs, join promoted … members now 3 in discovery's), and all 20,000 keys read back from the application tier. Going the other way is the same command with a smaller count: ECS sends SIGTERM, the node runs its planned leave inside --drain-timeout, and stopTimeout gives it the room (graceful scale-in). What that looks like in a leaver's log, from this cluster's teardown:

INFO shutdown signal received — decommissioning (budget 20s)
INFO decommission: handed off 13288 entr(ies)
INFO decommission: forwarding window open for 10s

(At teardown every task, discovery included, is stopped at once, so the leavers also log a WARN decommission: leave notification to disc.nanocached.local:8357 failed — the record was already withdrawn. Under a normal scale-in discovery is up and logs node left the cluster instead.)

The same drain with a bug in it is instructive: this guide's first run hit a bug in which an authenticated node could not fetch the roster on the way out (#171, fixed in the images pinned above). The leavers exited 0 on time, ECS was satisfied, and 2,736 of the 20,000 keys were gone. Watch the leaver's log, not just the exit code.

Teardown

for s in node discovery; do
  aws ecs update-service --cluster $NAME --service $s --desired-count 0
  aws ecs delete-service --cluster $NAME --service $s --force
done
aws ecs wait services-inactive --cluster $NAME --services node discovery
aws ec2 terminate-instances --instance-ids $APP_ID

for arn in $(aws ecs list-task-definitions --family-prefix $NAME- --query taskDefinitionArns --output text); do
  aws ecs deregister-task-definition --task-definition $arn
  aws ecs delete-task-definitions --task-definitions $arn
done
aws ecs delete-cluster --cluster $NAME

SD_ID=${DISC_SD_ARN##*/}
for inst in $(aws servicediscovery list-instances --service-id $SD_ID --query 'Instances[].Id' --output text); do
  aws servicediscovery deregister-instance --service-id $SD_ID --instance-id $inst
done
aws servicediscovery delete-service --id $SD_ID
aws servicediscovery delete-namespace --id $NS_ID
aws logs delete-log-group --log-group-name /ecs/$NAME
aws ssm delete-parameter --name /$NAME/auth-secret

aws iam delete-role-policy --role-name $NAME-ecs-execution --policy-name read-auth-secret
aws iam detach-role-policy --role-name $NAME-ecs-execution \
  --policy-arn arn:aws:iam::aws:policy/service-role/AmazonECSTaskExecutionRolePolicy
aws iam delete-role --role-name $NAME-ecs-execution
aws iam remove-role-from-instance-profile --instance-profile-name $NAME-ec2 --role-name $NAME-ec2
aws iam delete-instance-profile --instance-profile-name $NAME-ec2
aws iam detach-role-policy --role-name $NAME-ec2 \
  --policy-arn arn:aws:iam::aws:policy/AmazonSSMManagedInstanceCore
aws iam delete-role --role-name $NAME-ec2

aws ec2 wait instance-terminated --instance-ids $APP_ID
# stopped tasks' ENIs linger for a minute or two; the security group can't go until they have
until [ "$(aws ec2 describe-network-interfaces --filters Name=vpc-id,Values=$VPC_ID \
  --query 'length(NetworkInterfaces)' --output text)" = 0 ]; do sleep 5; done
aws ec2 delete-security-group --group-id $CLUSTER_SG
aws ec2 delete-security-group --group-id $APP_SG
for a in $(aws ec2 describe-route-tables --route-table-ids $RTB_ID \
  --query 'RouteTables[0].Associations[?!Main].RouteTableAssociationId' --output text); do
  aws ec2 disassociate-route-table --association-id $a
done
aws ec2 delete-route-table --route-table-id $RTB_ID
aws ec2 detach-internet-gateway --internet-gateway-id $IGW_ID --vpc-id $VPC_ID
aws ec2 delete-internet-gateway --internet-gateway-id $IGW_ID
aws ec2 delete-subnet --subnet-id $SUBNET_A
aws ec2 delete-subnet --subnet-id $SUBNET_B
aws ec2 delete-subnet --subnet-id $APP_SUBNET_ID
aws ec2 delete-vpc --vpc-id $VPC_ID

Confirm: aws ecs describe-clusters --clusters nanocached reports INACTIVE, aws servicediscovery list-namespaces no longer lists nanocached.local (the delete is asynchronous — give it half a minute), describe-log-groups, get-parameter and get-role for the names above all come back not found, and there is no non-default VPC.