The proxy tier

nanocached-proxy is optional. A cluster works without it: every SDK fetches the roster from discovery and talks to the nodes directly, one connection per node per client process (architecture). The proxy exists for the case where that stops scaling — a large or elastic application fleet, where the node-side connection count (processes × nodes) grows past what a node should hold — and for the case where the application side can only be given one stable address. It terminates many client connections, routes on their behalf with the same rendezvous hashing the SDKs use, and keeps one pipelined connection per node, so the node-side count becomes the proxy count. To a client it looks like a single node that owns every key.

This page is about where to run it and how many, not how it works. The ECS section was run in that order against a real account and torn down; the EKS section is the same design as Kubernetes manifests and was not run on a cluster; the sidecar section is the alternative, with the constraints that make it the alternative.

When you need one

Where it runs

The proxy is part of the cache cluster, not of the application: it is stateless, it registers itself with discovery, and it needs to reach discovery and the nodes — nothing about the application. So it runs with the cluster, in the cluster's own shape: a third ECS service next to discovery and node, a third Deployment on EKS. Applications on any platform in the same VPC reach it the same way. This keeps one operational model per cluster; a proxy tier on EC2 in front of an ECS cluster would work, but it would mean AMIs, SSM and manual replacement for one component and task definitions for the rest.

How applications find it. A proxy registers its own address with discovery on start. An SDK in proxy mode (via_proxy / viaProxy / ViaProxy — see the SDK docs) is still configured with discovery's address; it fetches the proxy roster instead of the node roster and connects to one proxy, chosen at random, reconnecting to another if it goes away. No Cloud Map entry or Service is needed for the proxies themselves. If you want a single address instead, put an internal Network Load Balancer in front of the proxy tasks and configure the SDK with that address as a plain single node.

How many. The three roles scale for three different reasons, which is why they are three services with three desired counts:

rolecount is set byrule of thumb
discoverynothing — it is the coordinator, joins are staged one at a time through it, a second copy would not help. Its loss stops new client bootstraps until the scheduler replaces it; nodes and proxies keep working on their last roster.1
nodedata: dataset × R ÷ the task's --max-memory, even across zones≥ 2, even
proxyclient connections and throughput: concurrent application connections ÷ --max-connections (1024 per proxy) with headroom; at least one per zone for availability. Node count does not enter into it. A proxy also caps connections per source IP (--max-connections-per-ip, 256 by default, the node's own rule): when application tasks reach it through one NAT or egress IP, that cap is the fleet's ceiling — raise it.≥ 2, one per zone; grow when nanocached_proxy_client_connections approaches _max

For example, 40 application tasks × 8 connections = 320 client connections: two proxies (capacity 2,048) carry that, and growing the dataset changes the node count, not the proxy count.

On ECS: a third service

Everything below assumes the cluster from Cluster on ECS exists as that page builds it — the same variables ($NAME, $CLUSTER_SG, $APP_SG, $EXEC_ARN, $NETCFG), the same subnets, the same secret in Parameter Store. It was run that way: the ECS guide's blocks first, then these.

Security group

Two more rules on the cluster's group: the proxy's client port from the application tier, and its operations port from inside the VPC. Proxy → node and proxy → discovery are already covered by the rule that lets the cluster's group talk to itself on 8356–8357.

aws ec2 authorize-security-group-ingress --group-id $CLUSTER_SG --ip-permissions \
  "IpProtocol=tcp,FromPort=8358,ToPort=8358,UserIdGroupPairs=[{GroupId=$APP_SG,Description=proxy from app tier}]" \
  "IpProtocol=tcp,FromPort=9358,ToPort=9358,IpRanges=[{CidrIp=10.1.0.0/16,Description=proxy operations endpoint from inside the VPC}]"

Task definition

td-proxy.json:

[{
  "name": "nanocached-proxy",
  "image": "ghcr.io/nanocached/nanocached-proxy@sha256:268707fa53a4791ee07f493c8ead20f481106226a95de4ab2cb25db009bddc46",
  "essential": true,
  "command": ["--host", "0.0.0.0", "--port", "8358", "--discovery", "disc.nanocached.local:8357",
              "--drain-timeout", "20", "--metrics-port", "9358"],
  "secrets": [{"name": "NANOCACHED_AUTH_SECRET",
               "valueFrom": "arn:aws:ssm:us-east-1:123456789012:parameter/nanocached/auth-secret"}],
  "portMappings": [{"containerPort": 8358}, {"containerPort": 9358}],
  "stopTimeout": 45,
  "healthCheck": {"command": ["CMD-SHELL", "wget -qO- http://127.0.0.1:9358/readyz || exit 1"],
                  "interval": 10, "timeout": 5, "retries": 3, "startPeriod": 30},
  "logConfiguration": {"logDriver": "awslogs", "options": {"awslogs-group": "/ecs/nanocached",
                       "awslogs-region": "us-east-1", "awslogs-stream-prefix": "proxy"}}
}]
aws ecs register-task-definition --family $NAME-proxy --network-mode awsvpc \
  --requires-compatibilities FARGATE --cpu 256 --memory 512 --execution-role-arn $EXEC_ARN \
  --tags key=Project,value=$NAME --container-definitions file://td-proxy.json

Substitute your account ID in valueFrom, as in the node's definition. The choices that matter:

Service

aws ecs create-service --cluster $NAME --service-name proxy \
  --task-definition $NAME-proxy --desired-count 2 --launch-type FARGATE \
  --network-configuration "$NETCFG" --tags key=Project,value=$NAME
aws ecs wait services-stable --cluster $NAME --services proxy

Observed: the service was stable 43 s after create-service (the cluster it joined had taken 149 s to bring discovery and two nodes up), with the two tasks placed one per zone like the nodes.

Verifying

From anywhere in the VPC, each proxy task's operations port says whether it is ready and how busy it is, and discovery's counts how many are registered — the task IPs come from describe-tasks:

for ip in $(aws ecs describe-tasks --cluster $NAME \
  --tasks $(aws ecs list-tasks --cluster $NAME --service-name proxy --query taskArns --output text) \
  --query 'tasks[].attachments[0].details[?name==`privateIPv4Address`].value' --output text); do
  curl -s -o /dev/null -w "$ip readyz %{http_code}\n" http://$ip:9358/readyz
  curl -s http://$ip:9358/metrics | grep -E '^nanocached_proxy_(client_connections|backend_connections|requests_total) '
done
DISC_IP=$(getent hosts disc.nanocached.local | awk '{print $1}')
curl -s http://$DISC_IP:9357/metrics | grep ^nanocached_discovery_proxies
# nanocached_discovery_proxies 2

From the application tier, the SDK's configuration is the same disc.nanocached.local:8357 and secret as before, plus proxy mode. This cluster was exercised that way from an instance in the app subnet with the Python, TypeScript and Go SDKs: set / get / delete and INCR succeed through a proxy, a wrong secret is rejected at the proxy, and keys written through one SDK read back through the others. Two things are worth seeing once: with proxy mode's random choice, one proxy carried 21,607 requests and the other 200 while three clients ran (per-client stickiness, not per-request balancing — the skew evens out with many clients); and each node's nanocached_node_connections read 2, the proxy count, after all of that traffic. On the 0.4.1 image, a minute with no client connected at all produced WARN backend connection to <node> poisoned; will redial on next request on the first request afterward, for each node it touched — the node had closed the idle connection (its 60 s idle timeout) and the proxy redialed. The get/set churn that produced it saw no failure, but a non-idempotent first request (INCR, CAS) would have been answered with an error, since a frame already written can't be retried — the proxy's own backend connections are exactly as exposed to that 60 s timeout as an SDK's, just with no traffic of their own to keep them busy while no client is attached. Re-run on the 0.4.2 image, the release that shipped the fix (issue #514): a key set, then 100 s — well past the node's 60 s idle timeout — with the proxy otherwise untouched, then a fresh connection sending INCR +1. The log group carried nothing but the three startup lines for the whole window: no poisoned, no redial, and the increment came back correct (5 to 6, proving the node received it) — the proxy's own 30 s probe (the same one the SDKs send) kept both backend connections live under the node's timeout the entire time. nanocached_proxy_backend_connections read 2 throughout, never dropping to redial. A node that closes an idle backend connection anyway — a restart, or a scale-in the proxy's roster hasn't caught up with — is logged at INFO instead of WARN, since no client request failed; that path is the fix's documented fallback, not something this run exercised.

Scaling the proxy tier

One more proxy is --desired-count plus one; it registers with discovery on start and clients pick it up on their next roster fetch. Going the other way is the drain. Verified with six via-proxy churn clients running against the two proxies:

aws ecs update-service --cluster $NAME --service proxy --desired-count 1
aws ecs wait services-stable --cluster $NAME --services proxy

Observed, with three of the six clients on each proxy: the service was stable at one task 7 s after the command; the stopped task's log shows the whole drain —

INFO drain: stop signal received — deregistering and finishing in-flight work
INFO drain complete

— 2 ms apart (nothing was mid-flight at that instant; the budget is for what is), and it exited 0. ECS reported the task stopped 35 s after it began stopping, which is ENI teardown after the exit, not the drain. nanocached_discovery_proxies read 1 while the service was at one task, and the survivor's nanocached_proxy_client_connections went from 3 to 6: the three clients that had been on the stopped proxy reconnected to it. All six churn clients finished their 150 s with 0 failures (5,528–5,592 operations each), and the 20,000 keys written before the scale-in read back complete afterwards. Setting the count back to 2 brought a fresh task up, registered, in 23 s on one run; on another it took 119 s because the first placement failed to pull the image (CannotPullContainerError … not found from ghcr.io for a digest that pulled fine four other times that hour — a registry transient) and ECS placed a replacement a minute later, which was ready on its first check. Discovery counted 2 again both times.

Teardown

The ECS guide's teardown with proxy added to the service loop — it is stopped first, so its drain runs while the nodes are still there to answer what is in flight:

for s in proxy node discovery; do
  aws ecs update-service --cluster $NAME --service $s --desired-count 0
  aws ecs delete-service --cluster $NAME --service $s --force
done
aws ecs wait services-inactive --cluster $NAME --services proxy node discovery

The rest is unchanged: the $NAME- task-definition loop already covers $NAME-proxy, and the security-group rules go with the group.

On EKS: a third Deployment

The same design as Kubernetes manifests, for the cluster from Cluster on EKS. Not run on a cluster — this is the ECS-verified design transcribed, with the same probes, grace period and counts; the EKS-specific parts (pod IP reachability through the VPC CNI, an internal NLB Service) are described, not measured.

apiVersion: apps/v1
kind: Deployment
metadata:
  name: proxy
  namespace: nanocached
spec:
  replicas: 2
  selector:
    matchLabels: {app: nanocached-proxy}
  template:
    metadata:
      labels: {app: nanocached-proxy}
    spec:
      terminationGracePeriodSeconds: 45
      topologySpreadConstraints:
        - maxSkew: 1
          topologyKey: topology.kubernetes.io/zone
          whenUnsatisfiable: DoNotSchedule
          labelSelector:
            matchLabels: {app: nanocached-proxy}
      containers:
        - name: proxy
          image: ghcr.io/nanocached/nanocached-proxy@sha256:268707fa53a4791ee07f493c8ead20f481106226a95de4ab2cb25db009bddc46
          args: ["--host", "0.0.0.0", "--port", "8358", "--discovery", "disc:8357",
                 "--drain-timeout", "20", "--metrics-port", "9358"]
          envFrom:
            - secretRef: {name: nanocached-auth}
          ports:
            - {containerPort: 8358, name: data}
            - {containerPort: 9358, name: ops}
          readinessProbe:
            httpGet: {path: /readyz, port: ops}
            periodSeconds: 5
            failureThreshold: 6         # 30 s of not-ready before it counts
          livenessProbe:
            httpGet: {path: /healthz, port: ops}
            periodSeconds: 10
          resources:
            requests: {cpu: 250m, memory: 256Mi}
            limits: {memory: 512Mi}
---
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
  name: proxy
  namespace: nanocached
spec:
  minAvailable: 1
  selector:
    matchLabels: {app: nanocached-proxy}

The sidecar alternative

A proxy container inside each application task (ECS) or pod (EKS), with the application configured for 127.0.0.1:8358 as a plain single node — not proxy mode, which would pick a random proxy from discovery's roster, most likely on another host. What it buys: no extra network hop, and an application configuration that is one fixed address. What it costs:

Sensible for a small fleet that wants the single-address configuration; not the way to bring the connection count down. On ECS it is one more entry in the application's container definitions, identical to the task definition above minus essential (set it false so an application container's exit, not the proxy's, decides the task's fate) — on EKS, a second container in the application pod, or a DaemonSet with hostPort and application pods pointed at status.hostIP (managed node groups only; not available on Fargate).