Wire protocol

Each request starts with an ASCII header terminated by \n; key and value bodies follow immediately, read by their declared byte lengths. Bodies have no terminator and may contain arbitrary bytes. One-byte commands keep overhead minimal while frames stay readable during development.

From 1.0.0 the client-facing protocol on this page is frozen; the compatibility policy says what that covers.

Multiple requests per connection and pipelining are both supported. A <key-length> of 0 is rejected for every command; there is no dedicated key or value length limit beyond the overall request-size limit (1 MiB per request on a node).

A — authenticate

A <secret-length>\n<secret>

Responds On\n on success or En\n followed by the server closing the connection (Od\n/Ed\n from a discovery server — the second byte always names the server type, see the note below). If the server has no auth secret configured, A always succeeds. If it does, every other command is rejected with E\n (and the connection closed) until a matching A has been sent.

The second byte (n/d — node/discovery) is how a client discovers what kind of server it reached and adapts, with no configuration; it is sent whether or not the secret was accepted, and whatever form of A was used. SDKs additionally append a T field (A <secret-length> T\n) to request response tags; a server that supports them echoes the capability as OnT/OdT.

G — get

G <key-length>\n<key>

Responds with the value:

V <value-length>\n<value>

or N\n when the key is missing or expired.

S — set

S <key-length> <value-length>\n<key><value>
S <key-length> <value-length> <ttl-seconds>\n<key><value>

Responds S\n. Without a TTL the entry lives until evicted (LRU under the memory bound) or deleted; with one it also expires <ttl-seconds> after the write — a literal 0 is a TTL that has already elapsed, so the entry expires immediately (the next access sees a miss).

D — delete

D <key-length>\n<key>

Responds D\n if the key existed, N\n otherwise.

g / s / d — namespaced get, set, delete

A namespace is a flat, opaque byte string that scopes a key: the same key name in two namespaces is two independent entries. The lowercase command letters are the namespaced counterparts of G/S/D: the header gains one leading <namespace-length> field and the namespace bytes lead the body.

g <namespace-length> <key-length>\n<namespace><key>
s <namespace-length> <key-length> <value-length>\n<namespace><key><value>
s <namespace-length> <key-length> <value-length> <ttl-seconds>\n<namespace><key><value>
d <namespace-length> <key-length>\n<namespace><key>

Responses are exactly those of the uppercase commands. A <namespace-length> of 0 addresses the default namespace — the same one every G/S/D addresses, so g 0 4\nname and G 4\nname are the same request. <key-length> 0 is rejected as always. The namespace is never interpreted: there is no delimiter, no escaping, no hierarchy — it is sliced by its declared length like every other body field, and may contain any bytes.

Namespaces also enter routing (clusters only). Each node's score is fmix64(fnv1a(node-name) XOR key-hash), where the key hash is

Every SDK and the server pin the same test vectors for both forms. Namespaced frames need a server that understands them; a pre-namespace node answers E\n and closes the connection, so upgrade every node before clients start using namespaces.

c / F — clear a namespace, flush everything

c <namespace-length>\n<namespace>
F\n

Both respond C\n. c drops every entry in one namespace — a <namespace-length> of 0 clears the default namespace — and F drops every namespace, the default one included. On the node this is a single sub-map drop (O(1), no key scan, memory reclaimed immediately), so neither command stalls other requests the way a scan-and-delete would. Neither is key-addressed, so W never applies: in a cluster a client sends the command to every member node (a namespace's keys are spread over all of them by rendezvous hashing) and each node clears its own share. A node that is mid-handoff to a joining node replays the clear there too, in order with the keys it is still sending, so a clear racing a join can't leave stale entries on the joiner.

i — increment/decrement a counter

i <namespace-length> <key-length> <delta>\n<namespace><key>

Adds <delta> — a signed decimal i64; a negative <delta> decrements, so there is no separate decrement command — to the key's stored value, in place, and returns the new value:

I <value-length> [<ttl-seconds>]\n<value>

or N\n if the key is missing or expired (INCR never creates a key), or T\n if the key exists but its stored value isn't a decimal-ASCII i64 — or applying <delta> would overflow one. Values must match the same grammar INCR itself writes back: an optional leading -, then digits with no leading zero (other than a lone 0) — no +, no internal sign, no whitespace. <namespace-length> is always present, unlike G/S/D: INCR has no pre-namespace legacy form, so a <namespace-length> of 0 addresses the default namespace like everywhere else. The optional <ttl-seconds> field on a successful reply is the entry's remaining TTL, rounded down to whole seconds, present only when the entry has one — INCR never resets or clears an existing TTL, unlike S. The TTL isn't cosmetic: a client or nanocached-proxy replicating the result to a key's other owners (see below) needs it to avoid silently making a TTL'd counter immortal on every replica.

This is exactly as volatile as a plain SET. LRU eviction and TTL expiry reclaim an incremented value the same as any other entry — INCR makes an operation atomic against concurrent requests on the node that owns the key, not durable or resistant to memory pressure. It is a good fit for rate limiting and approximate counters; it is not a fit for billing or inventory counts, which need an update to survive.

In a cluster, replication is client-side (see Architecture) and INCR is not fanned out like a plain write: only the key's primary owner runs the increment. Once it succeeds, the SDK (or nanocached-proxy in proxy mode) writes the result — the new value and its TTL — to the remaining owners as an ordinary S/s, never by replaying i itself. Re-running the increment on a replica could drift it from the primary (an eviction, or a replica leg that missed an earlier write, leaves it starting from a different base), so every replica always ends up with the exact bytes the primary computed.

k / x — compare-and-set

k <namespace-length> <key-length> <value-length> <cond> [<ttl-seconds>]\n<namespace><key><value>

Stores <value> only if <cond> holds against the key's current stored bytes; otherwise nothing changes. <cond> is one of three bare tokens (not length-prefixed — its own shape identifies it):

On success: S\n, the same acknowledgement a plain S gives. On a condition mismatch: N\n, reusing the same "nothing here to act on" status G/D already use for a miss — k introduces no new response marker. <ttl-seconds> means exactly what it means for S — omitted = no expiry, a literal 0 is a TTL that has already elapsed and the entry expires immediately; unlike INCR, the new value is supplied whole by the caller, so there is no old TTL to preserve.

x <namespace-length> <key-length> <cond> [<tag>]\n<namespace><key>

Removes the key only if <cond> holds — this is the two-argument remove(key, old). <cond> here is always a digest: an absent- or present-only conditioned delete is already the plain, unconditional D. On success: D\n, the same acknowledgement a plain D gives for a key that existed. On a mismatch or a missing key: N\n, the same status D already gives when there was nothing to delete.

Both are always namespaced, same as INCR: neither op has a pre-namespace legacy form, so <namespace-length> is unconditionally present (0 = the default namespace).

The digest. SHA-256 of the key's exact stored bytes — the same bytes a V/I response body would carry (for a compression-enabled client, that includes its marker byte, since the server never decompresses) — truncated to the first 16 bytes (128 bits), lowercase hex-encoded (32 characters). Computed identically by the server and every SDK; a fixed cross-language test vector pins the agreement (SHA-256 of the UTF-8 bytes nanocached-cas-vector truncates to 36287141940ca57acbd7695ccdde9d43).

A client obtains an expected digest from an ordinary GET by hashing the response body itself — there is no separate wire round-trip to fetch one, and GET's response shape is unchanged. This is content-based CAS, so it is exactly as sensitive to encoding as memcached's own value-based CAS: an expected digest reconstructed by re-serializing/re-compressing a value a caller already holds (rather than one taken from a real prior read) is only correct if that reconstruction produces byte-identical output to what the server actually stores — true within one client sharing one serializer/compressor, not guaranteed across languages with client-side compression enabled.

This is not a distributed lock. LRU eviction reclaims a key exactly as it would after a plain SET, CAS or not: if a key used as a lock (add to acquire, a TTL to eventually release) is evicted under memory pressure, a second caller's A-conditioned k succeeds while the first caller still believes it holds the lock — a silent double-acquisition CAS cannot detect. k/x are atomic against concurrent requests on the node that currently owns the key, the same guarantee INCR makes and no stronger.

Replication follows INCR's rule exactly: only the key's primary owner evaluates <cond>. Once it succeeds, the SDK (or nanocached-proxy in proxy mode) writes the result to the remaining owners as an ordinary S/s (for k) or D/d (for x) — never by replaying k/x itself. A replica evaluating the same condition against its own possibly-different copy could reach a different outcome than the primary just did.

m / o — batched get and set

m <namespace-length> <n> <key-length-1> ... <key-length-n>\n<namespace><key-1>...<key-n>

Gets n keys under one frame and, on the node that owns them, one round trip through the cache — instead of n independent g frames. Replies:

M <n> <result-1> ... <result-n>\n<values of the hits, concatenated in request order>

Each <result-i> is one of three tokens, in the same order as the request's keys: a decimal byte length (a hit — that many bytes of the trailing body belong to this key, in order), - (a miss), or W (this node doesn't own this particular key). Issue #125's R and the fatal E still apply to the whole frame, but W never does — a <key-length> being wrong for one key out of a thousand doesn't invalidate the other 999.

o <namespace-length> <n> <key-length-1> <value-length-1> ... <key-length-n> <value-length-n> [<ttl-seconds>]\n<namespace><key-1><value-1>...<key-n><value-n>

Sets n keys the same way, and <ttl-seconds> means exactly what it means for S — omitted = no expiry, a literal 0 is a TTL that has already elapsed and every key in the batch expires immediately. One <ttl-seconds> for the whole batch, not per key — every real caller of a batched set (Django's set_many, cache-manager's mset) already passes one TTL per call, so a per-key TTL field would complicate the frame for no consumer that exists. Replies:

O <n> <result-1> ... <result-n>\n

No body — a set has no value to echo back, unlike m's hits. Each <result-i> is S (stored) or W (wrong node), same per-key independence as m. O is never confused with the On/OnT identity reply A gives: that only ever appears as the very first reply on a connection, right after its A frame, so no other request's reply is ever mistaken for it.

Both are always namespaced, same reasoning as INCR/ k/x: neither has a pre-namespace legacy form, so <namespace-length> is unconditionally present (0 = the default namespace). Both are also restricted to one namespace per frame — rendezvous hashing routes on (namespace, key), so a frame mixing namespaces couldn't route as a single unit anyway, and every real many-key caller already works within one namespace at a time (a cache-manager store, a Django key prefix, a Spring cache name).

A batch never fails as a whole. Every key's outcome — hit/miss/wrong-node for m, stored/wrong-node for o — is independent of every other key's. A client (or nanocached-proxy in proxy mode) that sees a W for some subset of a batch refreshes its routing table and re-issues just those keys, exactly as it would for a single-key G/ S's own W — the rest of the batch's answers stand.

Through nanocached-proxy, a batch spanning multiple owners is split into one sub-frame per node and reassembled before answering the client — the one dispatch shape that turns a single client frame into more than one key's worth of backend traffic. A m sub-frame targets each key's primary only, same as a plain G; a o sub-frame reaches every owner a sub-batch's keys have (primary or replica — the same node can be primary for one key and a replica for another within one batch), and only the primary's outcome decides a key's client-facing status, replica outcomes being logged and swallowed exactly like a plain S's replication already is.

Response tags (tagged mode)

Pipelined clients match responses to requests by order alone — the frames above carry nothing to verify against. A connection whose A carried the T flag and was answered OnT runs in tagged mode: every G/S/D (and g/s/d/c/F/i/k/x/m/o) header must end with a client-chosen tag (a u32 in decimal), and the response echoes it as its own last field, so a client can verify the pairing before trusting the answer.

G <key-length> <tag>\n<key>
S <key-length> <value-length> <tag>\n<key><value>
S <key-length> <value-length> <ttl-seconds> <tag>\n<key><value>
D <key-length> <tag>\n<key>
g <namespace-length> <key-length> <tag>\n<namespace><key>
s <namespace-length> <key-length> <value-length> [<ttl-seconds>] <tag>\n<namespace><key><value>
d <namespace-length> <key-length> <tag>\n<namespace><key>
c <namespace-length> <tag>\n<namespace>
F <tag>\n
i <namespace-length> <key-length> <delta> <tag>\n<namespace><key>
k <namespace-length> <key-length> <value-length> <cond> [<ttl-seconds>] <tag>\n<namespace><key><value>
x <namespace-length> <key-length> <cond> <tag>\n<namespace><key>
m <namespace-length> <n> <key-length-1> ... <key-length-n> <tag>\n<namespace><key-1>...<key-n>
o <namespace-length> <n> <key-length-1> <value-length-1> ... <key-length-n> <value-length-n> [<ttl-seconds>] <tag>\n<namespace><key-1><value-1>...<key-n><value-n>

Responses become V <value-length> <tag>\n, S <tag>\n, D <tag>\n, N <tag>\n, W <tag>\n, and C <tag>\n; B\n stays bare (it is unsolicited, sent before authentication); a retryable error is R <tag>\n (see below). i's reply becomes I <value-length> [<ttl-seconds>] <tag>\n or T <tag>\n — the tag is always the last header field, after the optional TTL. k's reply is S <tag>\n or N <tag>\n; x's is D <tag>\n or N <tag>\n — both reuse the markers above, so tagging them needs no new shape. m's reply is M <n> <result-1> ... <result-n> <tag>\n<hit values>, o's is O <n> <result-1> ... <result-n> <tag>\n — same roster shape as the untagged form, with the tag appended as the header's own last field. Tagged mode is per connection and entirely opt-in — a plain A keeps every frame exactly as documented above. See Echoed response tags for the design.

What an untagged connection cannot promise

Without tags, an SDK pairs each response with the oldest request still waiting, by order alone. When a response of the wrong kind arrives (an S answering a G, say), the SDK knows the streams are misaligned, poisons the connection and redials — but the requests queued behind the mismatched one may already have been handed another request's reply, and a reply of the right kind for the wrong request is indistinguishable from a correct one. That window is inherent to matching by order; it is exactly what tagged mode removes, which is why every SDK asks for T and only falls back to an untagged connection against a server that predates tags (one that rejects the extended A). On such a server the window is the same one the pre-tags protocol always had.

Other statuses you may see

ReplyMeaning
B\nBusy. From a node: the connection limit is reached. From a discovery server: it is inside its startup grace after a restart, re-learning membership — retry shortly or try another replica.
W\nWrong node (clusters only): per the node's own view of membership it does not own this key, so the client's routing table is stale. SDKs respond by refreshing the node list and retrying once.
R\nRetryable error (issue #125): this one request failed transiently — an upstream briefly unreachable behind nanocached-proxy, typically — and should be retried shortly on the same connection; nothing is torn down. Sent only to clients that declared the capability by appending R to their A frame (A <len> [T] R); legacy clients keep getting the fatal E-and-close. Tagged connections receive R <tag>\n. Today the proxy is the only emitter — nodes and discovery accept the capability token but have no transient per-request failure to report.
T\nNot numeric (issue #129, i only): the key exists but its stored value isn't a decimal-ASCII i64, or applying <delta> would overflow one.
E\nError — authentication required or failed, or the request was malformed. Usually followed by the server closing the connection.

Cluster-internal commands

Nodes, discovery servers, and SDKs additionally exchange membership and migration frames — L (node list), J/H/P (join, heartbeat, announce), M/X/C (migrate, cancel, complete), Y/Z/Q (proxy announce, deregister, and list — how nanocached-proxy registers with discovery and how SDKs in proxy mode find one), U/u/V (handoff store, handoff delete, and node leave — the graceful-decommission frames a draining node uses to move its entries and exit membership; see Deployment), and T (node roster: a registered node fetching the member list with every member's membership token, self-identifying the way a heartbeat does — the frames below need those tokens, and discovery hands them only to proven members, never to a client that merely holds the shared secret; see Architecture for what that trust model implies). M, U, and u each carry the receiving node's membership token (U <ns-len> <key-len> <val-len> <token-len> [ttl] [A]), verified before the frame acts, since these are exactly the frames that skip the wrong-node check. U also carries an optional trailing A token: a survivor re-replicating a key after a membership change dropped a member (an eviction, or a leave it did not itself hand off for — see Architecture) sends it instead of a plain overwrite, so a re-replication racing a newer client write can never regress it — the receiver stores the entry only if it was absent, and acks S\n either way, since a key already present there is a success for the sender, not a conflict. An ordinary decommission handoff never sets it. These are versioned with the server and not part of the public caching API, so the compatibility policy does not freeze them; the server source documents them in detail.

A complete session

$ printf 'S 4 5\nnameAliceG 4\nname' | nc 127.0.0.1 8356
S
V 5
Alice

Two pipelined requests — a set and a get — answered in order on one connection.