Set Up Disaster Recovery Replication

Disaster Recovery (spec.activeRedis.mode: peerof) is the hot-standby topology: an upstream instance fans out directed replication links to one or more downstream instances, each ready to be promoted if the upstream datacenter is lost. Setting it up involves two parts — enabling cross-datacenter replication on each Redis instance, and establishing a connection from the downstream instance to the upstream instance.

For the alternative topology, in which every datacenter accepts writes, see Set Up Active-Active Replication.

Cluster Role Description

In a disaster recovery group, the role of each Redis instance is not fixed; it can be either an upstream or a downstream. The downstream does not restrict data writing, but data written to the downstream will be overwritten by data synchronized from the upstream or cause synchronization failure due to data type conflicts.

Choose the procedure for your Redis version

The setup procedure differs between the two module generations. Follow the section that matches your Redis version.

Redis versionProcedureLink authentication
7.2Redis 7.2 — new moduleA RedisUser bound on each instance
6.0Redis 6.0 — legacy moduleA Secret referenced by the connection
Both ends must run the same Redis version

Replication between Redis 6.0 and Redis 7.2 is not supported. See Module generations and version applicability.

Redis 7.2 — new module

The new module requires a peer-auth credential: every instance in the replication group binds a RedisUser through spec.activeRedis.redisUserName. The operator pushes that credential to every node, and both inbound and outbound peer links authenticate with it. Without it the module rejects all inbound peer replication.

The same username and password in every datacenter

The module accepts an inbound peer only when the username it presents equals the local peer-auth username. Every instance in the group must therefore bind a RedisUser with the same spec.username and the same password value. The RedisUser resource names may differ, but the credential values may not.

The peer-auth RedisUser must satisfy all of the following, otherwise link creation is rejected:

  • spec.redisName references the local instance.
  • spec.accountType is custom — the default and system accounts are refused, because binding them would let any client of that account act as a replication peer.
  • spec.passwordSecrets lists at least one Secret carrying a password key.
  • The Secret is dedicated to this RedisUser; sharing a password Secret across RedisUser resources is rejected by admission for peer-auth users.
  • status.phase is Success.

Step 1: Create the peer-auth credential in each datacenter

Perform these steps in both datacenters, using the same username and the same password value.

# The dedicated password secret for the peer-auth account
$ kubectl -n default create secret generic s72-dc1-peer-secret --from-literal=password='Peer@Repl123'

# The peer-auth RedisUser
$ cat << EOF | kubectl -n default create -f -
apiVersion: redis.middleware.alauda.io/v1
kind: RedisUser
metadata:
  name: s72-dc1-peer
spec:
  accountType: custom
  arch: sentinel
  redisName: s72-dc1
  username: aapeerrepl
  passwordSecrets:
  - s72-dc1-peer-secret
EOF

spec.arch must match the architecture of the referenced instance (sentinel, cluster, or standalone).

Wait for the user to be applied:

$ kubectl -n default get redisusers s72-dc1-peer
NAME           INSTANCE   USERNAME     PHASE     AGE
s72-dc1-peer   s72-dc1    aapeerrepl   Success   12s

Step 2: Enable replication on the upstream instance

$ kubectl -n default patch redis s72-dc1 --type=merge \
    --patch='{"spec": {"activeRedis": {"serviceID": 0, "mode": "peerof", "redisUserName": "s72-dc1-peer"}}}'
WARNING

The range of serviceID is [0-15] and it must be unique within the replication group. It cannot be changed later, and replication cannot be turned off once enabled.

The instance must also satisfy the new module's constraints, which are enforced by admission:

  • customConfig.appendonly must not be yes — the new module is RDB-only.
  • customConfig.databases must be <= 16 — the module refuses to load beyond 16 databases, and the pods would crash-loop.

Step 3: Expose the upstream proxy

After enabling replication, the proxy provides a NodePort access address by default. A single NodePort-based access address is not highly available. In a production environment, use a LoadBalancer address:

$ kubectl -n default patch redis s72-dc1 --type=merge \
    --patch='{"spec": {"activeRedis": {"proxy": {"service": {"type": "LoadBalancer"}}}}}'

Each instance creates a proxy Service named activeredis-proxy-<instance-name>.

For Sentinel instances, the instance name needs to be prefixed with rfr-.

$ kubectl -n default get svc activeredis-proxy-rfr-s72-dc1
NAME                            TYPE           CLUSTER-IP      EXTERNAL-IP     PORT(S)                         AGE
activeredis-proxy-rfr-s72-dc1   LoadBalancer   10.96.145.32    192.168.1.10    6379:31234/TCP,7379:32100/TCP   45s
Both ports must be reachable

The proxy Service of a Redis 7.2 instance exposes two ports: 6379 (RESP) and 7379 (peer). Replication rides the peer port exclusively, so the downstream must be able to reach both. Port 7379 must be exposed unchanged — the module always advertises its own local peer port. Connection admission performs a TCP dial against the peer port and rejects the connection if it is unreachable.

Step 4: Enable replication on the downstream instance

Repeat Step 1 in the downstream datacenter (same username, same password value), then enable replication with a different serviceID:

$ kubectl -n default patch redis s72-dc2 --type=merge \
    --patch='{"spec": {"activeRedis": {"serviceID": 1, "mode": "peerof", "redisUserName": "s72-dc2-peer"}}}'

Step 5: Create the connection on the downstream side

The ActiveRedisConnection is created in the downstream cluster. spec.instance names the local (downstream) instance, and spec.addresses points at the upstream proxy's RESP endpoint.

$ cat << EOF | kubectl -n default create -f -
apiVersion: redis.middleware.alauda.io/v1alpha1
kind: ActiveRedisConnection
metadata:
  name: conn-dc2-to-dc1
spec:
  instance: s72-dc2
  addresses:
  - 192.168.1.10:6379
  peerPort: 7379
  teardownPolicy: Detach
EOF
FieldDescription
instanceThe local (downstream) instance name.
addressesThe upstream proxy's RESP endpoint. The host of the first entry is also used for the peer port dial.
peerPortThe upstream proxy's peer port. Defaults to 7379; override it only when an external load balancer exposes that port under a different number.
secretNameIgnored on the new module. The link authenticates with the instance-level peer-auth binding.
teardownPolicyDetach (default) or Decommission. See Removing a connection.
Connection names are used verbatim as module peer names

The connection name becomes the module's peer name, which is stricter than Kubernetes naming:

  • it must start with a letter — 1conn is rejected;
  • reset, clear, all, and list are reserved words and are rejected.
One connection per instance

An instance may hold at most one ActiveRedisConnection — its own upstream link. A one-upstream-to-many-downstreams fan-out is built by creating a connection in each downstream cluster, all pointing at the same upstream.

Step 6: Verify

$ kubectl -n default get activeredisconnections conn-dc2-to-dc1
NAME              INSTANCE   STATUS    MESSAGE   AGE
conn-dc2-to-dc1   s72-dc2    Healthy             40s

$ kubectl -n default get activeredis
NAME                 INSTANCE   SERVICEID   SHARDS   PEERS   DPEERS   PHASE     MESSAGE   AGE
s72-dc2-activeredis  s72-dc2    1           1        1       0        Healthy             3m

The upstream instance's ActiveRedis resource reports the fan-out through status.downstreamPeerCount.

Inspect the connection for per-shard synchronization detail:

$ kubectl -n default get activeredisconnections conn-dc2-to-dc1 -o yaml
...
status:
  instance: s72-dc2
  credentialMode: peer-auth-global
  shards:
  - index: 0
    offset: "15320"
    opId: "412"
    status: Connected
    syncStatus: PartialSync
  status: Healthy
  upstreamPeer:
    service_id: 0
    service_metadata:
      instance: default/s72-dc1

status.shards[].status indicates the connection status of the shard, and status.shards[].syncStatus indicates the data synchronization status: PartialSync (incremental synchronization of Oplog) or FullSync (RDB is being synchronized).

status.credentialMode: peer-auth-global records that the module-side peer records hold no inline credential — outbound dials read the node-local peer-auth credential, so a rotation is picked up on the next re-dial.

The native Redis Sentinel mode does not have the concept of shards. Here, a master-replica pair of Sentinel is abstracted as shard 0 to be compatible with the same data structure as the Redis cluster mode.

Redis 6.0 — legacy module

Legacy implementation

Redis 6.0 carries the frozen legacy module, kept for compatibility with existing instances. It supports Disaster Recovery only — no Active-Active mode, and no module-level peer authentication. Enabling replication on Redis 6.0 returns an admission warning recommending Redis 7.2. For new deployments, use Redis 7.2.

Upstream side

You need to create a Redis instance first.

CLI
Web Console

Enable Disaster Recovery

# Enable ActiveRedis
$ kubectl -n default patch redis s6 --type=merge --patch='{"spec": {"activeRedis":{"serviceID":0 }}}'

Use LoadBalancer as the access address for the upstream Proxy

After enabling disaster recovery support for the instance, the disaster recovery Proxy provides a NodePort access address by default. A single NodePort-based access address is not highly available. In a production environment, you can use a LoadBalancer address to provide access to the Proxy.

CLI
# Enable ActiveRedis with LoadBalancer as Proxy access type
$ kubectl -n default patch redis s6 --type=merge --patch='{"spec": {"activeRedis":{"proxy": {"service": {"type": "LoadBalancer"}}}}}'

Each disaster recovery instance will create a Proxy Service with a name that follows this format: activeredis-proxy-<instance-name>.

For Sentinel instances, the instance name needs to be prefixed with rfr-

Downstream side

CLI
Web Console

Enable Disaster Recovery

# Enable ActiveRedis
$ kubectl -n default patch redis s6-dest --type=merge --patch='{"spec": {"activeRedis":{"serviceID":0 }}}'
WARNING

The range of serviceID is [0-15]. In the same disaster recovery cluster, the serviceID cannot be repeated.

Create a secret containing the password of the upstream default user

$ kubectl -n default create secret generic s6-src-auth-cred --from-literal=password=abc@123

Configure Disaster Recovery Connection

# create activeredisconnections
$ cat << EOF | kubectl -n default create -f -
apiVersion: redis.middleware.alauda.io/v1alpha1
kind: ActiveRedisConnection
metadata:
name: conn-s6-dest-to-s6-src
spec:
addresses:
- 192.168.1.10:30010
instance: s6-dest
secretName: s6-src-auth-cred
EOF

Note to replace the upstream address and secret.

INFO

spec.secretName is required on Redis 6.0 — it carries the credential the link authenticates with. The Secret must contain a password key, and may optionally contain a username key (the default user is assumed otherwise).

Check Disaster Recovery Connection Status

$ kubectl -n default get activeredisconnections conn-s6-dest-to-s6-src -o yaml
apiVersion: redis.middleware.alauda.io/v1alpha1
kind: ActiveRedisConnection
metadata:
annotations:
  cpaas.io/creator: admin
  cpaas.io/updated-at: "2025-08-26T10:16:49Z"
creationTimestamp: "2025-08-26T10:16:49Z"
generation: 1
labels:
  cpaas.io/activeredis: s6-dest-activeredis
  cpaas.io/activeredis-instance: s6-dest
name: conn-s6-dest-to-s6-src
namespace: default
resourceVersion: "18872971"
uid: 283ac8fa-d693-46ff-989f-68d018888584
spec:
addresses:
- 192.168.1.10:30010
instance: s6-dest
pause: false
secretName: s6-src-auth-cred
status:
instance: s6-dest
shards:
- index: 0
  offset: "0"
  opId: "0"
  status: Connected
  syncStatus: PartialSync
status: Healthy
upstreamPeer:
  service_id: 0
  service_metadata:
    instance: default/s6-src

Here status.shards[0].status indicates the connection status of the shard, and status.shards[0].syncStatus indicates the data synchronization status. The synchronization status can be PartialSync (indicating incremental synchronization of Oplog) or FullSync (indicating that RDB is being synchronized).

The native Redis Sentinel mode does not have the concept of shards. Here, a master-slave of Sentinel is abstracted as shard 0 to be compatible with the same data structure as the Redis cluster mode.

Pre-flight inspection

Before a connection is accepted, the platform runs a set of pre-flight checks against the upstream. The Web Console exposes them through the Inspect button; they also run automatically during ActiveRedisConnection admission, and a failing check rejects the connection with the corresponding message. The checks are backed by the ActiveRedisInspection resource.

Check ItemCheck ContentHandling Method
Network connectionNetwork connectivity check, password checkConfirm that the upstream address is correct; the address is accessible from the downstream side; the credential is correct.
ArchitecturesRedis instance architecture checkConfirm that the instance architectures of the upstream and downstream instances are consistent.
Cluster mode slices inspectCluster mode shard number and slot distribution checkConfirm that the number of shards and slot distribution of the upstream and downstream instances are the same. If not, you can refer to Initialize Cluster Instance Slot Distribution to create a new instance.
Requirements inspectInstance resource rule checkNeed to ensure that the memory resources of the downstream instance must be greater than or equal to the upstream instance.
Config inspectInstance Service ID checkNeed to ensure that in the replication group, the ServiceID is unique and within the range [0-15].

On Redis 7.2 the inspection dials with the peer-auth credential the data path will actually use, so a successful inspection also confirms that both datacenters carry the same credential. It additionally performs a TCP dial against the upstream peer port (7379).

Removing a connection

Deleting an ActiveRedisConnection tears the link down according to its spec.teardownPolicy:

PolicyEffectReversible
Detach (default)Aborts and removes the peer link. The module keeps the peer's bookkeeping, so re-creating the connection later resumes from the retained state.Yes
DecommissionPermanently decommissions the peer. In addition to removing the link, the module stops retaining Oplog, dead-key tombstones, and lag accounting for that peer.No
Decommission is one-way

Decommission is the correct teardown for a datacenter that is gone for good — a detached-but-never-returning peer keeps holding garbage-collection floors on the survivors. It cannot be changed back to Detach, and a decommissioned peer that later returns is treated as a brand-new peer requiring a fresh full synchronization.

Credential rotation

On Redis 7.2, peer links carry no inline credential: outbound dials read the node-local peer-auth credential that the operator re-pushes on every reconcile. To rotate, update the password Secret of each datacenter's peer-auth RedisUser in place, to the same new value, one datacenter at a time. The RedisUser controller replays the ACL, the operator re-pushes the credential, and the proxy accepts the new credential on the next rebuild. Existing links stay healthy, and a link that drops after the rotation re-dials with the current credential.

Rotating an instance's own password Secret is independent and does not affect peer links.