Dev and test ElastiCache replication groups paying for standby replica nodes
What does ZopNight detect here?
ZopNight flags ElastiCache replication groups with Multi-AZ enabled that carry an explicit dev or test environment tag. Multi-AZ itself is free; the cost is the replica nodes it requires, so the saving is the group cost times replicas per shard divided by replicas plus one, for example half the cost with one replica per shard.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-053 |
| Category | rightsizing |
| Severity | low |
| Metric | none — pure configuration read |
| Threshold | Multi-AZ enabled, dev/test env tag, replicas per shard > 0 |
| Source | ZopNight |
| Permissions used | elasticache:DescribeReplicationGroups · elasticache:ListTagsForResource |
Where it applies
Multi-AZ costs nothing; its replicas cost a node each
ElastiCache bills per node hour, and every node in a replication group is the same type. Multi-AZ with automatic failover is a setting that promotes a read replica when the primary fails, and it only works if replicas exist. The replicas are what you pay for. A cluster-mode-disabled group with one primary and one replica costs twice what a lone primary would.
For a production cache that trade is worth it. For a dev or test cache that can be rebuilt from its source of truth, the standby nodes are rarely used.
Finding Multi-AZ groups and their replica count
aws elasticache describe-replication-groups \ --query 'ReplicationGroups[?MultiAZ==`enabled`].[ReplicationGroupId,CacheNodeType,length(MemberClusters)]' \ --output table
aws elasticache list-tags-for-resource --resource-name REPLICATION_GROUP_ARNThe first lists Multi-AZ groups with their node type and node count; the second shows the tags that decide whether a group is non-production.
The environment test is strict
- Multi-AZ is enabled on the replication group.
- An environment tag (
env,environment,stageortier) exists. A production value vetoes the finding; a dev or test value is required for it to fire. - The group’s replicas per shard is known and above zero.
- The group’s cost can be priced from its member nodes.
The cache’s name is deliberately not used. A production cache named something like test-results
would otherwise be stripped of its failover.
Groups left alone
No environment tag, a production tag, or a tag with any other value means no finding. So do an unknown replica count and nodes that cannot be priced. Idle caches, whatever their environment, are covered by ElastiCache Cluster Idle.
Removing the replicas, priced by node count
saving = group monthly cost x replicas per shard / (replicas per shard + 1)With one replica per shard that is half the group cost; with two it is two thirds. The shard count cancels out because every shard has the same shape.
Dropping the standby nodes
- Confirm with the owner that the cache holds nothing that cannot be reloaded.
- Turn off Multi-AZ and automatic failover first:
aws elasticache modify-replication-group --replication-group-id cart-dev --no-multi-az-enabled --no-automatic-failover-enabled --apply-immediately - Remove the replicas:
aws elasticache decrease-replica-count --replication-group-id cart-dev --new-replica-count 0 --apply-immediately - Check the primary endpoint still serves traffic.