SageMaker HyperPod clusters created without a customer VPC
What does ZopNight detect here?
ZopNight flags a SageMaker HyperPod cluster that is in service or scaled to zero and whose `VpcConfig` lists no subnet, so its nodes run in the SageMaker platform VPC rather than yours. Only Slurm clusters can be created that way, since EKS orchestration requires a VPC, and the setting cannot be changed after creation. The finding carries a $0 saving.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-1627 |
| Category | compliance |
| Severity | medium |
| Metric | none — pure configuration read |
| Threshold | no subnet in VpcConfig |
| Source | ZopNight |
| Permissions used | sagemaker:ListClusters · sagemaker:DescribeCluster |
Where it applies
Where a HyperPod cluster runs if you do not choose
A HyperPod cluster is a fleet of accelerated instances for long training runs. Its network placement is decided once, at creation. According to the HyperPod prerequisites, when no VPC configuration is provided, HyperPod defaults to one subnet from the platform VPC. VPC configuration is mandatory for EKS orchestration and optional for Slurm, so the clusters at risk are Slurm clusters created without one.
A cluster outside your VPC cannot be reached through your security groups, cannot use your VPC endpoints to reach S3 or FSx for Lustre privately, and does not appear in your VPC flow logs. For teams training on sensitive data, that is usually a policy violation on its own.
Inspecting cluster network settings
aws sagemaker list-clusters --query 'ClusterSummaries[].ClusterName' --output text
aws sagemaker describe-cluster --cluster-name my-cluster \ --query '{vpc: VpcConfig, groups: InstanceGroups[].[InstanceGroupName, OverrideVpcConfig]}'An empty VpcConfig confirms the cluster is not in a VPC you own.
Which clusters are evaluated
The cluster must be in service, or scaled down to zero nodes; either way the network placement is fixed and non-compliant. ZopNight then checks the subnet it recorded from the cluster’s VPC configuration and flags the cluster when that subnet is confirmed empty.
Known gaps in the check
Clusters that are creating, deleting or failed are ignored. When the cluster description could not be read, no finding is produced.
ZopNight reads only the cluster-level VPC configuration. HyperPod also lets each instance group carry
an OverrideVpcConfig with its own subnets, which is how multi-AZ clusters are set up, but for Slurm
clusters AWS allows an override only when a cluster-level VpcConfig exists. An empty cluster-level
configuration therefore means no instance group is in your VPC either.
The exposure this reflects
This is a network-isolation finding with a $0 saving. The cost of fixing it is migration effort, since the cluster has to be rebuilt.
Rebuilding the cluster inside your VPC
- Plan a window between training jobs; AWS states that once a cluster is created, its VpcConfig settings cannot be modified.
- Prepare private subnets with enough free IP addresses and ENI quota, a restrictive security group, and VPC endpoints for S3 and the SageMaker API.
- Create a new cluster with
aws sagemaker create-cluster ... --vpc-config '{"SecurityGroupIds":["sg-0123"],"Subnets":["subnet-0abc"]}'. - Move checkpoints and jobs across, then delete the old cluster.