AWS Glue jobs allocated more DPUs than their peak demand for Spark executors needs
What does ZopNight detect here?
ZopNight compares an AWS Glue job's configured DPUs with the peak `numberMaxNeededExecutors` it reports to CloudWatch over 30 days. When the job asks for fewer executors than it is given, ZopNight prices the unused DPUs at the job's DPU-hour rate and recommends a smaller MaxCapacity that never drops below 10 DPUs or measured peak need.
Signal and threshold
| Field | Value |
|---|---|
| Rule IDs | RC-098 |
| Category | rightsizing |
| Severity | low |
| Metric | glue.driver.ExecutorAllocationManager.executors.numberMaxNeededExecutors |
| Threshold | peak needed executors below allocated executors |
| Evaluation window | 30d |
| Source | ZopNight |
| Permissions used | glue:GetJobs · glue:GetJobRuns · cloudwatch:GetMetricStatistics |
Where it applies
Every allocated DPU is billed for the whole run
AWS Glue bills an hourly rate per Data Processing Unit, and one DPU is 4 vCPU and 16 GB of memory. Billing is per second, with a 1-minute minimum on Glue 2.0 and later. AWS’s own example: a Spark job using 6 DPUs for 15 minutes at $0.44 per DPU-hour costs $0.66. By default Glue gives a Spark job 10 DPUs.
The bill follows the allocation, not the work. If a job is given 40 DPUs and its stages never need more than a dozen executors, the other capacity idles on every run and is still paid for.
Reading peak Spark executors from CloudWatch
Glue publishes two gauges that answer the question directly.
numberAllExecutors
is the number of executors running, and numberMaxNeededExecutors is the number of running plus
pending executors needed to satisfy the current load. Glue’s capacity-planning guide says to
compare the two. Job metrics must be enabled on the job for them to appear.
aws glue get-jobs \ --query 'Jobs[?MaxCapacity > `10`].[Name,GlueVersion,MaxCapacity,WorkerType]' --output table
aws cloudwatch get-metric-statistics --namespace Glue \ --metric-name glue.driver.ExecutorAllocationManager.executors.numberMaxNeededExecutors \ --dimensions Name=JobName,Value=my-job Name=JobRunId,Value=ALL Name=Type,Value=gauge \ --statistics Maximum --period 86400 \ --start-time 2026-08-26T00:00:00Z --end-time 2026-09-25T00:00:00ZEvidence the rule requires
- The job’s configured DPUs are known, from
MaxCapacityor from the worker count times the DPUs per worker. - The job is sized by
MaxCapacity, which is the setting ZopNight can recommend a new value for. Jobs sized by worker type and worker count are not resized by this rule. - ZopNight holds at least 7 days of true hourly peaks for
numberMaxNeededExecutorsinside the 30-day window, so a quiet week cannot pass for a small job. - A priced saving is available and is at least $5 a month.
Jobs the rule leaves alone
A job with no DPU figure, a worker-type job, a job without trusted peak data, and a job whose peak need matches its allocation all produce nothing. There is no fallback to a flat percentage of the bill: if any input is missing, no recommendation appears.
How the DPU saving is calculated
peak needed DPU = configured DPU x min(1, peak needed executors / allocated executors)saving = (configured DPU - peak needed DPU) / configured DPU x measured DPU-hours x DPU-hour rateThe saving is capped at the job’s priced monthly cost. The recommended new size is the larger of 10 DPUs and the peak needed DPU rounded up, so a job whose peak really was 32 DPUs is sized to 32, never below what it used.
Shrinking a Glue job’s DPU allocation
- Check run history for the job’s duration and any SLA it has to meet.
- Lower
MaxCapacityto the recommended value in the job details, or withaws glue update-job. The update replaces the job definition, so pass the full current definition with only the capacity changed. - Compare the next few run times and
numberAllExecutorsandnumberMaxNeededExecutorsagainst the old runs. - If you move the job to Glue 3.0 or later with G or R worker types, Auto Scaling can pick the worker count per stage up to a maximum you set. Standard DPUs are not supported for Auto Scaling.