Skip to main content
rightsizing · aws

AWS Glue jobs allocated more DPUs than their peak demand for Spark executors needs

resource types
1
rule IDs covered
1
severity
low

What does ZopNight detect here?

ZopNight compares an AWS Glue job's configured DPUs with the peak `numberMaxNeededExecutors` it reports to CloudWatch over 30 days. When the job asks for fewer executors than it is given, ZopNight prices the unused DPUs at the job's DPU-hour rate and recommends a smaller MaxCapacity that never drops below 10 DPUs or measured peak need.

Signal and threshold

How ZopNight evaluates AWS Glue jobs allocated more DPUs than their peak demand for Spark executors needs.
Field Value
Rule IDsRC-098
Categoryrightsizing
Severitylow
Metricglue.driver.ExecutorAllocationManager.executors.numberMaxNeededExecutors
Thresholdpeak needed executors below allocated executors
Evaluation window30d
SourceZopNight
Permissions usedglue:GetJobs · glue:GetJobRuns · cloudwatch:GetMetricStatistics

Every allocated DPU is billed for the whole run

AWS Glue bills an hourly rate per Data Processing Unit, and one DPU is 4 vCPU and 16 GB of memory. Billing is per second, with a 1-minute minimum on Glue 2.0 and later. AWS’s own example: a Spark job using 6 DPUs for 15 minutes at $0.44 per DPU-hour costs $0.66. By default Glue gives a Spark job 10 DPUs.

The bill follows the allocation, not the work. If a job is given 40 DPUs and its stages never need more than a dozen executors, the other capacity idles on every run and is still paid for.

Reading peak Spark executors from CloudWatch

Glue publishes two gauges that answer the question directly. numberAllExecutors is the number of executors running, and numberMaxNeededExecutors is the number of running plus pending executors needed to satisfy the current load. Glue’s capacity-planning guide says to compare the two. Job metrics must be enabled on the job for them to appear.

Terminal window
aws glue get-jobs \
--query 'Jobs[?MaxCapacity > `10`].[Name,GlueVersion,MaxCapacity,WorkerType]' --output table
aws cloudwatch get-metric-statistics --namespace Glue \
--metric-name glue.driver.ExecutorAllocationManager.executors.numberMaxNeededExecutors \
--dimensions Name=JobName,Value=my-job Name=JobRunId,Value=ALL Name=Type,Value=gauge \
--statistics Maximum --period 86400 \
--start-time 2026-08-26T00:00:00Z --end-time 2026-09-25T00:00:00Z

Evidence the rule requires

  1. The job’s configured DPUs are known, from MaxCapacity or from the worker count times the DPUs per worker.
  2. The job is sized by MaxCapacity, which is the setting ZopNight can recommend a new value for. Jobs sized by worker type and worker count are not resized by this rule.
  3. ZopNight holds at least 7 days of true hourly peaks for numberMaxNeededExecutors inside the 30-day window, so a quiet week cannot pass for a small job.
  4. A priced saving is available and is at least $5 a month.

Jobs the rule leaves alone

A job with no DPU figure, a worker-type job, a job without trusted peak data, and a job whose peak need matches its allocation all produce nothing. There is no fallback to a flat percentage of the bill: if any input is missing, no recommendation appears.

How the DPU saving is calculated

Terminal window
peak needed DPU = configured DPU x min(1, peak needed executors / allocated executors)
saving = (configured DPU - peak needed DPU) / configured DPU x measured DPU-hours x DPU-hour rate

The saving is capped at the job’s priced monthly cost. The recommended new size is the larger of 10 DPUs and the peak needed DPU rounded up, so a job whose peak really was 32 DPUs is sized to 32, never below what it used.

Shrinking a Glue job’s DPU allocation

  1. Check run history for the job’s duration and any SLA it has to meet.
  2. Lower MaxCapacity to the recommended value in the job details, or with aws glue update-job. The update replaces the job definition, so pass the full current definition with only the capacity changed.
  3. Compare the next few run times and numberAllExecutors and numberMaxNeededExecutors against the old runs.
  4. If you move the job to Glue 3.0 or later with G or R worker types, Auto Scaling can pick the worker count per stage up to a maximum you set. Standard DPUs are not supported for Auto Scaling.

See it fire on your bill.

Connect an account read-only. The first findings land in minutes.

472 rule families across 353 resource types on 22 platforms. Every threshold, metric, and IAM action is documented on these pages before you grant anything.

472 rule families documented
353 resource types covered
read-only default access level
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·