Skip to main content
Your progress
0 of 4 lessons complete0%
T0 / M0.3 / L4 OF 4 / Operator TIER / 10 min

When scheduling wins, when commitments win

Outcome

By the end of this lesson, you will be able to decide which lever (schedule, commit, or both) fits any given workload using a four-question decision tree.


TierOperator
JTBD”Stop debating and pick the right lever per workload, every time.”
PersonasAll five
PrerequisitesL1, L2, L3
Time10 minutes
Bloom verbDecide (Evaluate)

1. Concept

Every workload sits in one of five categories defined by two questions: Does it need to run 24/7? and How predictable is its capacity demand? The answers map cleanly to a lever choice.

The four-question decision tree

Terminal window
Q1: Does this workload need to be available 24/7?
YES → go to Q2
NO → SCHEDULE. Stop here.
Q2: Is the steady-state floor predictable for 1+ year?
YES → go to Q3
NO → ON-DEMAND. Skip commitments. Stay flexible.
Q3: Is the workload stateless and eviction-tolerant for SOME portion of its capacity?
YES → COMMIT FLOOR + SPOT BURST. Hybrid approach.
NO → COMMIT FLOOR + ON-DEMAND PEAK.
Q4: Is there an event-driven burst (specific dates, traffic events)?
YES → add EVENT READINESS pre-scaling on top of the base config.
NO → done.

The tree is short. Most workloads exit at Q1 (schedulable) or Q3 (commit-the-floor pattern). The trap is asking these questions in the wrong order: most commitment debates start at Q3 without ever testing Q1.

Workload archetype × right lever

Terminal window
ARCHETYPE RIGHT LEVER(S)
──────────────────────────────────────────────────────────────────
Dev / staging / test (non-prod) Schedule (default), no commitment
Production web app (24/7 steady) Commit floor, on-demand peak
Production with bursty traffic Commit floor, on-demand peak, spot for stateless work above
K8s clusters with stateless work Schedule non-prod, commit prod floor, spot for stateless pods
Batch / analytics / ML training Spot (default), commit only if predictable
Reporting DBs (business hours) Schedule (treat as non-prod for cost purposes)
Disaster recovery / cold standby On-demand, scale-to-zero between drills
Event-driven (marketing launches) Event Readiness pre-scaling on top of base

When commitments DO win on non-prod

The non-prod fallacy from L3 is the default. But there is a legitimate exception:

A non-prod environment that has been scheduled, observed, and has a stable post-schedule floor that nobody disputes is going away.

Example: a long-running performance test infrastructure that runs every weekday for two years. Scheduled off nights and weekends, but the weekday floor is stable. Buying a 1-yr commitment on the post-schedule floor is fine.

The exception applies when steps 1–3 of the decision tree have already been done. Before any of those steps, the default for non-prod is “schedule, do not commit.”

When scheduling does NOT apply

Production is the most common exception. Production usually needs to be available continuously. Scheduling production is rare and intentional; typically only for explicitly non-24/7 products (a B2B SaaS that publishes a 9-6 SLA and explicitly stops outside hours; some single-region developer tools that announce maintenance windows).

For most production, the right pattern is commit-the-floor + on-demand-peak, with autoscaling managing the variance. Production scheduling is a deliberate choice, not a default.

The combined view

A real estate uses all the levers, applied to different workload classes:

Terminal window
WORKLOAD CLASS PRIMARY LEVER SECONDARY
─────────────────────────────────────────────────────────────────
Non-prod (dev / test / stage) Schedule n/a
Production floor (steady) 1-yr Savings Plan n/a
Production peak (predictable) On-demand Autoscaling
Production burst (event) Event Readiness Autoscaling
Stateless batch Spot Fall back to on-demand
Disaster recovery On-demand Scale-to-zero

This is the picture a mature FinOps practice produces. Not one lever, not all levers everywhere: the right lever per class.

The 60-day baselining rule

A useful default rule:

For any new workload at non-trivial scale, do not buy commitments in the first 60 days. Schedule what is scheduleable. Observe the post-schedule steady-state. Commit only after the floor is stable.

This rule averts almost every over-commitment incident. The 60-day window is long enough to absorb the launch curve and short enough not to delay legitimate savings.


2. Demo

A real workload audit, walked through the decision tree:

Terminal window
WORKLOAD: "data-pipeline-staging"
TYPE: GKE cluster, 12 nodes, used by data team
Q1: Does this need to be 24/7?
Engineering: "Engineers only use it during business hours."
A: NO
→ DECISION: SCHEDULE
Pattern: 8-8 Mon-Fri schedule, scale-to-zero off-hours
Expected savings: ~64% of current cost
─────────────────────────────────────────────────────────
WORKLOAD: "checkout-service-prod"
TYPE: ECS service, 8 tasks at steady, 14 at peak afternoon
Q1: 24/7? YES
Q2: Steady-state floor predictable?
Engineering: "Yes, we've run 8 tasks at steady for 18 months."
A: YES → go to Q3
Q3: Some stateless / eviction-tolerant?
Engineering: "Tasks are stateless but eviction would cause user-visible errors."
A: NO
→ DECISION: COMMIT FLOOR + ON-DEMAND PEAK
Pattern: 1-yr Savings Plan covering 8 tasks, on-demand for peak
─────────────────────────────────────────────────────────
WORKLOAD: "data-warehouse-batch-jobs"
TYPE: EMR Serverless, ~80 vCPU-hours/day, stateless, retryable
Q1: 24/7? Sort of (batch fires multiple times daily)
Q2: Floor predictable? "Daily floor varies, but 40 vCPU-hours/day is the minimum."
A: YES → go to Q3
Q3: Stateless / eviction-tolerant?
A: YES (batch jobs checkpoint)
→ DECISION: COMMIT FLOOR + SPOT BURST
Pattern: Savings Plan for 40 vCPU-hours/day, Spot for the variable layer

Three workloads, three different lever combinations. The tree drives each.


3. Hands-on (7 min)

Run the decision tree on three of your own workloads:

Terminal window
WORKLOAD 1: ____________________
Q1 (24/7 needed?): Y / N
Q2 (floor predictable for 1yr?): Y / N / n/a
Q3 (eviction-tolerant burst?): Y / N / n/a
Q4 (event-driven burst?): Y / N / n/a
→ DECISION: ____________________
WORKLOAD 2: ____________________
...
WORKLOAD 3: ____________________
...

For each “schedule” decision, the next step is a real schedule. For each “commit” decision, the next step is the 60-day baselining rule before any purchase.


4. Knowledge check

Q1

A workload runs continuously, the floor is predictable, and the burst is stateless / retryable. The right lever combination is:

A. Schedule the workload off-hours
B. Buy a 1-yr Savings Plan for everything
C. Commit on the floor, use Spot for the stateless burst
D. Pure on-demand for flexibility

Show answer

Correct: C. This is the commit-floor + spot-burst pattern. Each layer gets the right instrument. (A is wrong because workload is 24/7; B is over-commitment because the burst is variable; D leaves savings on the table.)

Q2

A new workload was launched 30 days ago. The engineering team proposes a 1-yr commitment to lock in savings. The right response:

A. Agree
B. Decline: apply the 60-day baselining rule. Schedule first if applicable, observe steady-state, then revisit at day 60.
C. Buy a 3-yr commitment instead, for deeper savings
D. Wait for the next quarter

Show answer

Correct: B. The 60-day baselining rule averts most over-commitment. 30 days is too early to commit on a new workload.

Q3

Production is rarely scheduled because:

A. Production cannot be scheduled
B. Customers do not respect schedules. Production is available continuously by default. Exceptions exist for explicitly non-24/7 products but are deliberate, not the default.
C. Scheduling production is illegal
D. Production should be on Spot

Show answer

Correct: B. Production scheduling is deliberate when it happens (some single-region tools, scheduled-availability SaaS), not a default move.


5. Apply

ZopNight’s UI is the place this decision tree gets executed:

  • Q1 (schedule?)Schedules: build the schedule, attach resources or groups
  • Q3 (commit + burst design)Reports → Purchase Type for current mix; commitment purchase happens in provider console / via specialist tools
  • Q4 (event burst)Event Readiness: pre-scale for the event date

The decision tree is the question the team asks. The product is where they apply it.


Module quiz

You have now completed all four lessons of M0.3. The module quiz (10 questions, 80% pass) lives at /certifications/operator/m0.3-quiz. Pass to earn the Lever-Sequencer chip.


Glossary terms touched

Decision tree · 60-day baselining rule · Floor commit · Spot burst


Start with the bill.

Foundations takes about five hours. The first lesson is nine minutes.

Open curriculum. No login. No paywall. 237 lessons across 7 courses, three publicly verifiable credentials. Read it on the train, take the exam on a Saturday, list the credential on your résumé Monday.

5h median time to finish Foundations
0 logins, paywalls, or marketing forms
open curriculum, public credential verifier
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 30% average cloud cost cut· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 30% average cloud cost cut· 4 platforms · 1 console·