Outcome
By the end of this lesson, you will be able to choose the right write surface for a given action, execute the “agent helps; human writes” loop, and recognise when MCP is the wrong surface for a write even in an org where it is technically permitted.
| Tier | Engineer |
| JTBD | ”When the agent identifies an action I want to take, pick the right write surface and execute cleanly.” |
| Personas | Platform Engineer · SRE · FinOps Lead |
| Prerequisites | M6.5.L1 (the tiered write surface) · M6.2 (AI tool setup) |
| Time | 9 minutes |
| Bloom verb | Choose (Evaluate), Execute (Apply), Avoid (Apply) |
1. Concept
Most organisations leave writing switched off entirely, so an assistant reads and does nothing else.
Even where stopping a resource is permitted, an assistant is often still the wrong place to do it.
So when it tells you to stop i-0xyz123, where should that actually happen? There are five places it could, and this lesson is how to choose between them.
THE FIVE WRITE SURFACES: 1. ZopNight UI : humans, one-off actions 2. ZopNight API (PAT) : programmatic, automation 3. Auto-remediation rules : certified, recurring 4. Terraform / IaC : provisioning side 5. Cloud provider direct : emergency / bypassEach has a sweet spot. Picking the right one is the engineering skill.
Surface 1: ZopNight UI
USE CASE: one-off action, human-in-the-loop, low-to-medium volume
FLOW: 1. Agent identifies action via MCP 2. Agent returns ZopNight UI URL (deep link to the resource/rec) 3. Engineer opens the URL 4. UI shows context (cost, tags, owner, history) 5. Engineer clicks "Apply" or "Schedule" 6. Audit log records: user@email clicked at time T
WHEN TO USE: ✓ 1-5 actions ✓ Need visual confirmation ✓ Want audit clearly attributed to human ✓ Action requires judgment (not a clean rule)The UI is the highest-confidence write surface. The 30 seconds of “let me look” is worth it.
Surface 2: ZopNight API
USE CASE: scripted bulk action, automation, integration
FLOW: 1. Engineer reviews agent's suggestion + list of targets 2. Writes a script using ZopNight API 3. Runs script; PAT authenticates 4. Audit log records: PAT_X (description: "scripted bulk apply") applied N actions at time T
EXAMPLE (bulk apply recommendations): curl -X POST https://api.zop.dev/v1/recommendations/apply \ -H "Authorization: Bearer $ZN_PAT" \ -H "Content-Type: application/json" \ -d '{ "recommendation_ids": ["rec_1", "rec_2", "rec_3", ...], "actor_note": "Quarterly cleanup; reviewed by jane@platform" }'
WHEN TO USE: ✓ 10+ actions in one batch ✓ Automating a workflow ✓ Integrating with internal tools (CI, ticketing) ✓ Repeatable script (not one-shot)API + script is the right pattern for bulk operations. The script is the artifact; the agent’s suggestions become the input.
Surface 3: Auto-remediation
USE CASE: recurring action covered by a certified rule
PATTERN: Customer defines a rule: "delete unattached EBS volumes >30 days old" ZopNight runs it on schedule (daily/weekly) Each application is fully audited No human-in-the-loop for known-safe action types
WHEN TO USE: ✓ Action repeats (weekly+) ✓ Risk is low (safeguards built in) ✓ Customer has approved the rule type ✓ Volume too high for one-at-a-time review
RULE TYPES (certified, common): Delete unattached EBS volumes >30 days Delete unused EBS snapshots >90 days Stop idle non-prod instances overnight Delete idle Lambda versions Release unattached EIPsAuto-rem is “agent suggests the policy; ZopNight executes the policy.” See M5.3.L4 for full coverage.
Surface 4: Terraform / IaC
USE CASE: infrastructure change (right-size, schedule, lifecycle) that should be persistent in your IaC source of truth
PATTERN: Agent identifies right-sizing opportunity Engineer updates Terraform module (or asks agent to draft the diff) PR with cost estimation (M5.6.L3) shows the dollar impact Code review by team Merge → apply
EXAMPLE diff (drafted by agent, reviewed by engineer): resource "aws_instance" "web" { - instance_type = "m5.xlarge" + instance_type = "m5.large" # cost savings: $80/mo (agent rec) ... }
WHEN TO USE: ✓ Resource is managed by IaC ✓ Change should persist across re-provisions ✓ Want change-management review ✓ Cost delta deserves a PR-level decisionFor IaC-managed infra, the IaC path is mandatory: otherwise the next terraform apply reverts the change.
Surface 5: Cloud provider direct (emergency only)
USE CASE: incident response when ZopNight unavailable
PATTERN: Cost spike happens; ZopNight is down or slow Engineer kills resources directly via aws/gcloud/az CLI Audit log entry recorded manually in incident channel Post-incident: reconcile with ZopNight after restoration
CAVEATS: ✗ Loses ZopNight's audit + safeguards ✗ Risk of inconsistent state vs ZopNight's view ✗ Manual audit reconciliation
WHEN TO USE: ✓ True incident (cost runaway, security) ✓ ZopNight unavailable ✓ Action too time-critical to wait ✓ Document immediately; reconcile afterThis is the break-glass path. Use rarely; document thoroughly.
Decision matrix: pick the right surface
SITUATION RECOMMENDED PATH─────────────────────────────────────────────────────────────────One-off, low-stakes (single resource) ZopNight UIOne-off, high-stakes (production) ZopNight UI (with approval policy)Bulk action (10-100 resources) ZopNight API or auto-remBulk action (100+ resources) ZopNight API (script + review)Recurring action (weekly+) Auto-remediation ruleInfrastructure right-size Terraform/IaC + PRSchedule (recurring start/stop) ZopNight UI for schedule defEmergency (active incident) Cloud direct, audit afterMass tag application ZopNight API (bulk) + Terraform for new resourcesPrint this; tape to monitor.
Example: agent suggests right-sizing
AGENT (via MCP): "i-0abc123 is 20% CPU-utilized over last 14 days; recommend right-size m5.xlarge → m5.large estimated savings: $80/mo (~$960/yr)"
ENGINEER's decision tree:
Is this resource in Terraform? YES → Surface 4 (Terraform PR with cost-estimation) NO → continue
Is this production or dev? PROD → Surface 1 (UI with approval) or 4 (Terraform) DEV → Surface 1 (UI, fast) or 3 (auto-rem if pattern is common)
Frequent right-sizing in this team? YES → Surface 3 (define auto-rem rule) NO → Surface 1 or 4
DECISION (this case): Dev environment, IaC-managed, infrequent → Terraform PR. Agent drafts the diff; engineer reviews + merges. Total time: 15 minutes including review.The “right path” depends on context. The decision matrix gives the starting point.
Why not just expose writes via MCP
Recap from L1, in case anyone forgets:
1. Agents hallucinate (LLM picks wrong target or wrong action)2. Confirmation fatigue if every call needs approval3. Audit trail gets ambiguous ("did the agent intend this?")4. Existing surfaces cover the write needs5. One bad agent action in production = trust erosion org-wide
ZOPNIGHT'S BET: Better UX through dedicated write surfaces Than a single MCP "do anything" endpoint
The agent + 5 surfaces beat agent-with-writes.The “agent helps; human writes” loop
The full operating model:
1. Engineer asks agent a question "what's the biggest cost-saving opportunity in payment-team?"
2. Agent investigates via MCP (reads) list_resources, get_costs, get_recommendations
3. Agent recommends an action with justification "Top opportunity: stop rds-prod-staging-replica. Idle 75 days. Savings: $240/mo. Confirmed: no recent connections per CloudWatch."
4. Engineer reviews the recommendation "Is this in IaC? Who owns this DB?" (often the engineer asks the agent these follow-ups via MCP)
5. Engineer executes via the right write surface Surface 1 (UI) for a one-off non-IaC resource Surface 4 (Terraform PR) for an IaC-managed resource
6. Audit log records the human's action user@email clicked Stop at time T Action: stop rds-prod-staging-replica
LOOP CONTINUES: Engineer asks agent: "did the action take effect? cost recovered?" Agent checks via MCP read (audit log + current cost) Confirms or reports issueAgent is research + drafting. Human is decision + action. The split is the safety architecture.
Common mistake: fighting the contract
WRONG PATTERN: Engineer: "stop i-0xyz123" Agent: "I don't have a tool for that" Engineer: [annoyed, types "you're useless"] Agent: [apologetic, still cannot stop]
(And the variant that is worse, at tier 2: Agent: [stops it] Engineer: "wait, which one did you stop?")
THE PRODUCTIVE PATTERN: Engineer: "Should I stop i-0xyz123? It's been idle 30 days." Agent: investigates, confirms idle, checks owner, drafts justification with cost + risk assessment Engineer: reviews the justification opens ZopNight UI from agent's link clicks Stop with one click of confidence
REFRAME the request from imperative ("stop X") to investigative("should I stop X?"). The agent + UI together is FASTER than UI alonebecause the agent does the research the engineer would otherwise domanually.The mistake is treating MCP as a command surface. It’s an analysis + recommendation surface that pairs with write surfaces.
Tool integration: combining surfaces
SOME AI TOOLS support cross-surface workflows: Cursor: can edit Terraform files Claude Code: can run CLIs + edit code + open URLs Copilot Workspace: can open GitHub PRs Claude Desktop: limited to MCP and chat
COMBO PATTERN (best leverage): 1. Agent reads ZopNight data via MCP 2. Agent drafts the Terraform change (via Cursor/Claude Code) 3. Tool creates a Git branch + opens a PR (via Copilot Workspace/CLI) 4. Human reviews PR (cost estimate visible) 5. Human merges 6. Apply runs in CI 7. Engineer asks agent (MCP) to confirm savings recovered
EFFECTIVE WORKFLOW at tier none, and still the right shapeabove it. The agent + IDE + ZopNight UI + CI = the modernFinOps toolchain.This is the future: agents that orchestrate writes through deliberate, audited surfaces.
2. Demo
A quarterly cleanup workflow combining surfaces:
ENGINEER WANTS TO CLEAN UP IDLE RESOURCES (quarterly task):
STEP 1: Agent (via MCP): research > "list top 20 idle resources across non-prod accounts, sort by monthly cost descending" → Returns list: 20 resources, total $2,400/mo savings opportunity
STEP 2: Engineer reviews (with agent's help) > "for each, who's the owner? when was it last used?" → Agent enriches list with owner emails + last-used dates
STEP 3: Engineer decides Decides 14 are safe to stop now 4 need owner sign-off (Slack to owners) 2 are actually used (rare workflow); skip
STEP 4: Execute via right surfaces 6 IaC-managed → ask agent to draft Terraform diff; open as PR with cost estimate 8 manual resources → bulk apply via API: curl -X POST .../v1/resources/bulk-stop -d '{...}'
STEP 5: Agent (via MCP): verify > "compare yesterday's run rate to today's for non-prod" → Returns: $2,200/mo savings confirmed (98% of estimate)
TOTAL TIME: 45 minutes for $2,200/mo recovery = $26,400/yr.The agent did 30 minutes of research the engineer would havedone manually; engineer did 15 minutes of decisions + executions.This is the operating model in production. Each surface plays its role.
3. Hands-on (5 min)
Plan 3 cost-saving actions you want to take this week:
ACTION 1: __________ Recommended surface: __________ Why: __________
ACTION 2: __________ Recommended surface: __________ Why: __________
ACTION 3: __________ Recommended surface: __________ Why: __________
REFLECTION: Which actions are good candidates for auto-rem rules (i.e., they'll repeat)? → __________
Which need a Terraform PR (IaC-managed)? → __________Mapping actions to surfaces is a 5-minute exercise; saves hours of friction over the quarter.
Do it through MCP. The same task you just did in the console, asked in one sentence.
BEFORE A ZopNight account with one cloud connected. A non-production account, and two instances you are free to stop.ASK "Check what this would change, and if it is safe, stop the two idle instances."CHECK which actions needed your approval and which did not. That boundary is what this module is about, and the server enforces it, not the assistant.Tools behind it: preview_remediation (read, Optimize), stop_resource (write, tier 3, irreversible), bulk_stop_resources (write, tier 3, irreversible), create_schedule (write, tier 2, reversible). The full catalogue is at zop.dev/learn/mcp-tools.
4. Knowledge check
Q1
Agent suggests stopping one resource. Best write path:
A. ZopNight UI (one-off, low cognitive overhead, full audit, and you see exactly which resource you are acting on)
B. Via MCP, since the agent already holds the resource id and all the context needed to act on it correctly
C. AWS CLI directly, which is the fastest path and at least leaves the change in your own shell history
D. A scheduled auto-remediation rule, so that the same stop happens again next time without anyone asking
Show answer
Correct: A. At tier none MCP cannot do it; at tier 2 it can, and it is still the wrong choice for a one-off because the UI shows you the target. API is the alternative for scripted actions. UI for one-off. The 30 seconds of review is the safeguard.
Q2
Recurring write action (delete unattached EBS weekly):
A. By hand, one at a time
B. A one-off UI action
C. Auto-remediation rule
D. Skip it entirely
Show answer
Correct: C. No human-in-loop for the certified rule Define once; ZopNight applies on schedule with audit. See M5.3.L4 for the full pattern. Auto-rem rule. Manual repetition is the anti-pattern.
Q3
Persistent infrastructure right-sizing:
A. UI, for a one-off change
B. Terraform/IaC + PR
C. The cloud CLI, by hand
D. A ticket raised by hand first
Show answer
Correct: B. Audit + reproducibility + cost estimate pre-merge (M5.6.L3). Without IaC path, the next terraform apply reverts the manual change. IaC path for IaC-managed resources.
5. Apply
Match the write to the surface. Availability is not the same as suitability: even where the tier permits a write, pick the surface that shows a human what is about to change. Reframe imperative prompts (“stop X”) into investigative ones (“should I stop X?”).
For your team: post the decision matrix in the team wiki. Reference it the first 2-3 times until it becomes muscle memory.
Related lessons
- L1: The tiered write surface
- L3: Future: write approval roadmap (next)
- M5.3.L4: Auto-remediation rules
- M5.6.L3: Pre-merge cost estimation
Glossary terms touched
Write surface · Agent-helps-human-writes loop · Decision matrix · Break-glass path