Skip to main content
Your progress
0 of 3 lessons complete0%
T6 / M6.5 / L2 OF 3 / Engineer TIER / 9 min

Choosing the right write surface

Outcome

By the end of this lesson, you will be able to choose the right write surface for a given action, execute the “agent helps; human writes” loop, and recognise when MCP is the wrong surface for a write even in an org where it is technically permitted.


TierEngineer
JTBD”When the agent identifies an action I want to take, pick the right write surface and execute cleanly.”
PersonasPlatform Engineer · SRE · FinOps Lead
PrerequisitesM6.5.L1 (the tiered write surface) · M6.2 (AI tool setup)
Time9 minutes
Bloom verbChoose (Evaluate), Execute (Apply), Avoid (Apply)

1. Concept

Most organisations leave writing switched off entirely, so an assistant reads and does nothing else.

Even where stopping a resource is permitted, an assistant is often still the wrong place to do it.

So when it tells you to stop i-0xyz123, where should that actually happen? There are five places it could, and this lesson is how to choose between them.

Terminal window
THE FIVE WRITE SURFACES:
1. ZopNight UI : humans, one-off actions
2. ZopNight API (PAT) : programmatic, automation
3. Auto-remediation rules : certified, recurring
4. Terraform / IaC : provisioning side
5. Cloud provider direct : emergency / bypass

Each has a sweet spot. Picking the right one is the engineering skill.

Surface 1: ZopNight UI

Terminal window
USE CASE: one-off action, human-in-the-loop, low-to-medium volume
FLOW:
1. Agent identifies action via MCP
2. Agent returns ZopNight UI URL (deep link to the resource/rec)
3. Engineer opens the URL
4. UI shows context (cost, tags, owner, history)
5. Engineer clicks "Apply" or "Schedule"
6. Audit log records: user@email clicked at time T
WHEN TO USE:
✓ 1-5 actions
✓ Need visual confirmation
✓ Want audit clearly attributed to human
✓ Action requires judgment (not a clean rule)

The UI is the highest-confidence write surface. The 30 seconds of “let me look” is worth it.

Surface 2: ZopNight API

Terminal window
USE CASE: scripted bulk action, automation, integration
FLOW:
1. Engineer reviews agent's suggestion + list of targets
2. Writes a script using ZopNight API
3. Runs script; PAT authenticates
4. Audit log records: PAT_X (description: "scripted bulk apply")
applied N actions at time T
EXAMPLE (bulk apply recommendations):
curl -X POST https://api.zop.dev/v1/recommendations/apply \
-H "Authorization: Bearer $ZN_PAT" \
-H "Content-Type: application/json" \
-d '{
"recommendation_ids": ["rec_1", "rec_2", "rec_3", ...],
"actor_note": "Quarterly cleanup; reviewed by jane@platform"
}'
WHEN TO USE:
✓ 10+ actions in one batch
✓ Automating a workflow
✓ Integrating with internal tools (CI, ticketing)
✓ Repeatable script (not one-shot)

API + script is the right pattern for bulk operations. The script is the artifact; the agent’s suggestions become the input.

Surface 3: Auto-remediation

Terminal window
USE CASE: recurring action covered by a certified rule
PATTERN:
Customer defines a rule: "delete unattached EBS volumes >30 days old"
ZopNight runs it on schedule (daily/weekly)
Each application is fully audited
No human-in-the-loop for known-safe action types
WHEN TO USE:
✓ Action repeats (weekly+)
✓ Risk is low (safeguards built in)
✓ Customer has approved the rule type
✓ Volume too high for one-at-a-time review
RULE TYPES (certified, common):
Delete unattached EBS volumes >30 days
Delete unused EBS snapshots >90 days
Stop idle non-prod instances overnight
Delete idle Lambda versions
Release unattached EIPs

Auto-rem is “agent suggests the policy; ZopNight executes the policy.” See M5.3.L4 for full coverage.

Surface 4: Terraform / IaC

Terminal window
USE CASE: infrastructure change (right-size, schedule, lifecycle)
that should be persistent in your IaC source of truth
PATTERN:
Agent identifies right-sizing opportunity
Engineer updates Terraform module (or asks agent to draft the diff)
PR with cost estimation (M5.6.L3) shows the dollar impact
Code review by team
Merge → apply
EXAMPLE diff (drafted by agent, reviewed by engineer):
resource "aws_instance" "web" {
- instance_type = "m5.xlarge"
+ instance_type = "m5.large" # cost savings: $80/mo (agent rec)
...
}
WHEN TO USE:
✓ Resource is managed by IaC
✓ Change should persist across re-provisions
✓ Want change-management review
✓ Cost delta deserves a PR-level decision

For IaC-managed infra, the IaC path is mandatory: otherwise the next terraform apply reverts the change.

Surface 5: Cloud provider direct (emergency only)

Terminal window
USE CASE: incident response when ZopNight unavailable
PATTERN:
Cost spike happens; ZopNight is down or slow
Engineer kills resources directly via aws/gcloud/az CLI
Audit log entry recorded manually in incident channel
Post-incident: reconcile with ZopNight after restoration
CAVEATS:
✗ Loses ZopNight's audit + safeguards
✗ Risk of inconsistent state vs ZopNight's view
✗ Manual audit reconciliation
WHEN TO USE:
✓ True incident (cost runaway, security)
✓ ZopNight unavailable
✓ Action too time-critical to wait
✓ Document immediately; reconcile after

This is the break-glass path. Use rarely; document thoroughly.

Decision matrix: pick the right surface

Terminal window
SITUATION RECOMMENDED PATH
─────────────────────────────────────────────────────────────────
One-off, low-stakes (single resource) ZopNight UI
One-off, high-stakes (production) ZopNight UI (with approval policy)
Bulk action (10-100 resources) ZopNight API or auto-rem
Bulk action (100+ resources) ZopNight API (script + review)
Recurring action (weekly+) Auto-remediation rule
Infrastructure right-size Terraform/IaC + PR
Schedule (recurring start/stop) ZopNight UI for schedule def
Emergency (active incident) Cloud direct, audit after
Mass tag application ZopNight API (bulk)
+ Terraform for new resources

Print this; tape to monitor.

Example: agent suggests right-sizing

Terminal window
AGENT (via MCP):
"i-0abc123 is 20% CPU-utilized over last 14 days;
recommend right-size m5.xlarge → m5.large
estimated savings: $80/mo (~$960/yr)"
ENGINEER's decision tree:
Is this resource in Terraform?
YES → Surface 4 (Terraform PR with cost-estimation)
NO → continue
Is this production or dev?
PROD → Surface 1 (UI with approval) or 4 (Terraform)
DEV → Surface 1 (UI, fast) or 3 (auto-rem if pattern is common)
Frequent right-sizing in this team?
YES → Surface 3 (define auto-rem rule)
NO → Surface 1 or 4
DECISION (this case): Dev environment, IaC-managed, infrequent →
Terraform PR. Agent drafts the diff; engineer reviews + merges.
Total time: 15 minutes including review.

The “right path” depends on context. The decision matrix gives the starting point.

Why not just expose writes via MCP

Recap from L1, in case anyone forgets:

Terminal window
1. Agents hallucinate (LLM picks wrong target or wrong action)
2. Confirmation fatigue if every call needs approval
3. Audit trail gets ambiguous ("did the agent intend this?")
4. Existing surfaces cover the write needs
5. One bad agent action in production = trust erosion org-wide
ZOPNIGHT'S BET:
Better UX through dedicated write surfaces
Than a single MCP "do anything" endpoint
The agent + 5 surfaces beat agent-with-writes.

The “agent helps; human writes” loop

The full operating model:

Terminal window
1. Engineer asks agent a question
"what's the biggest cost-saving opportunity in payment-team?"
2. Agent investigates via MCP (reads)
list_resources, get_costs, get_recommendations
3. Agent recommends an action with justification
"Top opportunity: stop rds-prod-staging-replica.
Idle 75 days. Savings: $240/mo.
Confirmed: no recent connections per CloudWatch."
4. Engineer reviews the recommendation
"Is this in IaC? Who owns this DB?"
(often the engineer asks the agent these follow-ups via MCP)
5. Engineer executes via the right write surface
Surface 1 (UI) for a one-off non-IaC resource
Surface 4 (Terraform PR) for an IaC-managed resource
6. Audit log records the human's action
user@email clicked Stop at time T
Action: stop rds-prod-staging-replica
LOOP CONTINUES:
Engineer asks agent: "did the action take effect? cost recovered?"
Agent checks via MCP read (audit log + current cost)
Confirms or reports issue

Agent is research + drafting. Human is decision + action. The split is the safety architecture.

Common mistake: fighting the contract

Terminal window
WRONG PATTERN:
Engineer: "stop i-0xyz123"
Agent: "I don't have a tool for that"
Engineer: [annoyed, types "you're useless"]
Agent: [apologetic, still cannot stop]
(And the variant that is worse, at tier 2:
Agent: [stops it]
Engineer: "wait, which one did you stop?")
THE PRODUCTIVE PATTERN:
Engineer: "Should I stop i-0xyz123? It's been idle 30 days."
Agent: investigates, confirms idle, checks owner,
drafts justification with cost + risk assessment
Engineer: reviews the justification
opens ZopNight UI from agent's link
clicks Stop with one click of confidence
REFRAME the request from imperative ("stop X") to investigative
("should I stop X?"). The agent + UI together is FASTER than UI alone
because the agent does the research the engineer would otherwise do
manually.

The mistake is treating MCP as a command surface. It’s an analysis + recommendation surface that pairs with write surfaces.

Tool integration: combining surfaces

Terminal window
SOME AI TOOLS support cross-surface workflows:
Cursor: can edit Terraform files
Claude Code: can run CLIs + edit code + open URLs
Copilot Workspace: can open GitHub PRs
Claude Desktop: limited to MCP and chat
COMBO PATTERN (best leverage):
1. Agent reads ZopNight data via MCP
2. Agent drafts the Terraform change (via Cursor/Claude Code)
3. Tool creates a Git branch + opens a PR (via Copilot Workspace/CLI)
4. Human reviews PR (cost estimate visible)
5. Human merges
6. Apply runs in CI
7. Engineer asks agent (MCP) to confirm savings recovered
EFFECTIVE WORKFLOW at tier none, and still the right shape
above it. The agent + IDE + ZopNight UI + CI = the modern
FinOps toolchain.

This is the future: agents that orchestrate writes through deliberate, audited surfaces.


2. Demo

A quarterly cleanup workflow combining surfaces:

Terminal window
ENGINEER WANTS TO CLEAN UP IDLE RESOURCES (quarterly task):
STEP 1: Agent (via MCP): research
> "list top 20 idle resources across non-prod accounts,
sort by monthly cost descending"
→ Returns list: 20 resources, total $2,400/mo savings opportunity
STEP 2: Engineer reviews (with agent's help)
> "for each, who's the owner? when was it last used?"
→ Agent enriches list with owner emails + last-used dates
STEP 3: Engineer decides
Decides 14 are safe to stop now
4 need owner sign-off (Slack to owners)
2 are actually used (rare workflow); skip
STEP 4: Execute via right surfaces
6 IaC-managed → ask agent to draft Terraform diff;
open as PR with cost estimate
8 manual resources → bulk apply via API:
curl -X POST .../v1/resources/bulk-stop -d '{...}'
STEP 5: Agent (via MCP): verify
> "compare yesterday's run rate to today's for non-prod"
→ Returns: $2,200/mo savings confirmed (98% of estimate)
TOTAL TIME: 45 minutes for $2,200/mo recovery = $26,400/yr.
The agent did 30 minutes of research the engineer would have
done manually; engineer did 15 minutes of decisions + executions.

This is the operating model in production. Each surface plays its role.


3. Hands-on (5 min)

Plan 3 cost-saving actions you want to take this week:

Terminal window
ACTION 1: __________
Recommended surface: __________
Why: __________
ACTION 2: __________
Recommended surface: __________
Why: __________
ACTION 3: __________
Recommended surface: __________
Why: __________
REFLECTION:
Which actions are good candidates for auto-rem rules
(i.e., they'll repeat)?
→ __________
Which need a Terraform PR (IaC-managed)?
→ __________

Mapping actions to surfaces is a 5-minute exercise; saves hours of friction over the quarter.

Do it through MCP. The same task you just did in the console, asked in one sentence.

Terminal window
BEFORE A ZopNight account with one cloud connected. A non-production account, and two instances you are free to stop.
ASK "Check what this would change, and if it is safe, stop the two idle instances."
CHECK which actions needed your approval and which did not. That boundary is what this module is about, and the server enforces it, not the assistant.

Tools behind it: preview_remediation (read, Optimize), stop_resource (write, tier 3, irreversible), bulk_stop_resources (write, tier 3, irreversible), create_schedule (write, tier 2, reversible). The full catalogue is at zop.dev/learn/mcp-tools.


4. Knowledge check

Q1

Agent suggests stopping one resource. Best write path:

A. ZopNight UI (one-off, low cognitive overhead, full audit, and you see exactly which resource you are acting on)
B. Via MCP, since the agent already holds the resource id and all the context needed to act on it correctly
C. AWS CLI directly, which is the fastest path and at least leaves the change in your own shell history
D. A scheduled auto-remediation rule, so that the same stop happens again next time without anyone asking

Show answer

Correct: A. At tier none MCP cannot do it; at tier 2 it can, and it is still the wrong choice for a one-off because the UI shows you the target. API is the alternative for scripted actions. UI for one-off. The 30 seconds of review is the safeguard.

Q2

Recurring write action (delete unattached EBS weekly):

A. By hand, one at a time
B. A one-off UI action
C. Auto-remediation rule
D. Skip it entirely

Show answer

Correct: C. No human-in-loop for the certified rule Define once; ZopNight applies on schedule with audit. See M5.3.L4 for the full pattern. Auto-rem rule. Manual repetition is the anti-pattern.

Q3

Persistent infrastructure right-sizing:

A. UI, for a one-off change
B. Terraform/IaC + PR
C. The cloud CLI, by hand
D. A ticket raised by hand first

Show answer

Correct: B. Audit + reproducibility + cost estimate pre-merge (M5.6.L3). Without IaC path, the next terraform apply reverts the manual change. IaC path for IaC-managed resources.


5. Apply

Match the write to the surface. Availability is not the same as suitability: even where the tier permits a write, pick the surface that shows a human what is about to change. Reframe imperative prompts (“stop X”) into investigative ones (“should I stop X?”).

For your team: post the decision matrix in the team wiki. Reference it the first 2-3 times until it becomes muscle memory.


Glossary terms touched

Write surface · Agent-helps-human-writes loop · Decision matrix · Break-glass path


Start with the bill.

Foundations takes about five hours. The first lesson is nine minutes.

Open curriculum. No login. No paywall. 290 lessons across 7 courses, three publicly verifiable credentials. Read it on the train, take the exam on a Saturday, list the credential on your résumé Monday.

5h median time to finish Foundations
0 logins, paywalls, or marketing forms
open curriculum, public credential verifier
Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console· Multi-cloud automation· Production-ready in 30 min· SOC 2 · ISO 27001· 20–60% off the bill, first month· 4 platforms · 1 console·