Every AI Assistant Needs Its Own Grant and Its Own Audit Trail
A pasted access token is identical in every assistant that holds it. OAuth consent gives each one its own revocable grant, scoped to the orgs you pick.
Bableen works on the Kubernetes side of Zop.Dev, focused on cluster ops, autoscaling, and the long tail of pod-level reliability work. She writes about MTTR, OOMKill diagnosis, and what runbooks actually need to do.
A pasted access token is identical in every assistant that holds it. OAuth consent gives each one its own revocable grant, scoped to the orgs you pick.
Apply, dismiss and close all resolve a finding. Bookmarking is the state that says not yet, and it only works if the filter is resolved server-side.
Cost history must outlive the resource that created it. That principle sounds obvious until you watch a deleted API key take three months of spend attribution with it into the void.
Policy as Code works cleanly until it meets ten teams, and then it breaks in ways the pilot never predicted. The first three months feel like a governance win. Policies deploy, violations get caught,…
Point-in-time benchmarks produce misleading Kubernetes cost comparisons because cluster behavior, workload patterns, and pricing all shift across a 12-month horizon in ways a single snapshot cannot…
On-call engineers are routinely woken at 3am to execute the same five-step runbook they ran the night before, and the tooling to stop this pattern has existed in production environments for years.…
At 500 managed resources, infrastructure drift stops being a maintenance nuisance and becomes a misdiagnosis engine that corrupts incident response at the root.
A single production incident costs more than most engineering teams budget for the entire on-call program that prevents it. The $50,000 figure (ZopDev) is not a worst-case projection. It is a floor,…
At 500 Terraform resources, the bottleneck is never Terraform. It is the organization running it.
Visibility without workflow integration is a cost center, not a cost cure. Most engineering organizations have invested in dashboards, tagging policies, and cost explorer tools. The spend keeps…
Alert routing is a notification system, not a resolution system, and treating it as the end state of incident response is why on-call engineers burn out.
Egress charges are structurally invisible in most cloud billing workflows, and that invisibility is expensive. The clearest proof: $28,000 per month in egress costs were approved without explicit…
Alert-only monitoring does not reduce operational cost. It defers labor and compounds it.
Static runbooks fail modern infrastructure because the gap between when a failure begins and when a human reads an alert is measured in minutes your SLO does not have.
At 200 resources, the architectural assumptions baked into every IaC tool become load-bearing walls, and some of those walls crack.
Prompt engineering entered DevOps not as an experiment but as a pressure valve: teams shipping faster than their tooling could support needed a way to extract precise, repeatable outputs from AI…
The pipeline itself is not the risk. The risk is the gap between what the pipeline assumes is true and what is actually true in the environment it deploys into.
Most AIOps deployments stall because they stop at observation. The team gets a dashboard. Alerts fire. Engineers stare at graphs. Nothing closes.
The on-call model fails at the architectural level, not the execution level. Paging a human, waiting for acknowledgment, and then diagnosing a live incident introduces latency that compounds into…
Most IDPs ship as friction-reducers and land as a new category of sprint tax. The promise is a self-service portal that abstracts infrastructure complexity. The reality, in production, is a platform…
Every FinOps initiative follows the same arc: a burst of recoverable savings in the first weeks, then a structural decay that accelerates past month 3 (ZopDev, "Why FinOps Savings Decay Faster After…
IaC tools built for single-team deployments fail structurally at 200 accounts because the failure modes are architectural, not configurational.
It simply becomes the baseline because nobody stopped to question it. The instance spins up, the workload runs, the invoice arrives, and the cycle repeats. By the time a team audits its compute…
Cost allocation depends on tags, but a flat key=value dropdown is unreadable at scale. Grouping values by key turns the tag picker into a real FinOps control.
Most tools that promise to manage your existing workloads ask for one thing first: redeploy everything through us. That is the wrong price. A re-rollout against production carries a real downtime…
Your cloud bill shows a Kubernetes cluster as a single number. EKS, GKE, and AKS all roll node compute into one rolled-up figure. Finance sees the total. Nobody sees which namespace or deployment…
CPU throttling has a visibility problem that the Kubernetes community partially fixed. exposes the throttle. Grafana dashboards flag it. Engineers know to look for it.
The Fargate Tax: Why Serverless Kubernetes Costs 38% More Past 200 vCPU-Hours Fargate is appealing because the pitch is clean: no AMI patching, no node group sizing, no cluster autoscaler tuning. You…
A team ships a generative-AI summarisation feature. The first month it costs $400 in Bedrock invocations. The second month it costs $1,200 as adoption grows. The third month it costs $9,200 because…
A cloud bill increased 22% from March to April. Without context the number looks alarming. With context — traffic grew 35% in the same window, MAU grew 28%, completed transactions grew 41% — the…
The observability bill at a 50-engineer org goes from $8,000/month in year one to $90,000/month by year three. The growth never gets a budget review because each individual instrumentation change…
A right-sized EKS cluster should not run at 40 percent node utilization. The pods declare requests that sum to 78 percent of node capacity. The cluster autoscaler provisions nodes to fit those…
Three years of cost retrospectives across mixed AWS fleets keep landing on the same finding. Teams that pick one compute commitment model and apply it across the whole fleet (all-Savings-Plan,…
Backstage has a $0 license fee and requires 2-3 senior engineers full-time. This piece works through the build vs buy decision with real TCO numbers and a decision framework.
OPA Gatekeeper rejects a pod before it ever runs. Here is how to write admission policies that block oversized resource requests, missing cost labels, and non-prod images at deploy time, not billing time.
Every DR design decision has a precise dollar figure. Active-active vs active-passive, cross-region replication cadence, failover automation. Here is the full cost breakdown.
70-80% of S3 objects are never accessed after upload yet sit in Standard at $0.023/GB. Here's the cost math, when Intelligent-Tiering breaks even, and how to automate lifecycle policies with guardrails.
Most cloud governance platforms have a blind spot. They enforce access controls on your infrastructure but leave their own control plane wide open. Here's how to fix it with graduated RBAC.
Terraform, OpenTofu, or Pulumi, the tool matters less than how you use it. Here's what actually works in production: state management, module design, testing layers, secrets hygiene, and drift detection in 2026.
Most Kubernetes clusters run at 35% CPU utilization, and nobody revisits node sizing after launch. Here's the measurement process, blue-green rotation pattern, and 72-hour soak checklist to right-size safely without production incidents.
Cloud spend is heading to $1 trillion in 2026, but 27% of it, $182 billion, is wasted. Here's what the latest data from Flexera, Gartner, and FinOps Foundation says about where the money goes and what actually works to recover it.
GKE Autopilot isn't always cheaper than Standard. Learn exactly when each mode wins, how Autopilot's minimum pod billing surprises teams, and which workload profiles justify the switch.
One post a week. Sundays. No "10 ways to think about cloud" listicles, just the engineering and FinOps notes we'd want to read.
See. Find. Fix. Automatic.
Connect your first cloud account in under 5 minutes. See your first remediation in under 7. No credit card required.