23 de septiembre de 2026 · 5 min de lectura
Rehearse disasters before production does it for you
A hands-on GitOps runbook: minikube, Terraform, Argo CD app-of-apps, and three disaster drills to run before a real migration.
The problem with untested runbooks
Every team has a disaster recovery doc somewhere — a wiki page, a RUNBOOK.md nobody opens outside onboarding. It says what to do when a deployment breaks or scaling misbehaves. Nobody has ever run it. Before migrating a client's production infrastructure to a new cluster, I built a throwaway local setup specifically to break on purpose. Here's the walkthrough — steal it before your next migration, not during it.
What you need
minikube,kubectl,terraform, theargocdCLI- A git repo you control (a throwaway GitHub repo is fine)
Step 1 — Spin up a disposable cluster
minikube start --cpus=4 --memory=8192 --driver=docker
kubectl get nodes
Step 2 — Install Argo CD with Terraform
Use Terraform instead of a raw manifest — it's likely the same tool you'll use to provision the real cluster's supporting resources, so tracking infra state starts here.
# main.tf
terraform {
required_providers {
kubernetes = { source = "hashicorp/kubernetes", version = "~> 2.27" }
helm = { source = "hashicorp/helm", version = "~> 2.12" }
}
}
provider "kubernetes" { config_path = "~/.kube/config" }
provider "helm" {
kubernetes { config_path = "~/.kube/config" }
}
resource "kubernetes_namespace" "argocd" {
metadata { name = "argocd" }
}
resource "helm_release" "argocd" {
name = "argocd"
repository = "https://argoproj.github.io/argo-helm"
chart = "argo-cd"
namespace = kubernetes_namespace.argocd.metadata[0].name
version = "6.7.3"
}
terraform init
terraform apply -auto-approve
kubectl -n argocd get secret argocd-initial-admin-secret \
-o jsonpath="{.data.password}" | base64 -d
kubectl -n argocd port-forward svc/argocd-argocd-server 8080:443 &
argocd login localhost:8080 --username admin --password <password> --insecure
Step 3 — Structure the deployment as app-of-apps
Layout in your git repo:
apps/
root-app.yaml
weather/
argocd-app.yaml
deployment.yaml
service.yaml
apps/root-app.yaml — one Application that manages the rest from Git:
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: root-app
namespace: argocd
spec:
project: default
source:
repoURL: https://github.com/<you>/gitops-sandbox.git
targetRevision: main
path: apps
destination:
server: https://kubernetes.default.svc
namespace: argocd
syncPolicy:
automated:
prune: true
selfHeal: true
apps/weather/argocd-app.yaml — the child app:
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: weather-api
namespace: argocd
spec:
project: default
source:
repoURL: https://github.com/<you>/gitops-sandbox.git
targetRevision: main
path: apps/weather
destination:
server: https://kubernetes.default.svc
namespace: weather
syncPolicy:
automated:
prune: true
selfHeal: true
kubectl create namespace weather
kubectl apply -f apps/root-app.yaml
argocd app list
Step 4 — Deploy something disposable
The app doesn't matter — a trivial weather API works fine. What matters is that it has replicas, a rollout strategy, and a real sync policy, so it fails the way a production deployment fails.
# apps/weather/deployment.yaml
apiVersion: apps/v1
kind: Deployment
metadata:
name: weather-api
namespace: weather
spec:
replicas: 2
strategy:
type: RollingUpdate
rollingUpdate: { maxUnavailable: 1, maxSurge: 1 }
selector: { matchLabels: { app: weather-api } }
template:
metadata: { labels: { app: weather-api } }
spec:
containers:
- name: weather-api
image: ghcr.io/example/weather-api:1.0.0
ports: [{ containerPort: 8080 }]
git add . && git commit -m "deploy weather-api" && git push
argocd app sync weather-api
kubectl -n weather get pods -w
Drill 1 — Manual fix vs. self-heal
With selfHeal: true, GitOps reconciles continuously. That also means an emergency manual fix gets silently undone.
# Emergency scale-down directly on the cluster
kubectl -n weather scale deployment weather-api --replicas=0
# Argo CD notices the drift and reverts it within seconds
kubectl -n weather get pods -w
Now disable self-heal on purpose and repeat:
argocd app set weather-api --self-heal=false
kubectl -n weather scale deployment weather-api --replicas=0
kubectl -n weather get pods
# this time it sticks
# resume normal operation
argocd app set weather-api --self-heal=true
argocd app sync weather-api
Knowing which flag to flip, and remembering to flip it back, is not obvious the first time. It's obvious the second time, if the first was a drill.
Drill 2 — Reverting a bad scaling change
Push a bad replica count through Git, the way a real change would arrive:
sed -i '' 's/replicas: 2/replicas: 0/' apps/weather/deployment.yaml
git commit -am "oops" && git push
argocd app sync weather-api
kubectl -n weather get pods
Revert it the correct way — through Git, so the desired state stays truthful:
git revert HEAD --no-edit && git push
argocd app sync weather-api
Then time the fast-but-wrong path for comparison:
kubectl -n weather scale deployment weather-api --replicas=2
The manual command is faster in the moment and gets erased by the next sync. Knowing that tradeoff before an incident changes what your team reaches for under pressure.
Drill 3 — Rolling back a bad deployment
Break the image tag on purpose:
sed -i '' 's/1.0.0/does-not-exist/' apps/weather/deployment.yaml
git commit -am "broken release" && git push
argocd app sync weather-api
kubectl -n weather get pods
# ImagePullBackOff
Roll back through Argo CD's own history:
argocd app history weather-api
argocd app rollback weather-api <REVISION_ID>
Or, as the emergency fallback when you can't wait for Git or CI:
kubectl -n weather rollout undo deployment/weather-api
Time both. If the Argo CD rollback takes ten minutes on a sandbox with nothing on the line, it will not magically take two during a real migration weekend.
Turning the drills into a runbook
Write the exact commands you just ran — including the flags you had to look up — into a RUNBOOK.md next to the app-of-apps manifests. Not the theory. The commands, in the order that worked.
These drills are what shaped decisions on the real migration that followed: event-driven autoscaling backed by a message queue instead of static replica counts (manual scaling reverts were the clumsiest drill), a managed Postgres instance instead of in-cluster storage, a dedicated workflow engine instead of ad-hoc jobs, and a remote Terraform state backend chosen upfront instead of as an afterthought. None of that came from a whiteboard — it came from watching what was awkward to recover from in the sandbox.
What this costs
A local cluster and a disposable app cost an afternoon. A rollback that fails under pressure during a real migration costs a lot more — in downtime, and in explaining afterward why the documented procedure didn't work when it mattered. Run the drills on a system with nothing to lose, before production runs them for you on a schedule it controls.
Conoce mi trabajo
Si esto te fue útil, mira las plataformas que he construido o escríbeme sobre un proyecto.