Home lab
A Kubernetes lab as code
A cluster you can delete and bring back by following a runbook — together with its data.
Why
At work I used Docker and docker-compose, but the market wants Kubernetes. I decided not just to read about it, but to build a cluster the way it’s done in production: everything described as code, everything recoverable.
How it works
- Infrastructure as code: Terraform and libvirt create 3 virtual machines on KVM.
- Cluster as code: an idempotent Ansible playbook installs Kubernetes with kubeadm.
- Platform as code: everything in the cluster comes from Git through ArgoCD — 17 applications now.
- Secrets in Git — only encrypted, with SealedSecrets.
- Network and storage: ingress-nginx and MetalLB, volumes on Longhorn.
- Observability: Prometheus, Grafana, Loki, alerts described as code.
- Backups: Velero with a copy outside the cluster — more in the backups case.
- My own services go through the full cycle: GitHub Actions (tests, Semgrep, govulncheck, Trivy) → image in GHCR → CI commits the tag to the GitOps repo → ArgoCD.
The DR drill
I tested the main thing: can I lose everything and get it back? The scenario: terraform destroy → new VMs → a cluster from scratch → restore from backup. The drill found problems I would never have found otherwise:
- the new SealedSecrets controller generated a new key, so the old secrets didn’t decrypt — I had to restore the key;
- ArgoCD’s access to the private repository lived only in the old cluster;
- ArgoCD created the PVCs before Velero restored them, so the Postgres data didn’t come back. Lesson: databases need their own backup —
pg_dump.
All of this is in the runbook and the “known issues” table in the repository README.
What I learned
A DR plan that was never tested doesn’t exist. And the more is described as code, the fewer manual steps there are in a restore — steps where you can make a mistake at 3 a.m.