sadqwes Русский
← All cases

Home lab

Backing up the backups: a dashboard found 97%

I built a Grafana dashboard for my backups — and it immediately showed a problem that was invisible in the terminal.

ResultBackup size 8.65 GiB → 427 MiB, storage 2% full instead of 97%
VeleroMinIOGrafanaPrometheuspg_dumpCronJob

Before

In my cluster, Velero makes a backup to MinIO every 6 hours, and a script on my Mac copies the bucket outside the cluster. Everything “worked”: backups were created, and there were no errors.

What the dashboard showed

I built a backup dashboard in Grafana, and it showed right away: MinIO was 96.7% full, and every new backup was bigger than the previous one: 1.8 → 5.4 → 8.7 GiB.

The backup table showed the reason. The schedule saved all namespaces, including minio-system — the volume where the backups themselves live. Every backup copied all the previous ones into itself. The second reason: the schedule had been created by hand with 30-day retention, and the fix to 7 days had only made it into the README.

What I did

  • Expanded the MinIO volume first, so the current backups were safe.
  • Moved the schedule into Git, so ArgoCD manages it now: 7-day retention, without minio-system and monitoring.
  • Deleted old backups and unused kopia repositories — first the records in the cluster, then the data.

The result

A backup is 427 MiB instead of 8.65 GiB, and MinIO is 2.1% full instead of 96.7%.

Next — database backups

A volume snapshot is not a database backup yet. So for PostgreSQL I added a second level:

  • a CronJob with pg_dump puts a dump into S3 (MinIO) every day and deletes dumps older than 14 days;
  • a restore test: a Job downloads the newest dump, restores it into a temporary database and compares the row count of every table with production;
  • both Jobs push metrics to a Pushgateway, and the dashboard shows the age of the last backup and of the last restore test.

At work I used the same idea: PostgreSQL backups went to two places at once — DigitalOcean S3 and a local NAS in the office.

What I learned

  • Don’t back up the backup storage into itself.
  • Everything that creates backups should live in Git, not in manual commands.
  • Monitoring shows what you can’t see in the terminal.
  • A backup that was never restored is a hope, not a backup.

← All cases