Skip to main content

Backups and restore

What to back up​

WhatWhy
The databaseEverything except files.
The Entrosity Hub database (platform)Users, credentials, organizations and roles.
The bucket (rmm-packages)Package installers, script output, agent and connector MSIs.
deploy/.envWithout RMM_MASTER_KEY (and PLATFORM_MASTER_KEY for the Hub), encrypted secrets in a backup are lost.

Nightly backup​

deploy/scripts/backup.sh:

  1. pg_dump -Fc, verified by reading it back with pg_restore --list, into deploy/backups/db/ (with a .meta file holding the schema version). Dumps are written with mode 0600. When the Hub's database platform exists, it is dumped too, as platform-<same time>.dump.
  2. An incremental mc mirror of the bucket to deploy/backups/objects/.
  3. Deletes dumps older than RMM_BACKUP_KEEP_DAYS (default 14).
  4. Optionally copies everything to an off-site S3 bucket (RMM_BACKUP_S3_ENDPOINT, _BUCKET, _ACCESS_KEY, _SECRET_KEY) and applies the same retention there.
15 2 * * * root /opt/entrosity/deploy/scripts/backup.sh >> /var/log/rmm-backup.log 2>&1

On Entrosity's production host this line is not installed yet; every deploy takes a backup before it upgrades (Continuous deployment).

Keep an off-site copy

A backup on the same disk does not survive the loss of the host.

Point-in-time recovery​

Set PG_ARCHIVE_MODE=on and PG_ARCHIVE_COMMAND in .env to archive WAL segments to backups/wal/. Restore them with standard PostgreSQL PITR, starting from a base backup made with pg_basebackup.

Restore​

deploy/scripts/restore.sh deploy/backups/db/rmm-20260924T021500Z.dump

The script:

  1. stops the portal and the backends;
  2. recreates the database (and the rmm_app role if missing), and the Hub's database from the platform- dump of the same run when there is one (and the platform_app role);
  3. restores the dump (in parallel when it lies inside the backup directory);
  4. mirrors the objects back;
  5. runs migrate up, which moves an older dump forward and sets the rmm_app password, and the Hub's platform-migrate;
  6. starts everything.

Agents reconnect by themselves.

On a new host: install Docker, copy deploy/.env (the same keys!) and the backup directory, then run restore.sh. It starts only postgres and minio before restoring.

Moving production to another host (done on 2026-09-27, see CUTOVER.md in entrosity-infra): prepare the new host first (deploy/, .env with the new SPHERE_MEDIA_PUBLIC_IP, release-public-key.txt, the images, the caddy-data and caddy-config volumes so the certificates come along, rmm.service, the backup cron line and the deploy key). Then stop web, the backends and sphere-media on the old host, run backup.sh, copy the new dumps and backups/objects/, run restore.sh on the new host, compare row counts and point the DNS records at it. Finally disable rmm.service on the old host and change DEPLOY_HOST and DEPLOY_KNOWN_HOSTS (Continuous deployment).

Restore drill​

The drill (2026-09-24) used the production compose on a lab host with 20 simulated agents connected over TLS. It deleted the PostgreSQL and MinIO volumes, then restored from the nightly dump.

StepTime
Backup (dump + verify + object mirror)2 s
Stop services, start empty postgres/minio6 s
Recreate database + pg_restore2 s
Object mirror back1 s
migrate up + start backends and portal (healthy)10 s
Total restore19 s

Row counts matched, the admin could sign in, the package object was present, and all 20 agents were back online within 30 s without any action on them.

The drill database was small: plan on roughly 1 minute per GB of dump for pg_restore -j 4 on SSD. Repeat the drill after major upgrades, on a scratch host or with RMM_COMPOSE_PROJECT=rmm-drill.