Backups and restore
What to back up
| What | Why |
|---|---|
| The database | Everything except files. |
The Entrosity Hub database (platform) | Users, credentials, organizations and roles. |
The bucket (rmm-packages) | Package installers, script output, agent and connector MSIs. |
deploy/.env | Without RMM_MASTER_KEY (and PLATFORM_MASTER_KEY for the Hub), encrypted secrets in a backup are lost. |
Nightly backup
deploy/scripts/backup.sh:
pg_dump -Fc, verified by reading it back withpg_restore --list, intodeploy/backups/db/(with a.metafile holding the schema version). Dumps are written with mode 0600. When the Hub's databaseplatformexists, it is dumped too, asplatform-<same time>.dump.- An incremental
mc mirrorof the bucket todeploy/backups/objects/. - Deletes dumps older than
RMM_BACKUP_KEEP_DAYS(default 14). - Optionally copies everything to an off-site S3 bucket
(
RMM_BACKUP_S3_ENDPOINT,_BUCKET,_ACCESS_KEY,_SECRET_KEY) and applies the same retention there.
15 2 * * * root /opt/entrosity/deploy/scripts/backup.sh >> /var/log/rmm-backup.log 2>&1
On Entrosity's production host this line is not installed yet; every deploy takes a backup before it upgrades (Continuous deployment).
A backup on the same disk does not survive the loss of the host.
Point-in-time recovery
Set PG_ARCHIVE_MODE=on and PG_ARCHIVE_COMMAND in .env to archive WAL
segments to backups/wal/. Restore them with standard PostgreSQL PITR,
starting from a base backup made with pg_basebackup.
Restore
deploy/scripts/restore.sh deploy/backups/db/rmm-20260924T021500Z.dump
The script:
- stops the portal and the backends;
- recreates the database (and the
rmm_approle if missing), and the Hub's database from theplatform-dump of the same run when there is one (and theplatform_approle); - restores the dump (in parallel when it lies inside the backup directory);
- mirrors the objects back;
- runs
migrate up, which moves an older dump forward and sets thermm_apppassword, and the Hub'splatform-migrate; - starts everything.
Agents reconnect by themselves.
On a new host: install Docker, copy deploy/.env (the same keys!) and
the backup directory, then run restore.sh. It starts only postgres and
minio before restoring.
Moving production to another host (done on 2026-09-27, see
CUTOVER.md in entrosity-infra): prepare the new host first (deploy/,
.env with the new SPHERE_MEDIA_PUBLIC_IP, release-public-key.txt, the
images, the caddy-data and caddy-config volumes so the certificates
come along, rmm.service, the backup cron line and the deploy key). Then
stop web, the backends and sphere-media on the old host, run
backup.sh, copy the new dumps and backups/objects/, run restore.sh
on the new host, compare row counts and point the DNS records at it.
Finally disable rmm.service on the old host and change DEPLOY_HOST and
DEPLOY_KNOWN_HOSTS (Continuous deployment).
Restore drill
The drill (2026-09-24) used the production compose on a lab host with 20 simulated agents connected over TLS. It deleted the PostgreSQL and MinIO volumes, then restored from the nightly dump.
| Step | Time |
|---|---|
| Backup (dump + verify + object mirror) | 2 s |
| Stop services, start empty postgres/minio | 6 s |
Recreate database + pg_restore | 2 s |
| Object mirror back | 1 s |
migrate up + start backends and portal (healthy) | 10 s |
| Total restore | 19 s |
Row counts matched, the admin could sign in, the package object was present, and all 20 agents were back online within 30 s without any action on them.
The drill database was small: plan on roughly 1 minute per GB of dump for
pg_restore -j 4 on SSD. Repeat the drill after major upgrades, on a
scratch host or with RMM_COMPOSE_PROJECT=rmm-drill.