Scaling
More backends
Backends are stateless, and sticky sessions are not needed:
docker compose -f docker-compose.prod.yml --env-file .env up -d --scale backend=2
- Caddy discovers replicas through DNS and balances least-connections.
- Jobs for an agent connected to another replica are dispatched with
PostgreSQL
LISTEN/NOTIFY(channelagent_dispatch). - Periodic jobs (retention, offline sweep, winget index, rollouts, alert evaluation) run on the elected leader only.
- Sign-in limits are counted in the database, so N replicas do not give N times the attempts.
- Revocations and tenant suspensions reach every replica.
- Remote desktop streams are too large
for
NOTIFY. When a technician's viewer reaches a different replica than the device's stream, the viewer's replica connects straight to the device's replica at itsRMM_NODE_URL(defaulthttp://<container hostname>:8080, which resolves on the compose network) and relays through it. Caddy refuses this internal path from outside.
The default node URL is plain http://: replica-to-replica relays then
carry screen content unencrypted over the network between them. That is
fine within one Docker host (the compose network), but when replicas run
on different hosts, set RMM_NODE_URL to an https:// address that the
other replicas reach (with a certificate they trust), or keep the replicas
on a private, encrypted network.
Database connections
Replicas × RMM_DATABASE_MAX_CONNS must stay below max_connections
(200) minus about 20 for maintenance. Raise PG_MAX_CONNECTIONS with the
memory it needs, or add PgBouncer in session mode. Transaction mode is
not supported: the tenant scope travels in session settings.
Retention and partitions
- Retention runs daily in batches of 10,000 rows; progress is in
rmm_retention_rows_deleted_total. device_metricsis partitioned by month. Partitions are created two months ahead and dropped by retention.- Per-tenant job and audit windows are under Tenant settings.
Load test reference
5,000 agents on one host: a 1,000-device deployment completed in 28 s, and the portal answered with p95 127 ms at 50 requests per second. Heartbeats are batched into one update per second, unchanged inventory sections are skipped, and ingestion concurrency is capped at a quarter of the pool. Details: Performance.