Monitoring
Prometheus and Grafana
docker compose -f docker-compose.prod.yml -f docker-compose.monitoring.yml --env-file .env up -d
- Prometheus (
127.0.0.1:9090) scrapes every backend replica on:9091. - Grafana (
127.0.0.1:3000, admin passwordGRAFANA_ADMIN_PASSWORD) loads the Entrosity Axis overview dashboard. - Both listen on localhost only; reach them through an SSH tunnel.
- Connect Prometheus to your Alertmanager (see the comment in
deploy/monitoring/prometheus.yml).
Alert rules
deploy/monitoring/alerts.yml:
| Alert | Fires when |
|---|---|
RMMQueueBacklog | Job queue depth above 1,000 for 10 minutes |
RMMAgentConnectionsDropping | WebSocket connections drop by more than 20 % in 5 minutes |
RMMHighErrorRate | HTTP 5xx above 2 % |
RMMDatabasePoolSaturated | Database pool above 80 % |
RMMBackendDown | A backend stops answering scrapes |
RMMSlowAPI | API p95 latency too high |
RMMJobDispatchSlow | Job dispatch too slow |
RMMPlatformSyncStale | The copy of Entrosity Hub's users, tenants and roles is older than 5 minutes for 5 minutes |
Metrics
| Metric | Meaning |
|---|---|
rmm_ws_connections{kind} | Connected agents and connectors per replica |
rmm_ws_disconnects_total | WebSocket disconnects |
rmm_agents_online | Agents online |
rmm_jobs_total{type,status} | Finished jobs by type and result |
rmm_job_dispatch_seconds | Time from job creation to delivery |
rmm_http_request_duration_seconds{method,route,status} | Latency by route |
rmm_river_queue_depth{state} | Background job queue |
rmm_db_pool_connections{state} | Pool usage |
rmm_deployments_active | Running deployments |
rmm_sse_subscribers | Open portal event streams |
rmm_adsync_runs_total{status} | AD sync runs |
rmm_retention_rows_deleted_total{category} | Retention progress |
rmm_platform_sync_age_seconds | Age of the replica's copy of Entrosity Hub's data |
rmm_platform_syncs_total{result} | Snapshot pulls: applied, unchanged, error |
Profiles are at :9091/debug/pprof/ (internal only). There is no public
/metrics route.
Health checks
| Endpoint | Meaning |
|---|---|
/healthz (also /api/v1/healthz) | The process is up. |
/readyz | The database is reachable (used by Caddy and the container health check). checks.platform_sync also reports ok, stale (<age>) or never; a stale copy does not fail the check, so a Hub outage does not take Axis out of rotation. |
The container health check runs rmm-server healthcheck.
Logs
docker compose -f docker-compose.prod.yml logs -f backend
Logs are JSON in production, rotated by Docker (20 MB × 5 per container).
Secrets are redacted. PostgreSQL logs statements slower than 1 s with
their plan (auto_explain), and pg_stat_statements is enabled.