Skip to main content

Monitoring

Prometheus and Grafana​

docker compose -f docker-compose.prod.yml -f docker-compose.monitoring.yml --env-file .env up -d
  • Prometheus (127.0.0.1:9090) scrapes every backend replica on :9091.
  • Grafana (127.0.0.1:3000, admin password GRAFANA_ADMIN_PASSWORD) loads the Entrosity Axis overview dashboard.
  • Both listen on localhost only; reach them through an SSH tunnel.
  • Connect Prometheus to your Alertmanager (see the comment in deploy/monitoring/prometheus.yml).

Alert rules​

deploy/monitoring/alerts.yml:

AlertFires when
RMMQueueBacklogJob queue depth above 1,000 for 10 minutes
RMMAgentConnectionsDroppingWebSocket connections drop by more than 20 % in 5 minutes
RMMHighErrorRateHTTP 5xx above 2 %
RMMDatabasePoolSaturatedDatabase pool above 80 %
RMMBackendDownA backend stops answering scrapes
RMMSlowAPIAPI p95 latency too high
RMMJobDispatchSlowJob dispatch too slow
RMMPlatformSyncStaleThe copy of Entrosity Hub's users, tenants and roles is older than 5 minutes for 5 minutes

Metrics​

MetricMeaning
rmm_ws_connections{kind}Connected agents and connectors per replica
rmm_ws_disconnects_totalWebSocket disconnects
rmm_agents_onlineAgents online
rmm_jobs_total{type,status}Finished jobs by type and result
rmm_job_dispatch_secondsTime from job creation to delivery
rmm_http_request_duration_seconds{method,route,status}Latency by route
rmm_river_queue_depth{state}Background job queue
rmm_db_pool_connections{state}Pool usage
rmm_deployments_activeRunning deployments
rmm_sse_subscribersOpen portal event streams
rmm_adsync_runs_total{status}AD sync runs
rmm_retention_rows_deleted_total{category}Retention progress
rmm_platform_sync_age_secondsAge of the replica's copy of Entrosity Hub's data
rmm_platform_syncs_total{result}Snapshot pulls: applied, unchanged, error

Profiles are at :9091/debug/pprof/ (internal only). There is no public /metrics route.

Health checks​

EndpointMeaning
/healthz (also /api/v1/healthz)The process is up.
/readyzThe database is reachable (used by Caddy and the container health check). checks.platform_sync also reports ok, stale (<age>) or never; a stale copy does not fail the check, so a Hub outage does not take Axis out of rotation.

The container health check runs rmm-server healthcheck.

Logs​

docker compose -f docker-compose.prod.yml logs -f backend

Logs are JSON in production, rotated by Docker (20 MB × 5 per container). Secrets are redacted. PostgreSQL logs statements slower than 1 s with their plan (auto_explain), and pg_stat_statements is enabled.