Skip to main content

Performance

The detailed reports stay in the repository under docs/perf/ (2026-09-baseline.md, 2026-09-load-test.md). This page summarises them and shows how to measure.

Load test (September 2026)​

One 2-vCPU / 8 GB VM with PostgreSQL 16 (2 GB database, 70,000 other devices, 5.6 M software rows), one backend replica (pool of 20, RLS on), 5,000 simulated agents and k6. The simulated agents share the CPUs, so the numbers are conservative.

MeasureResult
Enrollment of 5,000 agents (80 programs each)all enrolled in 59 s
Reconnect of 5,000 agents after a backend restartall back in 12 s
Memory per WebSocket connection≈ 44 KB (19 KB heap + 25 KB stacks)
Heartbeat writes (5,000 agents)2 batched statements per second, ≈ 20 ms DB time per second
1,000-device deployment, concurrency 20028 s; create call 6.4 s; no deadlocks
Portal API at 50 req/sp50 18 ms, p95 127 ms, p99 305 ms, 0 errors

What the test fixed: enrollment no longer serialises on the token row; inventory sections are fingerprinted and unchanged ones skipped; ingestion concurrency is capped at a quarter of the pool; heartbeats are batched; deployment result handling is deadlock-safe.

API baseline​

Measured on a tenant with 50,000 devices. Target: device list p95 under 200 ms. Met for the list and every filter except the software filter (p95 ≈ 540 ms) and the deployment preview with a software condition (p95 ≈ 890 ms), which go through rmm_devices_with_software.

Measuring​

  1. A separate Postgres with the statistics extensions:
    docker run -d --name rmm-perf -p 55432:5432 -e POSTGRES_USER=rmm -e POSTGRES_PASSWORD=rmm -e POSTGRES_DB=rmm \
    postgres:16-alpine -c shared_preload_libraries=pg_stat_statements,auto_explain \
    -c auto_explain.log_min_duration=100ms -c auto_explain.log_analyze=on -c shared_buffers=1GB
    docker exec rmm-perf createdb -U rmm platform
    export RMM_DATABASE_URL=postgres://rmm:rmm@localhost:55432/rmm?sslmode=disable
    export PLATFORM_DATABASE_URL=postgres://rmm:rmm@localhost:55432/platform?sslmode=disable
    go run ./backend/cmd/server migrate up
    go run ./platform/backend/cmd/server migrate up
  2. Seed it (about 8 minutes on 2 vCPUs):
    go run ./backend/cmd/seed --tenants 20 --devices 70000 --largest 50000 \
    --software-per-device 80 --jobs 500000 --audit 200000
    Seeding needs both databases (RMM_DATABASE_URL and PLATFORM_DATABASE_URL, or --database-url and --platform-database-url): every tenant is also created as an organization with Axis on the Hub, with the organization admin admin@<slug>.seed.test (password --password, default seed-password-123), who is tenant admin in Axis. Axis's copy is written at the same time, so the first sync changes nothing.
  3. Run the Hub (make dev-platform) and the server with RMM_API_RATE_PER_SECOND=10000 and RMM_API_RATE_BURST=10000, then:
    go run ./backend/cmd/apibench --api http://localhost:8080 --hub http://localhost:8090 \
    --email admin@<slug>.seed.test --tenant <tenant id> --n 40
    It signs in on the Hub (--hub, default http://localhost:8090), renews its five-minute product token as needed, and prints p50/p95/max per endpoint as a Markdown table.
  4. Agents and portal traffic: simagent --count 5000 … and k6 run -e BASE_URL=http://localhost:8080 -e HUB_URL=http://localhost:8090 -e EMAIL=… -e PASSWORD=… -e TENANT=<tenant id> loadtest/k6/portal.js. The k6 script signs in on the Hub once; each virtual user mints its own product token from that session and renews it before it expires.

Slow queries​

SELECT round(mean_exec_time::numeric, 1) AS mean_ms, calls,
round(total_exec_time::numeric / 1000) AS total_s,
left(regexp_replace(query, '\s+', ' ', 'g'), 200) AS query
FROM pg_stat_statements ORDER BY total_exec_time DESC LIMIT 20;

To see a plan as the application sees it, SET ROLE rmm_app; SET app.tenant_id = '<id>'; first. Under RLS, non-leakproof operators (LIKE, ILIKE, @>) cannot use an index before the policy is checked; use a SECURITY DEFINER function like rmm_devices_with_software.