Hubuum operator package¶
The same seven Grafana dashboards and Prometheus rules serve single-host and
distributed installations: overview/SLO, API, PostgreSQL, tasks/workers, events,
identity/integrations, and storage/recovery. Use assets from your server's Git
tag; main describes development. Older releases contain the initial overview
and alerts. The complete manifest and installer integration first appear after
0.0.16. Thresholds are starting points for tuning to your workload.
Single-host installation¶
Add --monitoring to scripts/install-single-host.sh, or enable it on an
existing installation:
If this reports Unknown argument: --monitoring, the installed updater predates
monitoring support. Arguments are parsed before any script refresh, and older
updaters did not all refresh the management scripts. Run the installer from your
deployed server release once; for v0.0.17:
curl -fsSL https://raw.githubusercontent.com/hubuum/hubuum/v0.0.17/scripts/install-single-host.sh \
| sudo bash -s -- --dir /opt/hubuum --script-ref v0.0.17 --monitoring
This reuses the saved installation settings and secrets, replaces the management
scripts, and performs the normal application update with monitoring enabled.
It retains your application image choices and pins future management-script
refreshes to v0.0.17; change --script-ref when adopting another release.
For v0.0.17 with a Podman Compose provider that reports
missing services [hubuum-migrate], use the
administration-profile workaround.
Both all and backend modes support Docker Compose and rootful Podman Compose.
Monitoring stays disabled unless requested. In all mode the frontend domain
serves /grafana/ and /prometheus/; in backend mode the API domain serves
both. Shared-host bff, direct, and prefixed modes reserve these same paths.
Caddy terminates TLS and preserves application prefixes, redirects, asset URLs
and Grafana Live connections.
Grafana uses its native login, with anonymous access and signup disabled.
Prometheus requires separate HTTP Basic credentials at Caddy. Both initially
use username admin, with independently generated passwords. These accounts
are separate from Hubuum authentication. No monitoring ports are published
directly. Retrieve initial credentials locally as root:
Do not paste this output into logs or tickets. Change the Grafana password in
Grafana; its database is authoritative after first startup. The .env value
remains the bootstrap password. To change the Prometheus password, update
PROMETHEUS_PASSWORD in the root-readable .env and run the updater, which
regenerates the Caddy hash. Use a long random hexadecimal value. Grafana's
authentication options
can support separately managed SSO; the installer does not grant Hubuum users
Grafana access automatically.
Prometheus scrapes hubuum-api and hubuum-api-standby directly every 15 seconds.
The primary includes workers; the standby serves HTTP only. Both share a stable
deployment label, initially the API hostname, and distinct instance labels.
They never fail over to one another as scrape targets.
Setting in .env |
Default | Purpose |
|---|---|---|
MONITORING_ENABLED |
false |
Persisted opt-in; set by --monitoring |
MONITORING_DEPLOYMENT |
API hostname | Stable identifier for one database |
MONITORING_ASSETS_REF |
auto |
Derive package from backend tag; --monitoring-ref overrides |
PROMETHEUS_IMAGE |
Prometheus 3.13.1, digest-pinned | Override with --prometheus-image |
GRAFANA_IMAGE |
Grafana 13.2.3, digest-pinned | Override with --grafana-image |
PROMETHEUS_RETENTION_TIME |
31d |
Allows a full 30-day SLO window once enough data is collected |
PROMETHEUS_RETENTION_SIZE |
5GB |
TSDB retention bound, not a filesystem quota; reserve WAL/head space |
PROMETHEUS_MEMORY_LIMIT |
512m |
Container limit; tune for cardinality and queries |
GRAFANA_MEMORY_LIMIT |
512m |
Container limit |
Each monitoring container also has a one-CPU limit. Reserve additional host disk and memory before increasing workload or retention. The first retention limit reached applies, so the configured time window is not guaranteed by time alone.
Updates retain image choices, credentials, labels, named data volumes and
operator-edited configuration. They refresh maintained dashboards and rules
from the selected backend release and recreate monitoring containers to load
them. Source builds use their checkout's package. latest resolves the latest
server release; custom images and non-release tags require --monitoring-ref.
An unavailable/incomplete package stops refresh instead of silently selecting
another version. Inspect /opt/hubuum/monitoring/asset-source.txt for its source.
The initial monitoring/prometheus/prometheus.yml and Grafana provisioning files
are retained on updates. Add custom rule files under monitoring/prometheus/rules/
and dashboards under monitoring/grafana/dashboards/; reserve supplied filenames
for the maintained package. targets.json is regenerated for the configured
port and deployment label. Copy a supplied dashboard to a new UID to customize it.
To disable monitoring, set MONITORING_ENABLED=false in .env and run the
updater. This removes the monitoring containers and Caddy routes while retaining
configuration, credentials, Grafana's encryption key, and data volumes. Re-enable
with update-single-host.sh --monitoring to reuse that data and authentication.
Re-running the installer, including --recreate, also preserves these secrets.
Stop and ordinary uninstall preserve data and configuration. Explicit
uninstall-single-host.sh --purge removes Compose volumes and the installation
directory, including monitoring data. Application backups do not contain these
volumes; back up Grafana and required time series separately.
Direct Prometheus and Grafana installations¶
Copy prometheus/alerts.json and prometheus/recording-rules.json to any
Prometheus server. JSON is valid YAML. Load both through rule_files and
configure direct process targets as in
the example configuration. Use distinct
deployment labels for separate databases and stable unique instance labels
for every process. Do not scrape an application load balancer.
Import every dashboards/*.json into Grafana, then select a Prometheus
datasource and deployment. Datasource selection is a Grafana variable, so no
installer-specific UID or URL is embedded in the dashboards. Alternatively use
a file dashboard provider pointing to that directory. Dashboard UIDs and rule
group names remain stable across installations.
For Prometheus Operator, apply the equivalent resource:
Adjust its namespace and labels to match your Prometheus ruleSelector and
ruleNamespaceSelector. Configure ServiceMonitor or PodMonitor separately;
preserve deployment, instance, and the canonical job="hubuum" label.
Configure deduplication in the querying layer when using HA Prometheus replicas.
The resource contains exactly the groups in the directly consumed rule files.
An eventual Helm chart can package these files without another set of queries;
this package does not install Prometheus Operator or Helm.
Local Compose example¶
For the repository's development stack, use the optional overlay. Set its credentials in your shell; preserve them securely for later restarts:
export GRAFANA_ADMIN_PASSWORD="$(openssl rand -hex 24)"
export GRAFANA_SECRET_KEY="$(openssl rand -hex 32)"
export PROMETHEUS_PASSWORD="$(openssl rand -hex 24)"
export PROMETHEUS_PASSWORD_HASH="$(printf '%s\n' "$PROMETHEUS_PASSWORD" |
docker run --rm -i caddy:2-alpine caddy hash-password)"
# These public configuration assets must be readable by the container users.
chmod -R a+rX observability
docker compose -f docker-compose.yml -f observability/compose.monitoring.yml \
--profile monitoring up -d
First follow the development database/migration setup.
The local overlay exposes https://localhost:9443/grafana/ and
https://localhost:9443/prometheus/ through Caddy with a local development CA.
Trust only that local CA for browser use; public single-host installations use
the installer's ACME certificates. This example scrapes the development stack's
single hubuum process and mounts the same committed rule and dashboard files.
Compose down retains its named volumes; explicit down --volumes removes them.
Service objectives¶
| Objective | Initial SLI and target | Applicability and exclusions |
|---|---|---|
| API availability | 99.9% non-5xx over 30 days | /api/ route templates; denominator is 2xx, 3xx and 5xx. Excludes all 4xx, probes, metrics, Swagger, OpenAPI and unclassified routes. Review overload/authentication failures separately. |
| API latency | 99% of successful API responses within 1 second over 30 days | 2xx/3xx API responses only; HTTP time, not asynchronous execution. Split heavy routes into separate objectives when appropriate. |
| Task service | Oldest queued task below 300 seconds | Tune by kind. Histograms only observe claimed tasks, so oldest age and worker capacity also matter. This is an operational age objective, not a completion SLO. |
| Event delivery | Oldest actionable fanout/delivery below 300 seconds; no retained dead deliveries | Requires event processing. Net dead-letter growth is a gauge delta affected by retention, not a failure ratio. |
| Recovery readiness | Isolated restore verification succeeds within the configured maximum age | Requires the external job input below. Backup completion alone does not prove restorability. |
Availability alerts use two-window burn rates: 14.4 times budget over both 1 hour and 5 minutes, or 6 times over both 6 hours and 30 minutes. Latency uses the 1-hour/5-minute pair against its 1% budget. Idle/absent traffic yields no success ratio. Retain and collect 30 days before interpreting a monthly SLO.
Pool utilization and resources stay per instance. Counters are rated before
summing across processes. Database-wide task, inventory and event gauges use
max, never a replica sum. Missing values are not converted to healthy zeros.
Refresh-staleness checks cover failed inventory refreshes; target discovery and
Prometheus availability still require independent supervision.
Optional external inputs¶
external-metrics.json records dependencies outside the server metric contract. Missing optional inputs appear as no data, not success. Configure them for single-host or distributed deployments when applicable.
For readiness, use a Prometheus blackbox exporter with the HTTP module in
blackbox.example.yml. Probe every API /readyz
with job="hubuum-readiness", stable instance, and deployment labels. The
scrape example includes relabeling. HTTP status probes detect a reachable but
unready server; up alone cannot establish readiness.
For integration probes use job="hubuum-integrations" and bounded components
ldap, treetop, valkey, amqp, smtp, or webhook. Configure an appropriate
HTTP/TCP probe. Do not label by endpoint host, user, recipient, group or secret
reference. TCP reachability does not prove authentication or delivery. The
server exposes login/permission errors, secret resolution, remote HTTP and OTLP
results; provider-specific freshness and event-transport attempt/latency
counters are not currently emitted. Consult controlled logs and authorized
GET /api/v1/event-deliveries/health for that evidence.
The current contract also has no dedicated listener-health, lease-renewal, worker-shutdown, schema-readiness, maintenance-state, artifact-size or integrity gauges. Use readiness probes and administrator state for those checks. Task counts show active work rather than an invented active-worker gauge, and database error panels retain bounded caller/result categories rather than claiming to distinguish statement timeouts. Restore duration and outcome come from the external verification job, not from backup completion. Add native panels and alerts alongside the corresponding metric-contract additions.
For recovery, archive and externally supervised retention jobs, configure node exporter's textfile collector. Declare an expected job before its first run using Python 3.11+ (standard library only):
python3 scripts/observability.py record-job \
--directory /var/lib/node_exporter/textfile_collector \
--deployment production --operation restore_verify \
--max-age-seconds 86400 --init
Wrap your deployment's isolated verification command in its scheduler:
flock /run/hubuum-restore-verify.lock \
python3 scripts/observability.py record-job \
--directory /var/lib/node_exporter/textfile_collector \
--deployment production --operation restore_verify \
--max-age-seconds 86400 -- /usr/local/sbin/verify-hubuum-backup
The command must fail if restore or integrity checks fail. The helper preserves
its exit status, records duration/outcome atomically, and preserves the last
successful timestamp after failure. --init never claims a successful run.
Serialize runs per deployment/operation; the helper does not schedule jobs or
lock backup resources. Only restore_verify, event_archive, and retention
operations are accepted. Monitor the exporter's own scrape failures too;
removing its target is not evidence of health.
Walkthrough and notifications¶
- Open Hubuum operations in Grafana and choose the deployment. Check both
targets at
/prometheus/targetsand the API/worker role inventory. - Sign into Hubuum and load the Atlas example inventory from the getting-started documentation, or read existing collections. Request rate and pool activity should increase within two scrape intervals. Submit an export and inspect task completions and export phases.
- Open
/prometheus/alerts. On a disposable test installation, stop onlyhubuum-api-standby, wait five minutes and observeHubuumScrapeUnavailable. Start it again and confirm recovery; keep the primary serving traffic. - If external job recording is configured, wrap a harmless failing test command under a disposable deployment label and inspect the recovery dashboard. Remove only that test collector file afterwards.
Prometheus alerts do not send messages without Alertmanager. Add your Alertmanager targets to the preserved Prometheus configuration; configure routing, receivers, credentials, grouping and inhibition under your own on-call policy. Grafana contact points are separate. The installer does not choose destinations or reuse Hubuum event-sink credentials.
Monitoring on the same host cannot independently report total host failure. Use external probes and monitoring when that coverage matters.
Maintain and validate¶
Models and additional rules live in scripts/monitoring/generate.py; the
initial rule group remains maintained in prometheus/alerts.json. Run:
python3 scripts/observability.py generate
python3 scripts/observability.py check --promtool
python3 tests/python/run.py unit monitoring deployment.test_monitoring
python3 tests/python/run.py integration monitoring-fixture --engine docker
Validation checks every dashboard/rule metric, label and enum against the
server contract and explicit external list, checks runbooks in both directions,
rejects direct sums of shared gauges, checks SLI exclusions, and compares
Operator/direct groups. Pinned Prometheus parses every query and evaluates
firing, recovery, deduplication and SLI-exclusion fixtures. CI exercises installer
configuration and lifecycle. The integration monitoring-fixture command checks Docker and Podman
transport/routing with two independent metrics fixtures. The production-container
CI job also runs the real-server acceptance test:
python3 tests/python/run.py integration monitoring --image hubuum-server:ci \
--report target/monitoring-acceptance.json
Build the production image first. This test runs the actual single-host installer
in backend mode with PostgreSQL, both Hubuum processes, the restore executor,
Caddy, Prometheus and Grafana. Add --mode all to include the frontend and Valkey.
It verifies native Grafana login, all dashboard queries, Atlas import/export,
SQL versus metric counts, exact HTTP counter deltas, SLI exclusions, recording
rules, and the real five-minute scrape alert followed by recovery. It also creates
two webhook deliveries from one collection update through the API: one succeeds;
the other receives HTTP 503 responses, becomes retryable, and exhausts its two
configured attempts. Pending, failed, retryable, dead and recovered snapshots must
match exact SQL and health API counts on both Prometheus targets and the actual
Grafana event panel. The shared-database recording must deduplicate the targets.
Receiver access logs must show the same event UUID on every HTTPS attempt.
The production ten-minute dead-letter alert must become pending, fire once for the deployment, and recover after the receiver accepts an administrator-triggered retry. Queue rows and alert durations are never rewritten. A temporary worker uses the installed image and runtime database credentials, a two-minute retry backoff, the existing private-target setting and the installation's disposable CA; certificate verification remains enabled. The fixture briefly pauses that worker so the due retry remains observable. Database snapshots are allowed their documented cache and scrape intervals to converge. The receiver's extra internal Caddy host and the worker are removed before lifecycle checks. Allow approximately 25 minutes for the complete acceptance run (35-minute CI deadline). Updates, disable/re-enable, uninstall/restart and purge exercise credentials, application data, Grafana state and historical Prometheus samples.
Each run uses a unique Compose project and loopback port and purges its resources, including on failure. Only the root guard, download sources, global container names, published ports and bridge subnet are adapted in temporary copies; the rollout, health checks, database setup and alert hold remain unchanged. The provided server image and checkout assets are used throughout updates. Reports contain non-secret results and the failing stage, never generated credentials.
The CLI groups generation, validation, tests and external job recording under
python3 scripts/observability.py; its implementation uses normal modules in
scripts/monitoring/, with Python 3.11+ and no third-party packages.
Every metric-contract change requires reviewing this package in the same pull request. Update affected models, dependencies, fixtures and runbooks; do not suppress missing metrics. Version files with the server release and pin runbook links to that tag when immutable incident guidance is required.