Skip to content

Runtime Metrics

Hubuum exposes low-cardinality runtime metrics through a Prometheus scrape endpoint. The endpoint is enabled by default at /metrics.

Configuration

Variable Default Description
HUBUUM_METRICS_ENABLED true Enables the Prometheus metrics scrape endpoint
HUBUUM_METRICS_PATH /metrics Literal absolute non-root endpoint path; it must not contain route patterns or collide with API, probe, OpenAPI, or Swagger UI routes

The endpoint is subject to HUBUUM_CLIENT_ALLOWLIST and uses the configured TLS settings. Put it behind network-level access controls appropriate for operational data.

Processes with the all or api runtime role serve metrics on the main HTTP listener. A worker process serves only the configured metrics path on HUBUUM_BIND_IP:HUBUUM_PORT when metrics are enabled; it does not expose the application API or health probes. Give API and worker containers separate network namespaces or ports if they run on the same host.

Scrape Each Process Directly

Counters, histograms, connection-pool gauges, and worker configuration describe one process. Configure Prometheus with a stable target for every API and worker process. Do not scrape a load-balanced public URL: successive scrapes can land on different processes, producing apparent counter resets and hiding worker-only activity.

Use these metrics to identify a target:

  • hubuum_build_info{version,git_sha}
  • hubuum_runtime_info{role}
  • hubuum_storage_backend_info{backend}
  • hubuum_process_start_time_seconds

Check the target mix before interpreting process-local worker metrics:

count by (role) (hubuum_runtime_info)

An API-only target reports hubuum_runtime_info{role="api"} and zero configured task and event workers. It can still expose database-wide task and template gauges, but export, import, remote-call, task-execution, and worker-loop counters and histograms are produced by worker-enabled targets. If those families are absent from an API scrape, verify that Prometheus has a separate worker target before treating the absence as no activity.

Inventory, task backlog, export-template identity, and event-queue gauges are database-wide snapshots. Every replica connected to the same database reports the same logical values. Aggregate those series with max or avg, not sum, unless replica multiplication is intentional. Process-local gauges such as database-pool state should normally be summed when calculating deployment-wide capacity.

Database-backed gauges use a 30-second in-process cache and refresh on a best-effort basis. If a refresh fails, /metrics still returns process metrics and retains the last successful database snapshot when one exists. The refresh duration, last-success timestamp, failures, and skipped concurrent refreshes make stale values visible.

Single-host deployments provide deterministic convenience routes. Shared-host direct routing sends /metrics to the worker-enabled primary and /metrics/standby to the HTTP-only standby. With shared-host prefixed routing, use /hubuum-api/metrics and /hubuum-api/metrics/standby. Shared-host bff routing does not expose either backend metrics endpoint publicly.

Cardinality Rules

Metric labels must stay bounded. Hubuum metrics do not use usernames, user IDs, client IPs, raw URL paths, object IDs, class names, collection names, rendered remote URLs, task IDs, idempotency keys, or error messages.

HTTP metrics use Actix route templates such as /api/v1/classes/{class_id}/objects/{object_id}. Requests that do not resolve to a registered route use a coarse route group instead of their raw path.

Export phase timing is aggregated by phase and outcome. Only the total export histogram carries template_id; this keeps template identity useful without multiplying every query, hydration, and render bucket by every stored template. hubuum_export_template_info maps current database IDs to mutable names and is reset on every inventory snapshot, so renames and deletions do not leave stale series in a process.

Use admin JSON/API endpoints and task logs for per-task detail.

Duration Units And Histograms

Prometheus duration metric names and observations use seconds. A 4.7 millisecond request is recorded as 0.0047; using seconds does not discard millisecond or sub-millisecond precision.

Hubuum uses three explicit bucket profiles:

  • HTTP and database latency: 0.0005 seconds through 30 seconds.
  • Outbound remote calls: 0.01 seconds through 120 seconds.
  • Queued and background work: 0.01 seconds through 3600 seconds.

The long background buckets are intentional: most buckets should be empty for fast work, while imports, exports, backups, remote dependencies, or a constrained worker can legitimately take seconds or minutes. Empty upper buckets are cheap and preserve visibility when an exceptional slow operation occurs.

Prometheus histogram buckets are cumulative. This query calculates HTTP p95 latency by route over five minutes:

histogram_quantile(
  0.95,
  sum by (le, route) (
    rate(hubuum_http_request_duration_seconds_bucket[5m])
  )
)

The histogram sum divided by its count gives the mean. Apply rate before division for a time window:

sum by (route) (
  rate(hubuum_http_request_duration_seconds_sum[5m])
)
/
sum by (route) (
  rate(hubuum_http_request_duration_seconds_count[5m])
)

The exported _sum is therefore useful, but only together with _count. Classic Prometheus histograms do not expose exact minimum or maximum observations. histogram_quantile estimates percentiles from bucket boundaries and remains aggregatable across processes.

Readiness probes commonly dominate request counts. Exclude probes and the scrape endpoint when calculating application traffic or an overall application latency. The example uses the default scrape route, /metrics. If HUBUUM_METRICS_PATH is customized, replace /metrics in the matcher with the exact configured path; for example, use route!~"/(healthz|readyz)|/internal/metrics" for /internal/metrics. Prometheus regex matchers are fully anchored:

sum by (route) (
  rate(
    hubuum_http_requests_total{
      route!~"/(healthz|readyz|metrics)"
    }[5m]
  )
)

Counter and histogram series generally appear only after a process observes a matching event. The worker-error and failed-backup counters start at zero so alerts can observe their first failure after scraping the baseline. Seeing only /readyz means that target handled readiness probes but not application requests; it does not mean other routes are filtered out.

For stored-template exports, join the total-duration histogram to the current template-info gauge:

histogram_quantile(
  0.95,
  sum by (le, template_id) (
    rate(
      hubuum_export_duration_seconds_bucket{
        template_id!="none"
      }[15m]
    )
  )
)
* on (template_id) group_left (template_name)
max by (template_id, template_name) (
  hubuum_export_template_info
)

Metrics

The complete metric inventory—including type, unit, process/database scope, allowed label names, feature ownership, and explicit histogram buckets—is generated in Metric Reference. The generated reference and docs/operational-contract.json are the canonical compatibility sources; CI rejects drift from the typed registry.

Process, HTTP, And Database

The generated Metric Reference replaces the former hand-maintained table for this group. The notes below explain the bounded domains and aggregation semantics that are not part of each metric's structural signature.

The bounded refresh source values are database, events, inventory, login_limiter, process, tasks, and token_keys.

The bounded database caller values are event_delivery, event_fanout, event_retention, http_request, metrics_refresh, readiness, request_maintenance, restore_coordinator, task_lease, task_worker, token_retention, and unattributed.

Storage capability, operation, and result values come from closed Rust enums or fixed call sites. Expected domain failures are included in hubuum_storage_operation_errors_total; use the result label to distinguish them from database, unavailable, and internal backend failures. Storage metrics describe logical use cases, while hubuum_db_* metrics describe the PostgreSQL implementation. Do not sum the two as though they were the same operation count.

Distributed Tracing

Metric Labels Description
hubuum_tracing_info sampling_mode Effective bounded sampling mode; value is 1
hubuum_tracing_sample_ratio none Effective root trace sampling ratio
hubuum_tracing_queue_capacity none Configured bounded batch-processor queue capacity
hubuum_tracing_queue_utilization none Current spans admitted and waiting for export
hubuum_trace_spans_total category, state Sampled allowlisted spans observed at the processor's actual started or ended lifecycle callback
hubuum_trace_spans_dropped_total reason Spans dropped for classification or queue_saturation
hubuum_trace_export_batches_total outcome OTLP export batches by success or failure
hubuum_trace_export_spans_total outcome Spans submitted in OTLP batches by success or failure
hubuum_trace_flushes_total outcome Graceful-shutdown flushes by success or failure

Trace labels come only from fixed categories and outcomes. They never contain trace IDs, span IDs, endpoints, resource identities, or error messages. Alert on a sustained nonzero queue_saturation rate and export failure rate. Queue utilization is process-local and should be compared with that process's capacity. See Distributed Tracing for configuration and data policy.

Storage capability labels follow the singular capability trait vocabulary: the trait's Storage suffix is removed and its remaining stem becomes snake case. The bounded values are audit_event, authentication, authorization_data, backup_snapshot, catalog, class, class_relation, collection_authorization_query, collection, computed_field, computed_object, event_configuration, event_delivery_administration, event_delivery_worker, event_fanout, event_health, event_retention, export_template, external_identity, group, history, group_membership, identity_scope, import, inventory, local_identity_credential, metrics, object_aggregate, object_relation, object, operational_state, principal, relation_query, remote_target, restore, service_account, task_execution, task_queue, token_retention, token, transaction, unified_search, and user.

The standard unprefixed process_* families are available on Linux, macOS, and Windows. The names intentionally match the Prometheus ecosystem so existing process dashboards and alerts can be reused. File-descriptor or handle gauges remain zero if the operating system cannot return that information. On Windows, process_open_fds counts kernel handles and process_max_fds is zero because Windows has no comparable fixed per-process handle limit. Prefer resident memory for cross-platform pressure alerts: virtual address-space size is especially large and not directly comparable on macOS. process_start_time_seconds is the operating system's process start, while hubuum_process_start_time_seconds records when Hubuum initialized metrics.

Useful starting queries include:

rate(process_cpu_seconds_total[5m])
process_resident_memory_bytes
(process_open_fds / process_max_fds)
and
(process_max_fds > 0)

Tasks, Exports, Imports, And Remote Calls

Metric Labels Description
hubuum_task_worker_iterations_total outcome Worker iterations by claimed, idle, or error outcome
hubuum_task_claims_total kind Tasks claimed by workers
hubuum_task_lease_recoveries_total kind Tasks failed after their owning worker lease expired
hubuum_task_completions_total kind, final_status Tasks reaching a terminal status
hubuum_task_cancellation_requests_total kind, result Durable requests with queued, active, or unchanged result
hubuum_task_stop_acknowledgements_total kind, reason Terminal stops with cancel_requested or deadline_exceeded reason
hubuum_task_cancellation_acknowledgement_duration_seconds kind, reason Time from persisted cancellation request to terminal acknowledgement
hubuum_task_ambiguous_remote_stops_total reason Stops after remote dispatch where external effects may have occurred
hubuum_task_queue_wait_duration_seconds kind Time from task creation to claim
hubuum_task_execution_duration_seconds kind, final_status Time from task start to finish
hubuum_task_workers_configured none Task workers configured in this process; zero on API-only processes
hubuum_task_poll_interval_seconds none Configured task-worker poll interval
hubuum_tasks kind, status Current database-wide task counts
hubuum_task_oldest_age_seconds kind, state Oldest queued and active task age per task kind
hubuum_task_last_terminal_timestamp_seconds kind, status Most recent finish time for each task kind and terminal status; zero means none is retained
hubuum_task_output_cleanup_runs_total kind Stored output cleanup runs for export or backup artifacts
hubuum_task_output_cleanup_failures_total kind Stored output cleanup failures
hubuum_task_output_cleanup_deleted_total kind Stored outputs deleted by cleanup
hubuum_export_template_info template_id, template_name Current stored export-template identities from the shared database
hubuum_export_phase_duration_seconds phase, outcome Aggregate export query, hydration, render, and total phase duration
hubuum_export_duration_seconds template_id, outcome Total export duration by stored template ID; ad-hoc exports use none
hubuum_export_completions_total scope, content_type Successfully persisted export outputs
hubuum_export_truncations_total scope, content_type Successfully persisted truncated exports
hubuum_export_warnings_total scope, content_type Warning count on successfully persisted exports
hubuum_import_phase_duration_seconds phase, outcome Import planning, execution, and total phase duration, including failures
hubuum_import_processed_items_total none Items processed by terminal import tasks
hubuum_import_succeeded_items_total none Import items completed successfully
hubuum_import_failed_items_total none Import items completed with failure
hubuum_remote_call_duration_seconds method, status_family, outcome Remote HTTP execution duration
hubuum_remote_call_results_total method, status_family, outcome Remote outcomes such as success, failure, timeout, or validation rejection

Export timer phases are limited to total, query, hydration, and render; their outcomes are success, error, or timeout. Import timer phases are limited to total, planning, and execution; their outcomes are success, failed, partially_succeeded, or error.

Computed Fields, Security, Events, And Inventory

Metric Labels Description
hubuum_computed_field_evaluations_total scope, outcome Computed-field evaluations by shared, personal, or preview scope and outcome
hubuum_computed_field_errors_total scope, code Computed-field runtime errors by stable bounded code
hubuum_computed_field_live_fallbacks_total none Stale shared materializations evaluated live during reads
hubuum_computed_field_read_repairs_total outcome Guarded stale-materialization repairs by success or failure
hubuum_computed_field_rebuild_batches_total items Computed-field rebuild batches classified as empty or non-empty
hubuum_computed_field_rebuild_completions_total status Computed-field rebuild terminal outcomes
hubuum_computed_field_rebuild_duration_seconds status Computed-field rebuild duration histogram
hubuum_login_attempts_total outcome Login attempts by success, bad credentials, rate-limited, or internal error
hubuum_login_lockouts_total scope Login limiter lockout transitions by principal/IP, IP, or subnet scope
hubuum_login_limiter_backend_failures_total backend, operation Shared login-limiter failures while local enforcement remains active
hubuum_login_limiter_entries state Active and locked login-limiter entries in this process
hubuum_client_allowlist_rejections_total reason Requests rejected for a disallowed or missing client IP
hubuum_revision_conditions_total outcome Conditional writes classified as matched, wildcard, stale, unconditional, malformed, async_stale, or invariant_failure; never labelled by resource identity or revision
hubuum_event_queue_items queue, state Database-wide fan-out and delivery queue items by bounded state
hubuum_event_stale_claims queue Stale fan-out and delivery worker claims
hubuum_event_oldest_age_seconds queue Oldest actionable fan-out or delivery item age
hubuum_event_workers_configured worker Event workers configured in this process
hubuum_event_worker_batch_size worker Configured event-worker batch size
hubuum_event_worker_poll_interval_seconds worker Configured event-worker poll interval
hubuum_event_worker_lock_timeout_seconds worker Configured event-worker claim lock timeout
hubuum_event_worker_wakeups_total worker, kind Notification, poll, and notifications-sent wakeups observed by this process
hubuum_inventory_entities entity_type Database-wide collections, classes, objects, users, groups, service accounts, and remote targets

Alert Starting Points

These thresholds are deployment starting points, not universal defaults:

Signal Suggested alert
Missing target up == 0, grouped by expected API and worker target
Missing worker telemetry Queued tasks with sum(hubuum_task_workers_configured) == 0, or no expected worker/all role target
Counter reset Unexpected resets(hubuum_http_requests_total[15m]), correlated with process start time
Process resource pressure CPU, resident memory, or open-FD/handle ratio above the target's established baseline
Stale snapshots Current time minus hubuum_metrics_refresh_last_success_timestamp_seconds exceeds the cache and scrape tolerance
DB acquisition failures Any sustained non-zero hubuum_db_connection_acquire_failures_total rate
DB pool pressure Checked-out divided by configured connections above 0.8 for several minutes
HTTP 5xx rate 5xx status family above the normal route-specific baseline
Task queue age Oldest queued task age above the expected latency for that task kind
Recent task failure time() - max by (kind) (hubuum_task_last_terminal_timestamp_seconds{status="failed"}) below the alert window
Task worker errors Sustained non-zero worker iteration outcome="error" rate
Task lease recovery Any unexpected hubuum_task_lease_recoveries_total increase
Export or import failures Failure or timeout outcomes above the task-kind baseline
Login lockouts Sudden increase in lockouts or sustained locked entries
Shared limiter degradation Sustained non-zero login-limiter backend failure rate
Remote call failures Failure or timeout rate above the remote-call baseline
Event backlog Oldest fan-out or delivery age above the processing objective

Schema compliance

Schema mutation counters use bounded policy/result labels. Validation object counters distinguish valid, invalid, not-required, uninspectable, and stale results; duration histograms track terminal work. The compliance gauge exposes global valid/invalid/pending/not-required counts. Generic task metrics include schema_validation backlog and recovery. Per-class IDs, object IDs, paths, and schema content never become metric labels; administrator schema reports provide per-class detail. See class schema evolution.

Operator package

See the operator package for seven shared Grafana dashboards, recording rules, tested alerts, SLO definitions and response runbooks. The optional single-host installer and distributed installations consume the same assets; a matching Prometheus Operator resource is included. Review this package in the same pull request whenever the metric contract changes.