Runtime Metrics¶
Hubuum exposes low-cardinality runtime metrics through a Prometheus scrape
endpoint. The endpoint is enabled by default at /metrics.
Configuration¶
| Variable | Default | Description |
|---|---|---|
HUBUUM_METRICS_ENABLED |
true |
Enables the Prometheus metrics scrape endpoint |
HUBUUM_METRICS_PATH |
/metrics |
Literal absolute non-root endpoint path; it must not contain route patterns or collide with API, probe, OpenAPI, or Swagger UI routes |
The endpoint is subject to HUBUUM_CLIENT_ALLOWLIST and uses the configured TLS
settings. Put it behind network-level access controls appropriate for
operational data.
Processes with the all or api runtime role serve metrics on the main HTTP
listener. A worker process serves only the configured metrics path on
HUBUUM_BIND_IP:HUBUUM_PORT when metrics are enabled; it does not expose the
application API or health probes. Give API and worker containers separate
network namespaces or ports if they run on the same host.
Scrape Each Process Directly¶
Counters, histograms, connection-pool gauges, and worker configuration describe one process. Configure Prometheus with a stable target for every API and worker process. Do not scrape a load-balanced public URL: successive scrapes can land on different processes, producing apparent counter resets and hiding worker-only activity.
Use these metrics to identify a target:
hubuum_build_info{version,git_sha}hubuum_runtime_info{role}hubuum_storage_backend_info{backend}hubuum_process_start_time_seconds
Check the target mix before interpreting process-local worker metrics:
An API-only target reports hubuum_runtime_info{role="api"} and zero configured
task and event workers. It can still expose database-wide task and template
gauges, but export, import, remote-call, task-execution, and worker-loop
counters and histograms are produced by worker-enabled targets. If those
families are absent from an API scrape, verify that Prometheus has a separate
worker target before treating the absence as no activity.
Inventory, task backlog, export-template identity, and event-queue gauges are
database-wide snapshots. Every replica connected to the same database reports
the same logical values. Aggregate those series with max or avg, not sum,
unless replica multiplication is intentional. Process-local gauges such as
database-pool state should normally be summed when calculating deployment-wide
capacity.
Database-backed gauges use a 30-second in-process cache and refresh on a
best-effort basis. If a refresh fails, /metrics still returns process metrics
and retains the last successful database snapshot when one exists. The refresh
duration, last-success timestamp, failures, and skipped concurrent refreshes
make stale values visible.
Single-host deployments provide deterministic convenience routes.
Shared-host direct routing sends /metrics to the worker-enabled primary and
/metrics/standby to the HTTP-only standby. With shared-host prefixed
routing, use /hubuum-api/metrics and
/hubuum-api/metrics/standby. Shared-host bff routing does not expose either
backend metrics endpoint publicly.
Cardinality Rules¶
Metric labels must stay bounded. Hubuum metrics do not use usernames, user IDs, client IPs, raw URL paths, object IDs, class names, collection names, rendered remote URLs, task IDs, idempotency keys, or error messages.
HTTP metrics use Actix route templates such as
/api/v1/classes/{class_id}/objects/{object_id}. Requests that do not resolve
to a registered route use a coarse route group instead of their raw path.
Export phase timing is aggregated by phase and outcome. Only the total export
histogram carries template_id; this keeps template identity useful without
multiplying every query, hydration, and render bucket by every stored template.
hubuum_export_template_info maps current database IDs to mutable names and is
reset on every inventory snapshot, so renames and deletions do not leave stale
series in a process.
Use admin JSON/API endpoints and task logs for per-task detail.
Duration Units And Histograms¶
Prometheus duration metric names and observations use seconds. A 4.7
millisecond request is recorded as 0.0047; using seconds does not discard
millisecond or sub-millisecond precision.
Hubuum uses three explicit bucket profiles:
- HTTP and database latency:
0.0005seconds through30seconds. - Outbound remote calls:
0.01seconds through120seconds. - Queued and background work:
0.01seconds through3600seconds.
The long background buckets are intentional: most buckets should be empty for fast work, while imports, exports, backups, remote dependencies, or a constrained worker can legitimately take seconds or minutes. Empty upper buckets are cheap and preserve visibility when an exceptional slow operation occurs.
Prometheus histogram buckets are cumulative. This query calculates HTTP p95 latency by route over five minutes:
histogram_quantile(
0.95,
sum by (le, route) (
rate(hubuum_http_request_duration_seconds_bucket[5m])
)
)
The histogram sum divided by its count gives the mean. Apply rate before
division for a time window:
sum by (route) (
rate(hubuum_http_request_duration_seconds_sum[5m])
)
/
sum by (route) (
rate(hubuum_http_request_duration_seconds_count[5m])
)
The exported _sum is therefore useful, but only together with _count.
Classic Prometheus histograms do not expose exact minimum or maximum
observations. histogram_quantile estimates percentiles from bucket boundaries
and remains aggregatable across processes.
Readiness probes commonly dominate request counts. Exclude probes and the
scrape endpoint when calculating application traffic or an overall application
latency. The example uses the default scrape route, /metrics. If
HUBUUM_METRICS_PATH is customized, replace /metrics in the matcher with the
exact configured path; for example, use
route!~"/(healthz|readyz)|/internal/metrics" for /internal/metrics.
Prometheus regex matchers are fully anchored:
Counter and histogram series generally appear only after a process observes a matching
event. The worker-error and failed-backup counters start at zero so alerts can
observe their first failure after scraping the baseline. Seeing only /readyz means that target handled readiness probes but not
application requests; it does not mean other routes are filtered out.
For stored-template exports, join the total-duration histogram to the current template-info gauge:
histogram_quantile(
0.95,
sum by (le, template_id) (
rate(
hubuum_export_duration_seconds_bucket{
template_id!="none"
}[15m]
)
)
)
* on (template_id) group_left (template_name)
max by (template_id, template_name) (
hubuum_export_template_info
)
Metrics¶
The complete metric inventory—including type, unit, process/database scope,
allowed label names, feature ownership, and explicit histogram buckets—is
generated in Metric Reference. The generated reference
and docs/operational-contract.json are the canonical compatibility sources;
CI rejects drift from the typed registry.
Process, HTTP, And Database¶
The generated Metric Reference replaces the former hand-maintained table for this group. The notes below explain the bounded domains and aggregation semantics that are not part of each metric's structural signature.
The bounded refresh source values are database, events, inventory,
login_limiter, process, tasks, and token_keys.
The bounded database caller values are event_delivery, event_fanout,
event_retention, http_request, metrics_refresh, readiness,
request_maintenance, restore_coordinator, task_lease, task_worker,
token_retention, and unattributed.
Storage capability, operation, and result values come from closed Rust
enums or fixed call sites. Expected domain failures are included in
hubuum_storage_operation_errors_total; use the result label to distinguish
them from database, unavailable, and internal backend failures. Storage
metrics describe logical use cases, while hubuum_db_* metrics describe the
PostgreSQL implementation. Do not sum the two as though they were the same
operation count.
Distributed Tracing¶
| Metric | Labels | Description |
|---|---|---|
hubuum_tracing_info |
sampling_mode |
Effective bounded sampling mode; value is 1 |
hubuum_tracing_sample_ratio |
none | Effective root trace sampling ratio |
hubuum_tracing_queue_capacity |
none | Configured bounded batch-processor queue capacity |
hubuum_tracing_queue_utilization |
none | Current spans admitted and waiting for export |
hubuum_trace_spans_total |
category, state |
Sampled allowlisted spans observed at the processor's actual started or ended lifecycle callback |
hubuum_trace_spans_dropped_total |
reason |
Spans dropped for classification or queue_saturation |
hubuum_trace_export_batches_total |
outcome |
OTLP export batches by success or failure |
hubuum_trace_export_spans_total |
outcome |
Spans submitted in OTLP batches by success or failure |
hubuum_trace_flushes_total |
outcome |
Graceful-shutdown flushes by success or failure |
Trace labels come only from fixed categories and outcomes. They never contain
trace IDs, span IDs, endpoints, resource identities, or error messages. Alert on
a sustained nonzero queue_saturation rate and export failure rate. Queue
utilization is process-local and should be compared with that process's capacity.
See Distributed Tracing for configuration and data policy.
Storage capability labels follow the singular capability trait vocabulary: the
trait's Storage suffix is removed and its remaining stem becomes snake case.
The bounded values are audit_event, authentication, authorization_data,
backup_snapshot, catalog, class, class_relation,
collection_authorization_query, collection, computed_field,
computed_object, event_configuration,
event_delivery_administration, event_delivery_worker, event_fanout,
event_health, event_retention, export_template, external_identity,
group, history, group_membership, identity_scope, import, inventory,
local_identity_credential, metrics, object_aggregate, object_relation,
object, operational_state, principal, relation_query, remote_target,
restore, service_account, task_execution, task_queue, token_retention,
token, transaction, unified_search, and user.
The standard unprefixed process_* families are available on Linux, macOS, and
Windows. The names intentionally match the Prometheus ecosystem so existing
process dashboards and alerts can be reused. File-descriptor or handle gauges
remain zero if the operating system cannot return that information. On Windows,
process_open_fds counts kernel handles and process_max_fds is zero because
Windows has no comparable fixed per-process handle limit. Prefer resident memory
for cross-platform pressure alerts: virtual address-space size is especially
large and not directly comparable on macOS.
process_start_time_seconds is the operating system's process start, while
hubuum_process_start_time_seconds records when Hubuum initialized metrics.
Useful starting queries include:
Tasks, Exports, Imports, And Remote Calls¶
| Metric | Labels | Description |
|---|---|---|
hubuum_task_worker_iterations_total |
outcome |
Worker iterations by claimed, idle, or error outcome |
hubuum_task_claims_total |
kind |
Tasks claimed by workers |
hubuum_task_lease_recoveries_total |
kind |
Tasks failed after their owning worker lease expired |
hubuum_task_completions_total |
kind, final_status |
Tasks reaching a terminal status |
hubuum_task_cancellation_requests_total |
kind, result |
Durable requests with queued, active, or unchanged result |
hubuum_task_stop_acknowledgements_total |
kind, reason |
Terminal stops with cancel_requested or deadline_exceeded reason |
hubuum_task_cancellation_acknowledgement_duration_seconds |
kind, reason |
Time from persisted cancellation request to terminal acknowledgement |
hubuum_task_ambiguous_remote_stops_total |
reason |
Stops after remote dispatch where external effects may have occurred |
hubuum_task_queue_wait_duration_seconds |
kind |
Time from task creation to claim |
hubuum_task_execution_duration_seconds |
kind, final_status |
Time from task start to finish |
hubuum_task_workers_configured |
none | Task workers configured in this process; zero on API-only processes |
hubuum_task_poll_interval_seconds |
none | Configured task-worker poll interval |
hubuum_tasks |
kind, status |
Current database-wide task counts |
hubuum_task_oldest_age_seconds |
kind, state |
Oldest queued and active task age per task kind |
hubuum_task_last_terminal_timestamp_seconds |
kind, status |
Most recent finish time for each task kind and terminal status; zero means none is retained |
hubuum_task_output_cleanup_runs_total |
kind |
Stored output cleanup runs for export or backup artifacts |
hubuum_task_output_cleanup_failures_total |
kind |
Stored output cleanup failures |
hubuum_task_output_cleanup_deleted_total |
kind |
Stored outputs deleted by cleanup |
hubuum_export_template_info |
template_id, template_name |
Current stored export-template identities from the shared database |
hubuum_export_phase_duration_seconds |
phase, outcome |
Aggregate export query, hydration, render, and total phase duration |
hubuum_export_duration_seconds |
template_id, outcome |
Total export duration by stored template ID; ad-hoc exports use none |
hubuum_export_completions_total |
scope, content_type |
Successfully persisted export outputs |
hubuum_export_truncations_total |
scope, content_type |
Successfully persisted truncated exports |
hubuum_export_warnings_total |
scope, content_type |
Warning count on successfully persisted exports |
hubuum_import_phase_duration_seconds |
phase, outcome |
Import planning, execution, and total phase duration, including failures |
hubuum_import_processed_items_total |
none | Items processed by terminal import tasks |
hubuum_import_succeeded_items_total |
none | Import items completed successfully |
hubuum_import_failed_items_total |
none | Import items completed with failure |
hubuum_remote_call_duration_seconds |
method, status_family, outcome |
Remote HTTP execution duration |
hubuum_remote_call_results_total |
method, status_family, outcome |
Remote outcomes such as success, failure, timeout, or validation rejection |
Export timer phases are limited to total, query, hydration, and render;
their outcomes are success, error, or timeout. Import timer phases are
limited to total, planning, and execution; their outcomes are success,
failed, partially_succeeded, or error.
Computed Fields, Security, Events, And Inventory¶
| Metric | Labels | Description |
|---|---|---|
hubuum_computed_field_evaluations_total |
scope, outcome |
Computed-field evaluations by shared, personal, or preview scope and outcome |
hubuum_computed_field_errors_total |
scope, code |
Computed-field runtime errors by stable bounded code |
hubuum_computed_field_live_fallbacks_total |
none | Stale shared materializations evaluated live during reads |
hubuum_computed_field_read_repairs_total |
outcome |
Guarded stale-materialization repairs by success or failure |
hubuum_computed_field_rebuild_batches_total |
items |
Computed-field rebuild batches classified as empty or non-empty |
hubuum_computed_field_rebuild_completions_total |
status |
Computed-field rebuild terminal outcomes |
hubuum_computed_field_rebuild_duration_seconds |
status |
Computed-field rebuild duration histogram |
hubuum_login_attempts_total |
outcome |
Login attempts by success, bad credentials, rate-limited, or internal error |
hubuum_login_lockouts_total |
scope |
Login limiter lockout transitions by principal/IP, IP, or subnet scope |
hubuum_login_limiter_backend_failures_total |
backend, operation |
Shared login-limiter failures while local enforcement remains active |
hubuum_login_limiter_entries |
state |
Active and locked login-limiter entries in this process |
hubuum_client_allowlist_rejections_total |
reason |
Requests rejected for a disallowed or missing client IP |
hubuum_revision_conditions_total |
outcome |
Conditional writes classified as matched, wildcard, stale, unconditional, malformed, async_stale, or invariant_failure; never labelled by resource identity or revision |
hubuum_event_queue_items |
queue, state |
Database-wide fan-out and delivery queue items by bounded state |
hubuum_event_stale_claims |
queue |
Stale fan-out and delivery worker claims |
hubuum_event_oldest_age_seconds |
queue |
Oldest actionable fan-out or delivery item age |
hubuum_event_workers_configured |
worker |
Event workers configured in this process |
hubuum_event_worker_batch_size |
worker |
Configured event-worker batch size |
hubuum_event_worker_poll_interval_seconds |
worker |
Configured event-worker poll interval |
hubuum_event_worker_lock_timeout_seconds |
worker |
Configured event-worker claim lock timeout |
hubuum_event_worker_wakeups_total |
worker, kind |
Notification, poll, and notifications-sent wakeups observed by this process |
hubuum_inventory_entities |
entity_type |
Database-wide collections, classes, objects, users, groups, service accounts, and remote targets |
Alert Starting Points¶
These thresholds are deployment starting points, not universal defaults:
| Signal | Suggested alert |
|---|---|
| Missing target | up == 0, grouped by expected API and worker target |
| Missing worker telemetry | Queued tasks with sum(hubuum_task_workers_configured) == 0, or no expected worker/all role target |
| Counter reset | Unexpected resets(hubuum_http_requests_total[15m]), correlated with process start time |
| Process resource pressure | CPU, resident memory, or open-FD/handle ratio above the target's established baseline |
| Stale snapshots | Current time minus hubuum_metrics_refresh_last_success_timestamp_seconds exceeds the cache and scrape tolerance |
| DB acquisition failures | Any sustained non-zero hubuum_db_connection_acquire_failures_total rate |
| DB pool pressure | Checked-out divided by configured connections above 0.8 for several minutes |
| HTTP 5xx rate | 5xx status family above the normal route-specific baseline |
| Task queue age | Oldest queued task age above the expected latency for that task kind |
| Recent task failure | time() - max by (kind) (hubuum_task_last_terminal_timestamp_seconds{status="failed"}) below the alert window |
| Task worker errors | Sustained non-zero worker iteration outcome="error" rate |
| Task lease recovery | Any unexpected hubuum_task_lease_recoveries_total increase |
| Export or import failures | Failure or timeout outcomes above the task-kind baseline |
| Login lockouts | Sudden increase in lockouts or sustained locked entries |
| Shared limiter degradation | Sustained non-zero login-limiter backend failure rate |
| Remote call failures | Failure or timeout rate above the remote-call baseline |
| Event backlog | Oldest fan-out or delivery age above the processing objective |
Schema compliance¶
Schema mutation counters use bounded policy/result labels. Validation object
counters distinguish valid, invalid, not-required, uninspectable, and stale
results; duration histograms track terminal work. The compliance gauge exposes
global valid/invalid/pending/not-required counts. Generic task metrics include
schema_validation backlog and recovery. Per-class IDs, object IDs, paths, and
schema content never become metric labels; administrator schema reports provide
per-class detail. See class schema evolution.
Operator package¶
See the operator package for seven shared Grafana dashboards, recording rules, tested alerts, SLO definitions and response runbooks. The optional single-host installer and distributed installations consume the same assets; a matching Prometheus Operator resource is included. Review this package in the same pull request whenever the metric contract changes.