Distributed Deployment¶
Hubuum supports multiple API and background-worker replicas backed by one PostgreSQL database. Valkey or Redis is optional and adds shared login-throttle state; the default in-memory limiter remains fully functional for deployments that do not configure it.
clients -> load balancer -> API replicas ----+
+-> PostgreSQL
worker replicas ----+
+-> Valkey/Redis (optional)
Runtime Roles¶
Set HUBUUM_RUNTIME_ROLE on each long-running process:
| Value | HTTP API | Background workers | Intended use |
|---|---|---|---|
all |
Yes | Yes | Default; preserves the single-process behavior |
api |
Yes | No | Horizontally scaled stateless API replicas |
worker |
No | Yes | Independently scaled task and event workers |
An api process cannot start workers through a later queue notification. When
metrics are enabled, a worker process binds the configured address and port
with only HUBUUM_METRICS_PATH; it does not expose the application API or
health probes. Scrape every worker directly so worker counters and histograms
are not hidden behind an API load balancer. Use process/container liveness for
worker replicas rather than /healthz; the image's built-in Docker health check
accounts for the worker role. API replicas expose /healthz and /readyz as
documented in Quick Start. See
Runtime Metrics for aggregation
semantics.
Worker-only processes require at least one configured background worker and
supervise every worker thread they start. If any worker stops unexpectedly, the
server process exits so the container liveness check fails and the orchestrator
can replace the replica. The all role applies the same supervision whenever it
starts background workers.
The all role is the default, so existing single-instance deployments continue
to behave as before.
One-Shot Migrations¶
Every server container role checks schema readiness without applying
migrations. Run exactly one migration job before rolling out the new
application version. The production image contains hubuum-admin and embedded
migrations; it does not require the Diesel CLI or psql. See
PostgreSQL Database Roles for initial provisioning and the
complete least-privilege contract.
With the default HUBUUM_DATABASE_ROLE_MODE=single, give the Job
HUBUUM_DATABASE_URL, using the same Secret as the application. The example
below shows the opt-in split topology; set the mode on every workload and keep
the migration Secret out of API and worker Deployments.
apiVersion: batch/v1
kind: Job
metadata:
name: hubuum-migrate
spec:
template:
spec:
restartPolicy: OnFailure
containers:
- name: migrate
image: ghcr.io/hubuum/hubuum-server:VERSION
command: ["/usr/local/bin/hubuum-admin", "--migrate"]
env:
- name: HUBUUM_DATABASE_ROLE_MODE
value: split
- name: HUBUUM_MIGRATION_DATABASE_URL
valueFrom:
secretKeyRef:
name: hubuum-migration-database
key: database-url
Wait for the job to complete successfully before updating API or worker replicas. For the task-lease and task-provenance migrations, use this upgrade order:
- Stop old-version worker replicas, allowing their bounded graceful shutdown to finish or fail active tasks.
- Run the one-shot migration.
- Deploy new worker and API replicas.
For the lease migration, the drain prevents a new worker from treating a task
owned by an old, lease-unaware worker as abandoned. For the provenance
migration, it prevents an old worker from committing temporal mutations
without the root task context that only the new worker propagates. Old API
replicas may remain online during the migration: the expanded schema derives
task initiators from submitted_by and treats their legacy hubuum.actor_id
session setting as direct-user attribution.
For the resource-revision migration, stop old worker replicas before running the migration, but an adjacent old API replica may continue serving ordinary requests. CI certifies old and candidate API reads and writes against the migrated schema before an application rollback smoke test. Backup v3 and import v1 are still deliberately rejected by the candidate, so quiesce backup, restore, and import operations until every replica is on the new version.
The api and worker entrypoints wait until the database records the latest
migration required by the binary. API /readyz performs the same schema check.
This prevents a missed or incomplete migration job from making a replica appear
ready, while keeping migration ownership in the one-shot job.
When adopting split database roles, first verify maintenance is normal and block
POST /api/v1/restores/*/confirm at the ingress. The migration shares the
restore advisory lock and refuses to run if a confirmed restore is draining;
validated stages are preserved. After the Job succeeds, deploy and verify the
restore executor before rolling API and worker replicas or unblocking that
route. Existing single-role databases should follow the complete adoption and
rollback sequence in
PostgreSQL Database Roles.
Isolated Restore Executor¶
API confirmation queues a restore but never performs privileged SQL. Run one long-lived executor replica from the same image:
apiVersion: apps/v1
kind: Deployment
metadata:
name: hubuum-restore-executor
spec:
replicas: 1
selector:
matchLabels:
app: hubuum-restore-executor
template:
metadata:
labels:
app: hubuum-restore-executor
spec:
containers:
- name: restore-executor
image: ghcr.io/hubuum/hubuum-server:VERSION
command: ["/usr/local/bin/hubuum-admin", "--restore-executor"]
env:
- name: HUBUUM_DATABASE_ROLE_MODE
value: split
- name: HUBUUM_MIGRATION_DATABASE_URL
valueFrom:
secretKeyRef:
name: hubuum-migration-database
key: database-url
securityContext:
allowPrivilegeEscalation: false
readOnlyRootFilesystem: true
capabilities:
drop: ["ALL"]
Do not add a Service. In split mode, do not mount the migration secret into API or worker pods. The executor polls database control state, revalidates the staged document, and uses the configured single or migrator identity only for the closed, transaction-protected restore operation.
Kubernetes And Helm HTTP Availability¶
This repository does not currently publish a Helm chart. Direct Kubernetes manifests and externally maintained charts should implement the following contract for application upgrades without HTTP downtime:
- Run API and worker processes in separate Deployments with the
apiandworkerruntime roles. Do not putall-role pods behind the API Service. - Run at least two API replicas and use a
RollingUpdatestrategy withmaxUnavailable: 0andmaxSurge: 1or greater. The default 25% values do not state the availability requirement clearly and can behave differently as the replica count changes. - Use
/readyzfor readiness and/healthzfor startup and liveness. Only ready API pods should receive Service or ingress traffic. - Allow endpoint and ingress changes to propagate before terminating the
process. A short
preStopdelay is a practical starting point, and the pod's termination grace period must cover that delay plus Hubuum's 30-second graceful worker-shutdown bound. Start with 60 seconds and tune the drain delay for the actual ingress implementation. - Add a PodDisruptionBudget that keeps at least one API replica available, and spread replicas across nodes or failure domains. A disruption budget protects against voluntary evictions; the Deployment strategy still controls rolling updates.
The relevant API Deployment settings look like this. This is a fragment rather than a complete manifest:
spec:
replicas: 2
minReadySeconds: 5
strategy:
type: RollingUpdate
rollingUpdate:
maxUnavailable: 0
maxSurge: 1
template:
spec:
terminationGracePeriodSeconds: 60
containers:
- name: hubuum-api
env:
- name: HUBUUM_RUNTIME_ROLE
value: api
- name: HUBUUM_DATABASE_URL
valueFrom:
secretKeyRef:
name: hubuum-runtime-database
key: database-url
- name: HUBUUM_DATABASE_PRIVILEGE_MODE
value: strict
ports:
- name: http
containerPort: 8080
startupProbe:
httpGet:
path: /healthz
port: http
periodSeconds: 5
failureThreshold: 30
readinessProbe:
httpGet:
path: /readyz
port: http
periodSeconds: 5
failureThreshold: 1
livenessProbe:
httpGet:
path: /healthz
port: http
periodSeconds: 10
failureThreshold: 3
lifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "sleep 5"]
Before selecting a rollout strategy, run the candidate's
hubuum-admin --migration-mode with its migrator credentials. offline requires
stopping every old API, worker, and restore executor and taking a PostgreSQL
snapshot before migration. This applies to 0.0.16 to 0.0.17. Start only matching
candidate processes afterward. Recovery requires restoring that snapshot before
starting old processes; binary-only rollback is unsupported. Writes after the
snapshot are lost. Keep old binaries, credentials, and the snapshot until the
upgrade is accepted. A failed preflight must block deployment.
For a rolling migration plan, make the one-shot migration Job a blocking pre-install and
pre-upgrade hook, or run and await an equivalent uniquely named Job in the
release pipeline before helm upgrade:
metadata:
annotations:
"helm.sh/hook": pre-install,pre-upgrade
"helm.sh/hook-weight": "-5"
"helm.sh/hook-delete-policy": before-hook-creation,hook-succeeded
The migration Job must use the new application image. In single mode it may use the application database Secret. In split mode it must use a distinct migrator credential that is not mounted into any application Deployment. If the chart creates that credential, make it available before the hook runs or keep migration ownership in the release pipeline. A failed migration must stop the rollout while the old API replicas remain online.
Every migration used in this sequence must be compatible with both the old and new API versions. Use expand/backfill/switch/contract changes across releases; do not drop or reinterpret schema that the old pods still use. Helm rollback does not undo a completed database migration, so the previous application image must also remain compatible with the migrated schema. The task-lease and task-provenance migrations still require the worker drain order described above.
These settings protect the HTTP application rollout. They do not make the ingress controller, Kubernetes nodes, PostgreSQL, or Valkey highly available; those layers need their own redundancy and maintenance procedures.
Shared Configuration And Secrets¶
Environment-backed credentials are the default. For mounted credentials, set
HUBUUM_SECRET_SOURCE=file and HUBUUM_SECRET_FILE_ROOT on each workload, or
pass --secret-source file --secret-file-root DIRECTORY to either binary.
The secret-source mapping and precedence
apply to API replicas, workers, one-shot migrations, and restore executors.
In single-role mode those workloads can share database/url; in split mode
mount database/migration-url only into migration and restore workloads.
See split-role mounts for the layout.
All replicas must use the same values for settings that define cluster-wide identity or behavior. In particular:
- The effective database URL (
HUBUUM_DATABASE_URL,database/url, or an explicit--database-url) must point to the same PostgreSQL database. - The token hash key ring (active ID, ordered previous IDs, and every key's material) must be equivalent on all replicas except during the documented staged rotation boundary. Compare the redacted ring identity in runtime configuration before advancing a rollout.
- Authentication-provider configuration and referenced credentials must be equivalent on every API and worker replica.
- Task lease and event lock durations should be consistent across worker replicas.
- Client-visible pagination defaults and maxima must be consistent across API replicas.
- Reverse-proxy trust and client-allowlist settings should be consistent across API replicas.
The effective non-secret configuration is available through the existing running-configuration endpoint. Secret fields report only whether they are configured. Its database section also reports the selected complete storage backend and effective pool settings. Every replica in one deployment must report the same backend and compatible settings.
Use the token key-ring rotation procedure for multi-replica changes. Changing a single key in place is not a safe online rotation.
Shared Login Throttling¶
The default backend is local memory:
This has the same behavior as earlier releases and needs no external service. Each API replica enforces its own limits.
To coordinate attempts and lockouts across API replicas, configure Valkey or Redis:
HUBUUM_LOGIN_RATE_LIMIT_BACKEND=valkey
HUBUUM_LOGIN_RATE_LIMIT_VALKEY_URL=rediss://valkey.example:6379/0
HUBUUM_LOGIN_RATE_LIMIT_VALKEY_PREFIX=hubuum:login-rate-limit
HUBUUM_LOGIN_RATE_LIMIT_VALKEY_IO_TIMEOUT_MS=1000
The local limiter remains active as a safety net. If the shared backend is
unavailable, password logins continue under per-instance limits, a warning is
logged on the transition, and
hubuum_login_limiter_backend_failures_total{backend="valkey",operation="..."}
increments. Shared enforcement resumes automatically after Valkey recovers.
Existing bearer-token authentication never depends on Valkey.
Limiter administration remains strict during an outage: shared list, release, and clear operations return an error when they cannot operate on the canonical shared state. See Login Rate Limiting.
Task Ownership And Recovery¶
PostgreSQL remains the task queue. Claims use FOR UPDATE SKIP LOCKED, and each
claimed task now carries a durable random lease token and expiry. The owning
worker renews the lease while executing. State updates and terminal writes are
fenced by that token, so a stale worker cannot overwrite recovery performed by
another replica.
Configure the lease with:
| Variable | Default | Description |
|---|---|---|
HUBUUM_TASK_LEASE_SECONDS |
60 |
Lease duration for an active task |
HUBUUM_TASK_HEARTBEAT_SECONDS |
20 |
Renewal interval; must be shorter than the lease |
HUBUUM_TASK_RECOVERY_INTERVAL_SECONDS |
30 |
Minimum interval between recovery scans in one process |
HUBUUM_COMPUTED_REINDEX_BATCH_SIZE |
100 |
Objects processed per computed-field rebuild transaction; valid range is 1 through 1000 |
An expired task is failed, its request payload is redacted, and a system event records the prior state and recovery reason. Hubuum does not automatically replay abandoned tasks because imports and remote calls can have external side effects. Inspect the task history and submit a new task only when replay is known to be safe.
Capacity Planning¶
Each process owns its main database pool. A worker or all process with
HUBUUM_TASK_WORKERS > 0 also lazily creates a dedicated one-connection task
lease pool so lease renewals cannot be starved by task execution. Start with
this upper-bound budget:
connections = (api replicas + worker replicas + all-role replicas)
* HUBUUM_DB_POOL_SIZE
+ (worker replicas with task workers
+ all-role replicas with task workers)
* 1 task lease connection
+ migration/administration connections
+ operational headroom
Keep the result below PostgreSQL's connection limit and leave capacity for maintenance. API and worker replicas may use different pool sizes even though they share the same setting name in their respective process configuration. See Performance for measurement and pool metrics.
Scale task workers by increasing worker replicas or HUBUUM_TASK_WORKERS.
Correctness does not depend on optimizing the active-task capacity count query;
that query's higher-scale optimization remains follow-up work in
issue #67.
Storage And Networking Checklist¶
- Put a load balancer only in front of
apireplicas. - Configure
HUBUUM_TRUSTED_PROXIESfor the actual load-balancer networks before enabling forwarded-IP headers. - Do not use local event archive files across replicas. Keep retention archive disabled or deliver audit events to shared external storage.
- Mount identical TLS, CA, and authentication configuration where those files are used.
- Keep clocks synchronized for logs and external integrations. Valkey supplies the shared limiter clock, and PostgreSQL supplies the task-lease clock.
- Use graceful termination and allow at least the worker shutdown timeout before forcibly killing a pod.
For a single-host installation, continue to use Single-Host Container Deployment.
Operator monitoring¶
Use the shared operator package for Grafana dashboards, Prometheus recording and alerting rules, SLO definitions and response runbooks. The same assets work with the optional single-host stack, independently managed Prometheus/Grafana installations, and Prometheus Operator. Pin the package to your server release and scrape every process directly with deployment labels.