Buttons can run as a replicated, self-healing cluster on Kubernetes: every service triple-replicated across three nodes, with its own Postgres and Redis high-availability layer included. This is a separate, advanced deployment path from a single-host install, and requires an Enterprise license.
An Enterprise license: with HA_ENABLED set, any license below Enterprise is treated as no license at all, not merely capped to a lower tier.
A Kubernetes cluster with:
A storage class capable of ReadWriteMany, for the shared modules store.
An ingress controller: Traefik is what this deployment's manifests actually use (included by default with k3s).
Optionally, cert-manager for automatic certificate issuance.
The repository's kubernetes/ manifest set, checked out locally.
Note
As shipped, this deployment is built around exactly three nodes and three replicas per service; every StatefulSet's replica count is fixed at 3, and node-affinity rules and the cluster's own leader-election math are written around that number. Scaling beyond three nodes isn't a configuration this deployment currently supports out of the box. See Understand HA clustering for why three is the number.
Every Buttons service runs as a Kubernetes StatefulSet with 3 replicas. Most of them (chore, connection, discovery, mdns-relay, nmos, orchestrator, surface, tally, workflow, xy, and the watchdog itself) participate in the cluster's own internal leader election, so that exactly one replica is actively "in charge" at a time even though three are running. Three services don't need this: the render engine (bcsdd), the USB Relay bridge, and the web frontend (www) are all safely active on every replica at once.
Postgres and Redis are deployed inside the cluster, not as externally managed services you bring yourself:
PostgreSQL runs as a 3-instance CloudNativePG cluster with synchronous replication and automatic failover, fronted by a PgBouncer connection pooler on every node.
Redis runs as a 3-replica Valkey (Redis-compatible) set, with Sentinel managing automatic failover between primary and replicas.
Only two services need real persistent storage beyond ephemeral scratch space: the connection service and the www service both mount a single shared, ReadWriteMany-capable volume for the module cache. Postgres, Redis, and Redis Sentinel each have their own dedicated volume. Everything else uses ephemeral storage that doesn't need to survive a pod restart.
Warning
In this deployment's shipped configuration, the chore service (which is what actually writes scheduled backups) has no persistent volume at all; by default, a scheduled backup lands in ephemeral storage that's lost if that pod restarts or is rescheduled. If you rely on Configure and monitor scheduled backups in a Kubernetes deployment, mount a persistent volume into the chore service yourself and point its backup path at it. Don't assume backups are durable here without doing that.
Separately, the underlying storage layer can take its own periodic snapshots of the database and module-cache volumes. This is a lower-level safety net for the volumes themselves, and isn't a substitute for the application's own Configure and monitor scheduled backups feature described above.
Only the web frontend is exposed outside the cluster through the ingress: every other service stays internal-only, with one exception, covered below. TLS terminates at the ingress, using the same certificate-management flow already covered in Replace the HTTPS certificate: when Kubernetes is managing the certificate (for example, via cert-manager), Buttons' own Settings → Certificates page shows "Certificate managed by Kubernetes," and the certificate is delivered to the ingress the same way regardless of which service happens to be handling a given request.
The ingress accepts both plain HTTP (port 80) and HTTPS (port 443) by default, routed to the same backend: reaching Buttons over http:// works out of the box, with no automatic redirect to HTTPS.
Note
HSTS and an HTTP→HTTPS redirect are both included in the shipped manifests but commented out by default. To turn them on, uncomment the hsts and redirect-https middleware references in kubernetes/buttons/ingress/ingress.yaml and reapply it; enable both together, since HSTS alone doesn't stop port 80 from serving requests unencrypted. Once enabled, browsers that visit over HTTPS cache the HSTS setting and will only connect to Buttons over HTTPS afterward, so confirm your certificate is working correctly before turning it on.
Tally pods aren't hostNetwork, and their own Service exposes only the internal leader-election ports, so a Tally/UMD Server (Listen) connection (see Send or receive tally and labels over TSL/UMD) has nothing to be reachable through by default.
To fix this, the tally app's leader pushes one watchdog port forward per enabled TSL Server connection, each creating its own dedicated LoadBalancer Service (app-tally-pf-<connection id>) outside the ingress, exposing exactly that connection's configured port and transport. Plan a facility firewall rule for each such Service's external address the same way you would for the ingress. This Service follows the elected leader: on failover, the watchdog retargets its selector to the new leader pod, and the Service itself is never deleted on step-down, to avoid flapping the external load balancer during the handover.
A TSL Server connection can't be set to listen on 3131 (reserved for the health check) or 7110–7112 (reserved for tally leader election): Buttons rejects the port with an explicit "reserved by Buttons for..." error before you can save it.
This applies the infrastructure layer (storage classes, ingress, the database operator) first, then the Buttons application manifests, in the order the manifests are meant to be applied.
There's no separate migration job that runs and completes before the rest of the cluster starts: migrations run inside the chore service itself, guarded by a database lock so that only one of its three replicas actually performs them while the others wait. Other services don't wait on migrations finishing before they start. Plan upgrades with this in mind: a rolling update doesn't guarantee migrations are already complete before every service comes back up.
Kubernetes' own liveness and readiness probes govern low-level pod lifecycle (whether a container gets restarted, and whether it receives traffic) independently of anything in Buttons' own interface. Buttons' own System Health banner and per-service status, covered in Interpret system health and restart services safely, is a separate, richer signal layered on top: it factors in real pod readiness alongside the cluster's own leader-election state, so treat the two as related but not identical: a service can be "ready" at the Kubernetes level while Buttons' own view of the cluster still shows something worth attention.
For planned node maintenance, use Cordon & Drain from the Services page (also covered in that same guide) rather than manually stopping pods: it's the action designed for taking a node offline while keeping the cluster running on the remaining two.
Scheduled backups don't survive a chore pod restart.
This is expected in the default configuration: mount a persistent volume into the chore service and point its backup path there.
A rolling upgrade seems to run before migrations finish.
Expected: migrations run inside chore under a database lock, and other services don't wait for them. Give an upgrade a moment to fully settle before assuming something's wrong.
Buttons shows no active license, or license features are missing, in an HA deployment.
Confirm you actually have an Enterprise license: any lower tier is treated as no license at all once HA mode is on, not a reduced one.
You want to check individual pod status directly.
kubectl -n buttons get pods, and kubectl -n buttons logs -f statefulset/<service-name> for a specific service's logs.
You're considering more than three nodes.
This deployment's replica counts, node-affinity rules, and leader-election quorum math are all built around exactly three: treat a different node count as unsupported rather than just changing the replica field.
You're wondering why it's three nodes and not one or two.
See Understand HA clustering: quorum needs a real majority, and three is the smallest number that tolerates losing a node without any ambiguity over which replica is in charge.