Bitfocus AS
logo
logo
Bitfocus AS
logo
logo
Sign upSign in

Loading...

Bitfocus

Subscribe to our newsletter

The latest news, articles, and resources, sent to your inbox.

FacebookInstagramGitHubYouTubeLinkedIn

Products

  • Buttons
  • Companion

Integrations

  • Supported Devices
  • Developer Community
  • Connection Development

Support

  • Support Overview
  • Documentation
  • Video Tutorials
  • Community Forum

Sales

  • Resellers & Integrators
  • Buttons Pricing

Updates

  • Case Studies
  • Events & Trade Shows
  • Press Releases
  • Product Updates
  • Webinars

Legal

  • Legal Overview
  • Privacy Policy
  • Buttons EULA
  • Terms & Cookie Policy

Company

  • About us
  • Press kit
  • Careers

© 2026 Bitfocus AS. All rights reserved.

Kubernetes HA
Docs for
Overview
Getting started
What is Bitfocus Buttons?
Install Buttons and get started
Manage your Buttons license
Activate Buttons offline
Find your way around Buttons
Create your first backup
Add an ATEM connection
Choose a control method
Choose an installation path
Install Buttons on Debian or Ubuntu
Understand HA clustering
Kubernetes HA
Update or remove Buttons
Positions
Understand positions
Create a position
Add controls and sections to a position
Create your first button
Use a connection's presets
Build more capable button actions
Add more feedback to a button
Organize controls in a section
Shift Section
Organize controls with a Folder Section
Add a Popover Section
Build and reuse a Shared Section
Build a Router Section
Understand Custom Routers
Custom Router panel
Surfaces
Surface compatibility
Add and attach a surface
Device orientation
Connections
Update a connection's module safely
Monitor and troubleshoot a connection
Router integrations
VideoHub and AJA KUMO
Utah Scientific BPS
Generic SW-P-08
Nevion VideoIPath
Arkona BLADE//runner
Routing
Physical routing
Configure ports and labels
Take a physical route
Understand route status
Topology graph
Routing Presets
Get started with virtual routing
Configure Nested Shapes
Reverse routing
Tielines
Routing Projects
Routing settings
Troubleshoot a route
Tally
Understand the Tally system
Send ATEM tally and labels to a UMD
Interpret Active Tally state
TSL/UMD connections
Diagnose tally problems
NMOS
Understand NMOS in Buttons
Connect Buttons to an NMOS Registry
Built-in Registry Server
Configure NMOS connections
Discover and adopt
Browse the NMOS inventory
Manage NMOS multicast addresses
Diagnose NMOS problems
Understand Cuelists
Build a Cuelist
Read and advance a running Cuelist
Control a Cuelist from a Position
Workflows
Understand workflows
Build your first workflow
Reuse a group of workflow nodes safely
Troubleshoot a workflow
Recipes
Sequence a timed automation
Call an HTTP endpoint from a workflow
REST endpoint
Use variables
Understand variable scope
Understand nested variables
Update expressions for v1.8
Plan and use Tags
Access
Create and manage users
Create roles and assign permissions
Grant access to specific resources
Show different controls by role
Sessions
Set up PIN and NFC sign-in
SSO
Get started with SSO
Connect a generic OIDC provider
Connect LDAP or Active Directory
Map identity claims to roles
Secure a Buttons deployment
Integrations
External control
Connect to Bitfocus Listener
USB Relay
Install USB Relay on Windows
Install USB Relay on macOS
Install USB Relay on Linux
Install USB Relay on a Raspberry Pi
Get started with the Control API
Secure and monitor the Control API
Control API reference
API reference
Administration
Enable and manage installable features
Services and health
Configure and monitor scheduled backups
Restore a backup and verify it
Export or import Buttons configuration
Store and rotate connection secrets
Replace the HTTPS certificate
HA backup and recovery
Settings
Collect support information
Reference
Glossary
Button Inspector reference
Network ports reference
Expressions
Internal actions reference
Routing Presets panel reference
Startup configuration reference
Workflow nodes
Connection workflow nodes
Workflow workflow nodes
Internal workflow nodes
Position workflow nodes
API workflow nodes
Utility workflow nodes

Loading...

Previous
← Understand HA clustering
Next
Update or remove Buttons →
Contact support →
You are viewing documentation for Buttons 1.8.See the docs for Buttons 1.6
Buttons/Installation/Kubernetes HA

Kubernetes HA

Buttons can run as a replicated, self-healing cluster on Kubernetes: every service triple-replicated across three nodes, with its own Postgres and Redis high-availability layer included. This is a separate, advanced deployment path from a single-host install, and requires an Enterprise license.

Before you begin#

  • An Enterprise license: with HA_ENABLED set, any license below Enterprise is treated as no license at all, not merely capped to a lower tier.
  • A Kubernetes cluster with:
  • A storage class capable of ReadWriteMany, for the shared modules store.
  • An ingress controller: Traefik is what this deployment's manifests actually use (included by default with k3s).
  • Optionally, cert-manager for automatic certificate issuance.
  • The repository's kubernetes/ manifest set, checked out locally.

Note

As shipped, this deployment is built around exactly three nodes and three replicas per service; every StatefulSet's replica count is fixed at 3, and node-affinity rules and the cluster's own leader-election math are written around that number. Scaling beyond three nodes isn't a configuration this deployment currently supports out of the box. See Understand HA clustering for why three is the number.

Understand the topology#

Every Buttons service runs as a Kubernetes StatefulSet with 3 replicas. Most of them (chore, connection, discovery, mdns-relay, nmos, orchestrator, surface, tally, workflow, xy, and the watchdog itself) participate in the cluster's own internal leader election, so that exactly one replica is actively "in charge" at a time even though three are running. Three services don't need this: the render engine (bcsdd), the USB Relay bridge, and the web frontend (www) are all safely active on every replica at once.
Postgres and Redis are deployed inside the cluster, not as externally managed services you bring yourself:
  • PostgreSQL runs as a 3-instance CloudNativePG cluster with synchronous replication and automatic failover, fronted by a PgBouncer connection pooler on every node.
  • Redis runs as a 3-replica Valkey (Redis-compatible) set, with Sentinel managing automatic failover between primary and replicas.

Storage#

Only two services need real persistent storage beyond ephemeral scratch space: the connection service and the www service both mount a single shared, ReadWriteMany-capable volume for the module cache. Postgres, Redis, and Redis Sentinel each have their own dedicated volume. Everything else uses ephemeral storage that doesn't need to survive a pod restart.

Warning

In this deployment's shipped configuration, the chore service (which is what actually writes scheduled backups) has no persistent volume at all; by default, a scheduled backup lands in ephemeral storage that's lost if that pod restarts or is rescheduled. If you rely on Configure and monitor scheduled backups in a Kubernetes deployment, mount a persistent volume into the chore service yourself and point its backup path at it. Don't assume backups are durable here without doing that.

Separately, the underlying storage layer can take its own periodic snapshots of the database and module-cache volumes. This is a lower-level safety net for the volumes themselves, and isn't a substitute for the application's own Configure and monitor scheduled backups feature described above.

Ingress and TLS#

Only the web frontend is exposed outside the cluster through the ingress: every other service stays internal-only, with one exception, covered below. TLS terminates at the ingress, using the same certificate-management flow already covered in Replace the HTTPS certificate: when Kubernetes is managing the certificate (for example, via cert-manager), Buttons' own Settings → Certificates page shows "Certificate managed by Kubernetes," and the certificate is delivered to the ingress the same way regardless of which service happens to be handling a given request.
The ingress accepts both plain HTTP (port 80) and HTTPS (port 443) by default, routed to the same backend: reaching Buttons over http:// works out of the box, with no automatic redirect to HTTPS.

Note

HSTS and an HTTP→HTTPS redirect are both included in the shipped manifests but commented out by default. To turn them on, uncomment the hsts and redirect-https middleware references in kubernetes/buttons/ingress/ingress.yaml and reapply it; enable both together, since HSTS alone doesn't stop port 80 from serving requests unencrypted. Once enabled, browsers that visit over HTTPS cache the HSTS setting and will only connect to Buttons over HTTPS afterward, so confirm your certificate is working correctly before turning it on.

Tally TSL Server connections are the exception#

Tally pods aren't hostNetwork, and their own Service exposes only the internal leader-election ports, so a Tally/UMD Server (Listen) connection (see Send or receive tally and labels over TSL/UMD) has nothing to be reachable through by default.
To fix this, the tally app's leader pushes one watchdog port forward per enabled TSL Server connection, each creating its own dedicated LoadBalancer Service (app-tally-pf-<connection id>) outside the ingress, exposing exactly that connection's configured port and transport. Plan a facility firewall rule for each such Service's external address the same way you would for the ingress. This Service follows the elected leader: on failover, the watchdog retargets its selector to the new leader pod, and the Service itself is never deleted on step-down, to avoid flapping the external load balancer during the handover.
A TSL Server connection can't be set to listen on 3131 (reserved for the health check) or 7110–7112 (reserved for tally leader election): Buttons rejects the port with an explicit "reserved by Buttons for..." error before you can save it.

Deploy#

From the root of the repository's kubernetes directory:
kubectl apply -f ./infrastructure --recursive
kubectl apply -f ./buttons --recursive
This applies the infrastructure layer (storage classes, ingress, the database operator) first, then the Buttons application manifests, in the order the manifests are meant to be applied.

Upgrade#

There's no separate migration job that runs and completes before the rest of the cluster starts: migrations run inside the chore service itself, guarded by a database lock so that only one of its three replicas actually performs them while the others wait. Other services don't wait on migrations finishing before they start. Plan upgrades with this in mind: a rolling update doesn't guarantee migrations are already complete before every service comes back up.

Health and failover#

Kubernetes' own liveness and readiness probes govern low-level pod lifecycle (whether a container gets restarted, and whether it receives traffic) independently of anything in Buttons' own interface. Buttons' own System Health banner and per-service status, covered in Interpret system health and restart services safely, is a separate, richer signal layered on top: it factors in real pod readiness alongside the cluster's own leader-election state, so treat the two as related but not identical: a service can be "ready" at the Kubernetes level while Buttons' own view of the cluster still shows something worth attention.
For planned node maintenance, use Cordon & Drain from the Services page (also covered in that same guide) rather than manually stopping pods: it's the action designed for taking a node offline while keeping the cluster running on the remaining two.

If you get stuck#

What you see
What to try
Scheduled backups don't survive a chore pod restart.
This is expected in the default configuration: mount a persistent volume into the chore service and point its backup path there.
A rolling upgrade seems to run before migrations finish.
Expected: migrations run inside chore under a database lock, and other services don't wait for them. Give an upgrade a moment to fully settle before assuming something's wrong.
Buttons shows no active license, or license features are missing, in an HA deployment.
Confirm you actually have an Enterprise license: any lower tier is treated as no license at all once HA mode is on, not a reduced one.
You want to check individual pod status directly.
kubectl -n buttons get pods, and kubectl -n buttons logs -f statefulset/<service-name> for a specific service's logs.
You're considering more than three nodes.
This deployment's replica counts, node-affinity rules, and leader-election quorum math are all built around exactly three: treat a different node count as unsupported rather than just changing the replica field.
You're wondering why it's three nodes and not one or two.
See Understand HA clustering: quorum needs a real majority, and three is the smallest number that tolerates losing a node without any ambiguity over which replica is in charge.

Where to go next#

  • Understand HA clustering, for why this deployment is built around exactly three nodes.
  • Choose an installation path, if you're still deciding whether this is the right deployment for your situation.
  • Interpret system health and restart services safely
  • Replace the HTTPS certificate
  • Configure and monitor scheduled backups
  • Manage your Buttons license

Was this helpful?

Was this helpful?

0 of 0 users found this page helpful