Bitfocus AS
logo
logo
Bitfocus AS
logo
logo
Sign upSign in

Loading...

Bitfocus

Subscribe to our newsletter

The latest news, articles, and resources, sent to your inbox.

FacebookInstagramGitHubYouTubeLinkedIn

Products

  • Buttons
  • Companion

Integrations

  • Supported Devices
  • Developer Community
  • Connection Development

Support

  • Support Overview
  • Documentation
  • Video Tutorials
  • Community Forum

Sales

  • Resellers & Integrators
  • Buttons Pricing

Updates

  • Case Studies
  • Events & Trade Shows
  • Press Releases
  • Product Updates
  • Webinars

Legal

  • Legal Overview
  • Privacy Policy
  • Buttons EULA
  • Terms & Cookie Policy

Company

  • About us
  • Press kit
  • Careers

© 2026 Bitfocus AS. All rights reserved.

Understand HA clustering
Docs for
Overview
Getting started
What is Bitfocus Buttons?
Install Buttons and get started
Manage your Buttons license
Activate Buttons offline
Find your way around Buttons
Create your first backup
Add an ATEM connection
Choose a control method
Choose an installation path
Install Buttons on Debian or Ubuntu
Understand HA clustering
Kubernetes HA
Update or remove Buttons
Positions
Understand positions
Create a position
Add controls and sections to a position
Create your first button
Use a connection's presets
Build more capable button actions
Add more feedback to a button
Organize controls in a section
Shift Section
Organize controls with a Folder Section
Add a Popover Section
Build and reuse a Shared Section
Build a Router Section
Understand Custom Routers
Custom Router panel
Surfaces
Surface compatibility
Add and attach a surface
Device orientation
Connections
Update a connection's module safely
Monitor and troubleshoot a connection
Router integrations
VideoHub and AJA KUMO
Utah Scientific BPS
Generic SW-P-08
Nevion VideoIPath
Arkona BLADE//runner
Routing
Physical routing
Configure ports and labels
Take a physical route
Understand route status
Topology graph
Routing Presets
Get started with virtual routing
Configure Nested Shapes
Reverse routing
Tielines
Routing Projects
Routing settings
Troubleshoot a route
Tally
Understand the Tally system
Send ATEM tally and labels to a UMD
Interpret Active Tally state
TSL/UMD connections
Diagnose tally problems
NMOS
Understand NMOS in Buttons
Connect Buttons to an NMOS Registry
Built-in Registry Server
Configure NMOS connections
Discover and adopt
Browse the NMOS inventory
Manage NMOS multicast addresses
Diagnose NMOS problems
Understand Cuelists
Build a Cuelist
Read and advance a running Cuelist
Control a Cuelist from a Position
Workflows
Understand workflows
Build your first workflow
Reuse a group of workflow nodes safely
Troubleshoot a workflow
Recipes
Sequence a timed automation
Call an HTTP endpoint from a workflow
REST endpoint
Use variables
Understand variable scope
Understand nested variables
Update expressions for v1.8
Plan and use Tags
Access
Create and manage users
Create roles and assign permissions
Grant access to specific resources
Show different controls by role
Sessions
Set up PIN and NFC sign-in
SSO
Get started with SSO
Connect a generic OIDC provider
Connect LDAP or Active Directory
Map identity claims to roles
Secure a Buttons deployment
Integrations
External control
Connect to Bitfocus Listener
USB Relay
Install USB Relay on Windows
Install USB Relay on macOS
Install USB Relay on Linux
Install USB Relay on a Raspberry Pi
Get started with the Control API
Secure and monitor the Control API
Control API reference
API reference
Administration
Enable and manage installable features
Services and health
Configure and monitor scheduled backups
Restore a backup and verify it
Export or import Buttons configuration
Store and rotate connection secrets
Replace the HTTPS certificate
HA backup and recovery
Settings
Collect support information
Reference
Glossary
Button Inspector reference
Network ports reference
Expressions
Internal actions reference
Routing Presets panel reference
Startup configuration reference
Workflow nodes
Connection workflow nodes
Workflow workflow nodes
Internal workflow nodes
Position workflow nodes
API workflow nodes
Utility workflow nodes

Loading...

Previous
← Install Buttons on Debian or Ubuntu
Next
Kubernetes HA →
Contact support →
You are viewing documentation for Buttons 1.8.See the docs for Buttons 1.6
Buttons/Installation/Understand HA clustering

Understand HA clustering

Kubernetes high availability lets Buttons keep running through the loss of a single node (a crash, a reboot, planned maintenance) without an operator intervening and without losing data. This page explains the idea that makes that possible: quorum, why it needs a minimum of three nodes to work at all, and what it does and doesn't guarantee once it's running.

What high availability means here#

Running several copies of something isn't automatically "high availability." Buttons' web frontend and a couple of other stateless services already run active on all three nodes at once, mainly so there's always one nearby to answer a request: that's redundancy, not failover.
The harder problem is the stateful parts: the database, the shared cache, and "which replica is currently in charge" for services that can't safely have two copies acting at once. For those, losing a node has to end in exactly one clear outcome: a surviving replica takes over, cleanly, with no two replicas ever believing they're both in charge at the same time. That's what this page is actually about.

The idea behind quorum#

A quorum is just "more than half the group has to agree before it counts." Picture a three-person committee that requires a majority to approve anything: any two of the three agreeing is decisive, because it's mathematically impossible for the other two people to also form a majority around a different answer at the same time. Whichever pair actually agrees, their answer is the only one that can win.
That's the entire trick HA clustering relies on. If "more than half of the nodes agree on who's in charge" is the rule, only one node can ever satisfy it at once, so there's never a moment where two nodes both correctly believe they're the active one. That failure mode, two nodes acting as the leader simultaneously, is usually called split-brain, and it's the thing quorum exists to rule out.

Why one node can't do this#

A single node has nothing to reach quorum with: it's automatically 100% of a group of one. That's not high availability, it's just availability: if that one node goes down, there's no surviving copy to hand off to, and no vote to take, because there's no one else to vote. A single-node deployment is a single point of failure by definition, which is exactly what HA clustering is meant to remove.

Why two nodes don't solve it either#

Two might sound like it doubles your safety margin, but it doesn't. With two nodes, "more than half" means both: there's no way for one node alone to form a majority of two. So the moment either node goes down, the survivor can't establish quorum on its own, and the safe behavior is to stop rather than guess. A two-node deployment can't actually survive losing either node and keep working; it's still a single point of failure, just built from twice the hardware.
The alternative (letting one surviving node out of two carry on by itself) would remove that safety margin rather than add it: after a brief network hiccup that splits the two nodes apart without actually killing either one, both would satisfy "I'm the only node I can see," and both would carry on believing they're in charge. That's the split-brain scenario quorum exists to prevent, and it's exactly what "requires a real majority" rules out.

Why three is the answer#

Three is the smallest number where both things are true at once: the cluster can lose exactly one node and the two survivors still form a majority (2 of 3), and a majority can never form around more than one answer at a time. That combination (genuine fault tolerance, with no ambiguity about who's in charge) is why Buttons' Kubernetes HA deployment is built around three nodes rather than two, or one.

Where this shows up in Buttons#

The three-node HA deployment isn't one mechanism: it's three separate systems that each independently apply the same majority-quorum principle to their own job:
  • Buttons' own leader election decides which of the three replicas of each stateful service (chore, connection, surface, and the others listed in Understand the topology) is actively in charge. It's built on the same reasoning as the committee example above: a replica only takes over once a majority of the cluster agrees it should.
  • PostgreSQL, run as a CloudNativePG cluster, requires a write to be confirmed by at least one other node before it's considered committed, and that requirement holds even if the primary loses contact with the rest of the cluster. A primary that's been cut off can't quietly keep accepting writes on its own; it blocks until it can reach a standby, which is what stops the database itself from ever quietly diverging into two different versions of the truth.
  • Redis, run as a Valkey set with Sentinel, needs a majority of Sentinels to agree before promoting a new primary after a failure: the same rule, applied to the cache layer.
None of these three systems talk to each other to make this decision; each one separately needs a working majority of three nodes to do its job safely.

What quorum protects against, and what it doesn't#

  • It protects against exactly one node failing at a time: a crash, a reboot, or planned maintenance taken one node at a time (see Understand what "self-healing" actually means for how Cordon & Drain triggers this deliberately). The cluster keeps serving requests and accepting writes throughout, without an operator needing to step in.
  • It does not protect against losing two of the three nodes at once. That takes the cluster below the quorum every one of these systems needs: Buttons' own leader election stalls, Sentinel can't reach a majority to promote a new Redis primary, and an isolated Postgres primary (if it's the node still standing) blocks new writes rather than risk the cluster diverging. There's no margin left for a second simultaneous failure; the cluster needs at least one of the lost nodes back before it can resume.
  • It is not a way to scale out. Adding more nodes isn't a supported way to buy more redundancy today, as covered in Deploy Buttons with Kubernetes high availability: the current deployment's replica counts, node-affinity rules, and quorum math are all specifically built around three.

Where to go next#

  • Deploy Buttons with Kubernetes high availability
  • Choose an installation path
  • Interpret system health and restart services safely

Was this helpful?

Was this helpful?

0 of 0 users found this page helpful