Kubernetes high availability lets Buttons keep running through the loss of a single node (a crash, a reboot, planned maintenance) without an operator intervening and without losing data. This page explains the idea that makes that possible: quorum, why it needs a minimum of three nodes to work at all, and what it does and doesn't guarantee once it's running.
What high availability means here#
Running several copies of something isn't automatically "high availability." Buttons' web frontend and a couple of other stateless services already run active on all three nodes at once, mainly so there's always one nearby to answer a request: that's redundancy, not failover.
The harder problem is the stateful parts: the database, the shared cache, and "which replica is currently in charge" for services that can't safely have two copies acting at once. For those, losing a node has to end in exactly one clear outcome: a surviving replica takes over, cleanly, with no two replicas ever believing they're both in charge at the same time. That's what this page is actually about.
The idea behind quorum#
A quorum is just "more than half the group has to agree before it counts." Picture a three-person committee that requires a majority to approve anything: any two of the three agreeing is decisive, because it's mathematically impossible for the other two people to also form a majority around a different answer at the same time. Whichever pair actually agrees, their answer is the only one that can win.
That's the entire trick HA clustering relies on. If "more than half of the nodes agree on who's in charge" is the rule, only one node can ever satisfy it at once, so there's never a moment where two nodes both correctly believe they're the active one. That failure mode, two nodes acting as the leader simultaneously, is usually called split-brain, and it's the thing quorum exists to rule out.
Why one node can't do this#
A single node has nothing to reach quorum with: it's automatically 100% of a group of one. That's not high availability, it's just availability: if that one node goes down, there's no surviving copy to hand off to, and no vote to take, because there's no one else to vote. A single-node deployment is a single point of failure by definition, which is exactly what HA clustering is meant to remove.
Why two nodes don't solve it either#
Two might sound like it doubles your safety margin, but it doesn't. With two nodes, "more than half" means both: there's no way for one node alone to form a majority of two. So the moment either node goes down, the survivor can't establish quorum on its own, and the safe behavior is to stop rather than guess. A two-node deployment can't actually survive losing either node and keep working; it's still a single point of failure, just built from twice the hardware.
The alternative (letting one surviving node out of two carry on by itself) would remove that safety margin rather than add it: after a brief network hiccup that splits the two nodes apart without actually killing either one, both would satisfy "I'm the only node I can see," and both would carry on believing they're in charge. That's the split-brain scenario quorum exists to prevent, and it's exactly what "requires a real majority" rules out.
Why three is the answer#
Three is the smallest number where both things are true at once: the cluster can lose exactly one node and the two survivors still form a majority (2 of 3), and a majority can never form around more than one answer at a time. That combination (genuine fault tolerance, with no ambiguity about who's in charge) is why Buttons' Kubernetes HA deployment is built around three nodes rather than two, or one.
The three-node HA deployment isn't one mechanism: it's three separate systems that each independently apply the same majority-quorum principle to their own job:
- Buttons' own leader election decides which of the three replicas of each stateful service (chore, connection, surface, and the others listed in Understand the topology) is actively in charge. It's built on the same reasoning as the committee example above: a replica only takes over once a majority of the cluster agrees it should.
- PostgreSQL, run as a CloudNativePG cluster, requires a write to be confirmed by at least one other node before it's considered committed, and that requirement holds even if the primary loses contact with the rest of the cluster. A primary that's been cut off can't quietly keep accepting writes on its own; it blocks until it can reach a standby, which is what stops the database itself from ever quietly diverging into two different versions of the truth.
- Redis, run as a Valkey set with Sentinel, needs a majority of Sentinels to agree before promoting a new primary after a failure: the same rule, applied to the cache layer.
None of these three systems talk to each other to make this decision; each one separately needs a working majority of three nodes to do its job safely.
What quorum protects against, and what it doesn't#
- It protects against exactly one node failing at a time: a crash, a reboot, or planned maintenance taken one node at a time (see Understand what "self-healing" actually means for how Cordon & Drain triggers this deliberately). The cluster keeps serving requests and accepting writes throughout, without an operator needing to step in.
- It does not protect against losing two of the three nodes at once. That takes the cluster below the quorum every one of these systems needs: Buttons' own leader election stalls, Sentinel can't reach a majority to promote a new Redis primary, and an isolated Postgres primary (if it's the node still standing) blocks new writes rather than risk the cluster diverging. There's no margin left for a second simultaneous failure; the cluster needs at least one of the lost nodes back before it can resume.
- It is not a way to scale out. Adding more nodes isn't a supported way to buy more redundancy today, as covered in Deploy Buttons with Kubernetes high availability: the current deployment's replica counts, node-affinity rules, and quorum math are all specifically built around three.
Where to go next#