A
Kubernetes HA deployment survives losing a node without going down, because Postgres, Redis, and every Buttons service run replicated across three nodes. That replication protects you from a single failure: it does not, by itself, give you a backup. This page covers what's actually in place today for backing up and recovering configuration in that topology, and what you need to do yourself to get a durable, restorable backup out of it. It assumes you've already read
Configure and monitor scheduled backups,
Restore a backup and verify it, and
Export or import Buttons configuration: this page covers only what's different or additional in the Kubernetes HA topology, not the underlying backup and restore mechanism itself.
Before you begin#
- An Enterprise license: Kubernetes HA mode requires it; any lower tier is treated as no license at all once
HA_ENABLED is set, the same as elsewhere in the HA deployment. - Familiarity with the deployment's topology: this page assumes you know that Postgres and Redis run inside the cluster as CloudNativePG and Valkey/Sentinel, and that the chore service is what actually performs Buttons' own scheduled backups.
kubectl access to the cluster, for anything below that goes beyond Buttons' own Settings pages.
What actually protects your data today, and what doesn't#
Three separate layers exist in this deployment, and it's easy to assume any one of them is a backup when it isn't:
- Postgres replication (CloudNativePG). The database runs as three synchronously replicated instances with automatic failover. This protects against losing a node or a pod: a replica takes over and you keep running. It does not protect against corrupted data, an accidental deletion, or a bad migration: whatever happens to the primary replicates to the standbys too. As shipped, this deployment does not configure CloudNativePG's own backup capability (continuous WAL archiving or scheduled base backups to object storage): archiving is explicitly switched off, and no scheduled backup resource is defined for the database. So there's no built-in point-in-time recovery for Postgres here, only failover between healthy replicas.
- Underlying storage snapshots. The storage layer takes its own periodic, same-cluster snapshots of the database and shared module-cache volumes (daily for the database volumes, weekly for the module cache) as a lower-level safety net for the volumes themselves. These aren't off-site backups (the shipped configuration only takes local snapshots, not the kind of backup that copies data outside the cluster), and they aren't aware of Buttons' own data model: restoring one rolls back the whole volume, not a single configuration change. Restoring from one of these snapshots is a storage-administration operation outside Buttons' own interface; treat it as a last resort for your storage administrator or Kubernetes operator to perform using your storage layer's own tooling, not a Buttons-specific procedure this page can walk you through.
- Buttons' own scheduled and manual backups. This is the mechanism covered in Configure and monitor scheduled backups: the one that produces a portable, versioned archive you can actually restore through Buttons' own interface. In this deployment's shipped configuration, it has nowhere durable to write to by default. See the next section.
Scheduled backups have no durable place to land#
The
Kubernetes HA topology guide already flags that the chore service (which is what runs scheduled and manual backups) has no persistent volume in the shipped StatefulSet. It's worth being precise about what that means in practice, because it's a bit sharper than "backups don't survive a restart":
The chore container also runs with a read-only root filesystem, and the only writable location mounted into it is a small, temporary scratch volume that doesn't persist across a pod restart or reschedule. Because this deployment runs in the same container mode used for Docker installs, the backup path defaults to a location under the application's own working directory unless you explicitly set a different one, and that default location sits on the read-only part of the filesystem, not the writable scratch volume. In practice, that means: unless you configure a backup destination that points at a real, mounted, writable volume, scheduled and manual backups either fail to write at all (visible as
Last run error on the rule, per
Monitor whether backups are succeeding) or, if you happen to point them at the scratch volume, succeed but don't survive the pod being rescheduled.
To get a real, durable backup file out of this deployment today, you need to mount a persistent volume into the chore service yourself and configure its backup destination to write there. That isn't a configuration this deployment ships or tests out of the box: it's a customization you'd be making to the shipped manifests, not a documented, supported setup. Treat it as something to validate carefully (confirm a rule's Run Now actually produces a file, and that the file survives deleting and recreating the chore pod) before relying on it, rather than as a certified procedure.
Note
Whether Buttons should ship a first-class, durable backup destination for the chore service in Kubernetes HA deployments (for example, a dedicated persistent volume wired into the manifests by default) is an open product question this page can't resolve on its own. Until that's decided, durable backups in this topology are something you build yourself, not something the deployment gives you.
Redis and Sentinel: nothing to back up#
Redis (Valkey) and its Sentinel layer hold only short-lived coordination data in this deployment: service heartbeats, leader-election state, and pub/sub messages used to keep the three replicas of each service in sync, not durable configuration. Its own persistence is switched off in the shipped configuration, by design: it's meant to be rebuilt from scratch, not restored. The durable source of truth for everything you'd actually want backed up (connections, positions, routing, users, licenses) is Postgres. If a Redis or Sentinel pod loses its volume, it simply rejoins the cluster and rebuilds its state; there's nothing to restore and no recovery action needed on your part.
If you lose the Postgres cluster or its data#
A single Postgres pod or node failing is what the replication layer already handles: see
Health and failover for how to watch that happen safely. This section is about the less common case: the database itself is gone or corrupted beyond what failover can fix.
Your realistic options, in order:
- Restore from a Buttons backup archive, the same way as Restore a backup and verify it describes, once you have a working Postgres cluster and the www service back up. This only works if you already have a durable archive: either from a chore backup rule writing to a volume you configured as described above, or from a manual export you copied off the cluster beforehand.
- Restore the underlying storage snapshot, if one exists and your archive doesn't. This is a storage-layer operation for your Kubernetes or storage administrator to perform with your storage layer's own restore tooling, not a Buttons-specific procedure, and not something guaranteed to leave you at a consistent point for the application's own data model. Confirm with whoever manages your cluster's storage whether this is realistic before relying on it.
- If neither exists, there is currently no way back to your previous configuration: the same as a single-host deployment with no backup file, except that in this topology the gap in the first option is easy to hit without ever having realized it applied to you.
Warning
Rebuilding the Postgres cluster from a Buttons archive means reapplying the database manifests and importing the archive into a fresh installation: everyone connected is disconnected for the duration, the same as any restore, and any configuration created after the archive's export date is lost.
Migrations and recovery timing#
Migrations run inside the chore service under a database lock, with no separate step that gates the rest of the cluster until they finish: the same behavior
Upgrade describes for routine upgrades applies just as much to recovery. After restoring or rebuilding the Postgres cluster, give the chore leader a moment to finish migrating before trusting that other services are working correctly, and check the
System Health banner (see
Interpret system health and restart services safely) rather than assuming everything is consistent the instant pods report ready.
If you get stuck#
What you see | What to try |
|---|
A backup rule shows Last run error in this deployment, and you haven't configured a backup destination. | |
You need to recover from total loss of the Postgres cluster and have no backup archive. | |
Something looks inconsistent right after a database restore or rebuild. | Give chore's migration lock time to finish before judging the result: check System Health rather than acting immediately. |
Where to go next#