Buttons runs as a set of separate processes working together, watched over by a supervisor that keeps them alive. Most day-to-day health information shows up on the Dashboard; the full picture, and the only place to restart an individual process, is Settings → Services.
Read the Dashboard's health signals#
The Dashboard's Home navigation icon shows an amber count badge whenever there's something worth looking at, and the connection status indicator turns amber specifically for connection-related issues.
A dedicated System Warnings card appears on the Dashboard only when there's a standing issue in one of three categories: a process problem, a high-availability problem, or a browser-compatibility problem. Each entry is labeled by category: Service (a process issue), HA, or Browser, followed by the specific warning and its message. Only Service (process) warnings carry a restart control here, and clicking it restarts that process immediately, with no confirmation step.
Issues in other categories (surfaces, connections, workflows, NMOS) don't get a dedicated warnings card. Instead, the affected item shows an amber status dot in its own list on the Dashboard (Surfaces, Connections, Workflows, NMOS), and errored items sort to the top of their list so they're easy to spot.
Check and restart services in Settings → Services#
Open Settings → Services for the full list of running processes. In a standalone deployment the page is titled High Availability Status with the description "Status and control of all running services." In a high-availability deployment (multiple nodes), the description instead reads "Real-time status of all processes across nodes."
A green banner reading System Health: OK / All services are running normally appears when nothing needs attention. Each service is listed with a status dot (Running, Starting, Waiting, or Stopped) and an uptime figure.
To restart a single process, use its restart icon in the row. This takes effect immediately: there's no confirmation dialog and no warning about the interruption it causes, so only restart a process when you're ready for it to briefly go offline. One process is protected from this control: the key-value store (Valkey on macOS/Linux, Garnet on Windows) can't be restarted from here, since too much else depends on it staying up.
Understand what "self-healing" actually means#
Buttons doesn't automatically restart a crashed process for you outside of one specific case: in a high-availability deployment, taking a node offline with Cordon & Drain causes the services that were running on it to start automatically on other available nodes in the cluster. That's a deliberate failover action you take, not an automatic recovery from a crash. Outside of it, a stopped or stuck process stays stopped until you restart it yourself.
What you see in the editor during a brief outage#
A person actively working in the editor doesn't just see an error the moment a service they depend on (the database, for example) becomes briefly unreachable. Buttons keeps the page mounted, including any unsaved form you were editing, and shows a recovery state that blocks interaction and marks visible information as potentially stale while it checks the session and retries in the background, roughly once a second for the first minute and roughly every ten seconds after that. Once the service answers again, Buttons automatically refetches what failed and reconnects any live subscriptions, with no manual reload needed.
If the outage runs past about 20 seconds, a Refresh button also appears on the recovery dialog as a manual fallback, reloading the page immediately rather than waiting for the automatic recovery.
This only covers a genuine, temporary infrastructure problem: a session that's actually expired or been revoked still requires signing in again, and an ordinary error from a specific request doesn't trigger it. Buttons never replays a mutation (an action you tried to submit) on your behalf once recovery completes: if you attempted to save something during the outage, check whether it actually went through before assuming it did.
Restart the whole system#
There's no dedicated "restart everything" control. The only full-system reset is
Delete Configuration on the Configuration Data page, which resets the installation to factory settings and restarts the system as a side effect of that reset. It isn't a safe way to simply bounce every process, since it also wipes your configuration. See
Export or import Buttons configuration before considering it.
If you get stuck#
What you see | What to try |
|---|
The Dashboard shows an amber badge but no System Warnings card. | The issue is likely scoped to a surface, connection, workflow, or NMOS item rather than a process: check each resource list on the Dashboard for an amber status dot. |
A service shows Stopped and isn't coming back on its own. | Restart it manually from Settings → Services: outside of node failover during Cordon & Drain, Buttons doesn't restart a stopped process automatically. |
The restart icon is unavailable for a service. | The key-value store (Valkey/Garnet) can't be restarted from this page; if it needs attention, that's a deeper intervention outside routine service restarts. |
You want to recover from a broad, unclear problem. | Restart the specific process shown as unhealthy rather than reaching for Delete Configuration: that action resets your entire configuration, not just the running processes. |
The editor is blocked with a recovery state and won't clear. | |
The Watchdog itself appears to hang, unresponsive, during startup or while stopping every process. | A prior lock-ordering bug in the Watchdog's process manager could deadlock it in either direction; this is fixed. If you still see it on a current version, treat it as a new finding rather than a known issue. |
Where to go next#