Bitfocus AS
logo
logo
Bitfocus AS
logo
logo
Sign upSign in

Loading...

Bitfocus

Subscribe to our newsletter

The latest news, articles, and resources, sent to your inbox.

FacebookInstagramGitHubYouTubeLinkedIn

Products

  • Buttons
  • Companion

Integrations

  • Supported Devices
  • Developer Community
  • Connection Development

Support

  • Support Overview
  • Documentation
  • Video Tutorials
  • Community Forum

Sales

  • Resellers & Integrators
  • Buttons Pricing

Updates

  • Case Studies
  • Events & Trade Shows
  • Press Releases
  • Product Updates
  • Webinars

Legal

  • Legal Overview
  • Privacy Policy
  • Buttons EULA
  • Terms & Cookie Policy

Company

  • About us
  • Press kit
  • Careers

© 2026 Bitfocus AS. All rights reserved.

Services and health
Docs for
Overview
Getting started
What is Bitfocus Buttons?
Install Buttons and get started
Manage your Buttons license
Activate Buttons offline
Find your way around Buttons
Create your first backup
Add an ATEM connection
Choose a control method
Choose an installation path
Install Buttons on Debian or Ubuntu
Understand HA clustering
Kubernetes HA
Update or remove Buttons
Positions
Understand positions
Create a position
Add controls and sections to a position
Create your first button
Use a connection's presets
Build more capable button actions
Add more feedback to a button
Organize controls in a section
Shift Section
Organize controls with a Folder Section
Add a Popover Section
Build and reuse a Shared Section
Build a Router Section
Understand Custom Routers
Custom Router panel
Surfaces
Surface compatibility
Add and attach a surface
Device orientation
Connections
Update a connection's module safely
Monitor and troubleshoot a connection
Router integrations
VideoHub and AJA KUMO
Utah Scientific BPS
Generic SW-P-08
Nevion VideoIPath
Arkona BLADE//runner
Routing
Physical routing
Configure ports and labels
Take a physical route
Understand route status
Topology graph
Routing Presets
Get started with virtual routing
Configure Nested Shapes
Reverse routing
Tielines
Routing Projects
Routing settings
Troubleshoot a route
Tally
Understand the Tally system
Send ATEM tally and labels to a UMD
Interpret Active Tally state
TSL/UMD connections
Diagnose tally problems
NMOS
Understand NMOS in Buttons
Connect Buttons to an NMOS Registry
Built-in Registry Server
Configure NMOS connections
Discover and adopt
Browse the NMOS inventory
Manage NMOS multicast addresses
Diagnose NMOS problems
Understand Cuelists
Build a Cuelist
Read and advance a running Cuelist
Control a Cuelist from a Position
Workflows
Understand workflows
Build your first workflow
Reuse a group of workflow nodes safely
Troubleshoot a workflow
Recipes
Sequence a timed automation
Call an HTTP endpoint from a workflow
REST endpoint
Use variables
Understand variable scope
Understand nested variables
Update expressions for v1.8
Plan and use Tags
Access
Create and manage users
Create roles and assign permissions
Grant access to specific resources
Show different controls by role
Sessions
Set up PIN and NFC sign-in
SSO
Get started with SSO
Connect a generic OIDC provider
Connect LDAP or Active Directory
Map identity claims to roles
Secure a Buttons deployment
Integrations
External control
Connect to Bitfocus Listener
USB Relay
Install USB Relay on Windows
Install USB Relay on macOS
Install USB Relay on Linux
Install USB Relay on a Raspberry Pi
Get started with the Control API
Secure and monitor the Control API
Control API reference
API reference
Administration
Enable and manage installable features
Services and health
Configure and monitor scheduled backups
Restore a backup and verify it
Export or import Buttons configuration
Store and rotate connection secrets
Replace the HTTPS certificate
HA backup and recovery
Settings
Collect support information
Reference
Glossary
Button Inspector reference
Network ports reference
Expressions
Internal actions reference
Routing Presets panel reference
Startup configuration reference
Workflow nodes
Connection workflow nodes
Workflow workflow nodes
Internal workflow nodes
Position workflow nodes
API workflow nodes
Utility workflow nodes

Loading...

Previous
← Enable and manage installable features
Next
Configure and monitor scheduled backups →
Contact support →
You are viewing documentation for Buttons 1.8.See the docs for Buttons 1.6
Buttons/Administration/Services and health

Services and health

Buttons runs as a set of separate processes working together, watched over by a supervisor that keeps them alive. Most day-to-day health information shows up on the Dashboard; the full picture, and the only place to restart an individual process, is Settings → Services.

Read the Dashboard's health signals#

The Dashboard's Home navigation icon shows an amber count badge whenever there's something worth looking at, and the connection status indicator turns amber specifically for connection-related issues.
A dedicated System Warnings card appears on the Dashboard only when there's a standing issue in one of three categories: a process problem, a high-availability problem, or a browser-compatibility problem. Each entry is labeled by category: Service (a process issue), HA, or Browser, followed by the specific warning and its message. Only Service (process) warnings carry a restart control here, and clicking it restarts that process immediately, with no confirmation step.
Issues in other categories (surfaces, connections, workflows, NMOS) don't get a dedicated warnings card. Instead, the affected item shows an amber status dot in its own list on the Dashboard (Surfaces, Connections, Workflows, NMOS), and errored items sort to the top of their list so they're easy to spot.

Check and restart services in Settings → Services#

Open Settings → Services for the full list of running processes. In a standalone deployment the page is titled High Availability Status with the description "Status and control of all running services." In a high-availability deployment (multiple nodes), the description instead reads "Real-time status of all processes across nodes."
A green banner reading System Health: OK / All services are running normally appears when nothing needs attention. Each service is listed with a status dot (Running, Starting, Waiting, or Stopped) and an uptime figure.
To restart a single process, use its restart icon in the row. This takes effect immediately: there's no confirmation dialog and no warning about the interruption it causes, so only restart a process when you're ready for it to briefly go offline. One process is protected from this control: the key-value store (Valkey on macOS/Linux, Garnet on Windows) can't be restarted from here, since too much else depends on it staying up.

Understand what "self-healing" actually means#

Buttons doesn't automatically restart a crashed process for you outside of one specific case: in a high-availability deployment, taking a node offline with Cordon & Drain causes the services that were running on it to start automatically on other available nodes in the cluster. That's a deliberate failover action you take, not an automatic recovery from a crash. Outside of it, a stopped or stuck process stays stopped until you restart it yourself.

What you see in the editor during a brief outage#

A person actively working in the editor doesn't just see an error the moment a service they depend on (the database, for example) becomes briefly unreachable. Buttons keeps the page mounted, including any unsaved form you were editing, and shows a recovery state that blocks interaction and marks visible information as potentially stale while it checks the session and retries in the background, roughly once a second for the first minute and roughly every ten seconds after that. Once the service answers again, Buttons automatically refetches what failed and reconnects any live subscriptions, with no manual reload needed.
If the outage runs past about 20 seconds, a Refresh button also appears on the recovery dialog as a manual fallback, reloading the page immediately rather than waiting for the automatic recovery.
This only covers a genuine, temporary infrastructure problem: a session that's actually expired or been revoked still requires signing in again, and an ordinary error from a specific request doesn't trigger it. Buttons never replays a mutation (an action you tried to submit) on your behalf once recovery completes: if you attempted to save something during the outage, check whether it actually went through before assuming it did.

Restart the whole system#

There's no dedicated "restart everything" control. The only full-system reset is Delete Configuration on the Configuration Data page, which resets the installation to factory settings and restarts the system as a side effect of that reset. It isn't a safe way to simply bounce every process, since it also wipes your configuration. See Export or import Buttons configuration before considering it.

If you get stuck#

What you see
What to try
The Dashboard shows an amber badge but no System Warnings card.
The issue is likely scoped to a surface, connection, workflow, or NMOS item rather than a process: check each resource list on the Dashboard for an amber status dot.
A service shows Stopped and isn't coming back on its own.
Restart it manually from Settings → Services: outside of node failover during Cordon & Drain, Buttons doesn't restart a stopped process automatically.
The restart icon is unavailable for a service.
The key-value store (Valkey/Garnet) can't be restarted from this page; if it needs attention, that's a deeper intervention outside routine service restarts.
You want to recover from a broad, unclear problem.
Restart the specific process shown as unhealthy rather than reaching for Delete Configuration: that action resets your entire configuration, not just the running processes.
The editor is blocked with a recovery state and won't clear.
Give it time: it retries roughly once a second for the first minute, then roughly every ten seconds. If the underlying service is actually still down, fix that first, see Check and restart services in Settings → Services above.
The Watchdog itself appears to hang, unresponsive, during startup or while stopping every process.
A prior lock-ordering bug in the Watchdog's process manager could deadlock it in either direction; this is fixed. If you still see it on a current version, treat it as a new finding rather than a known issue.

Where to go next#

  • Secure a Buttons deployment, for how service health fits into a broader operational baseline.
  • Deploy Buttons with Kubernetes high availability, for how this health signal relates to Kubernetes' own pod probes in a clustered deployment.

Was this helpful?

Was this helpful?

0 of 0 users found this page helpful