I run about a dozen Linux servers at home: GitLab and its runner, Proxmox, Plex, a photo library, a web host and a Docker host for everything else. Every one of them reports into Home Assistant through my MQTT agent, which gives each server an Agent Online sensor and an Updates Pending count.
Having the sensors is the easy part. The hard part is getting alerts that are actually useful.
The problem with one automation per server
My first attempt was the obvious one. Each server got its own “offline” automation and its own “clean up” automation to clear the notification afterwards. With ten servers that’s twenty automations, all nearly identical, and every time I added a server I had to copy two more and remember to change every entity in them.
Worse, when the network blipped or the Docker host rebooted, I’d get a flood of separate notifications on my phone, one per server, and then a second flood as they came back.
So I replaced the lot with two automations:
- Server offline: one notification listing every server that’s currently down
- Server updates: one notification listing every server with updates waiting
Both clear themselves when everything is healthy again.
How it works
The trick is a fixed notification_id. Home Assistant’s persistent notifications replace any existing notification with the same id instead of stacking a new one. So whenever any server changes state, the automation rebuilds the full list from scratch and overwrites the notification. If the list is empty, it dismisses the notification instead.
That gives three nice properties:
- One notification, always current. Three servers down means one notification with three names in it, not three notifications.
- It clears itself. When the last server comes back, the notification disappears. No separate clean up automation.
- It survives restarts. Straight after Home Assistant starts, every agent sensor is unavailable until its server reports in. So the automation waits fifteen minutes after startup, then checks again.
The offline automation
Here’s my version, with the entity names swapped for generic ones. Paste it into a new automation in YAML mode and change the list to match your servers.
alias: Server offline notification
description: >-
One notification listing every offline server. Clears itself when they're
all back, and rechecks after a Home Assistant restart.
mode: parallel
triggers:
- trigger: state
entity_id:
- sensor.docker_host_agent_online
- sensor.gitlab_agent_online
- sensor.gitlab_runner_agent_online
- sensor.plex_agent_online
- sensor.proxmox_agent_online
for:
minutes: 10
id: sensor-change
- trigger: homeassistant
event: start
id: reboot
actions:
- if:
- condition: trigger
id: reboot
then:
- delay:
minutes: 15
- variables:
server_list:
- name: Docker host
entity: sensor.docker_host_agent_online
- name: GitLab
entity: sensor.gitlab_agent_online
- name: GitLab Runner
entity: sensor.gitlab_runner_agent_online
- name: Plex
entity: sensor.plex_agent_online
- name: Proxmox
entity: sensor.proxmox_agent_online
offline: >
{% set ns = namespace(items=[]) %}
{% for s in server_list %}
{% if states(s.entity) != 'online' %}
{% set ns.items = ns.items + [s.name] %}
{% endif %}
{% endfor %}
{{ ns.items }}
- choose:
- conditions:
- condition: template
value_template: "{{ offline | count > 0 }}"
sequence:
- action: persistent_notification.create
data:
notification_id: server-offline
title: Server status check
message: "The following servers are offline:\n{{ offline | join('\n') }}"
- action: notify.mobile_app_your_phone
data:
title: Server status check
message: "Offline: {{ offline | join(', ') }}"
default:
- action: persistent_notification.dismiss
data:
notification_id: server-offline
A few details worth pointing out:
- The trigger has no
to:, so it fires on any change that lasts ten minutes, whether a server went down or came back. Either way the list gets rebuilt, which is exactly what you want. - The ten minute
for:stops a quick reboot after patching from setting off an alert. My servers patch themselves with my AutoPatch script, and I don’t want a notification every time one of them restarts. != 'online'catches everything. The agent’s sensors go unavailable within two minutes if a server stops reporting, so a dead server, a crashed agent and a broken network all show up the same way.mode: parallelmeans a second server going down while the first run is still in its startup delay doesn’t get dropped.
The updates automation
The updates version is the same pattern with a different test. Instead of checking for online, it reads the Updates Pending count and lists any server above zero:
outstanding: >
{% set ns = namespace(items=[]) %}
{% for s in server_list %}
{% set count = states(s.entity) | float(0) %}
{% if count > 0 %}
{% set ns.items = ns.items + [s.name ~ ': ' ~ count | int ~ ' update(s)'] %}
{% endif %}
{% endfor %}
{{ ns.items }}
I don’t send updates to my phone, though. Pending updates aren’t urgent, so that one only creates the notification in Home Assistant, and it disappears once AutoPatch has done its job.
One gotcha: keep the two lists in step
Home Assistant won’t let you build the trigger list from a template, so each server appears twice: once under triggers and once in server_list. If you add a server to one and forget the other, it either never triggers the automation or never shows up in the message.
I know this because I did exactly that. When I looked back through my own setup while writing this post, my GitLab runner was in the triggers but missing from the list, and my newest server wasn’t in either. Both are fixed now, but it’s worth checking yours whenever you add a machine.
What about websites?
Servers aren’t the only thing I watch. Uptime Kuma keeps an eye on the websites I look after, including this one, and its Home Assistant integration turns each monitor into an entity. I pair those with Home Assistant’s alert integration, which keeps reminding me until the site is back up instead of telling me once and hoping I noticed.
Wrapping up
Two automations instead of twenty, one notification instead of a flood, and nothing to clean up by hand. If you’re already using my MQTT agent, this is the first automation I’d add.
More projects
total 8