Monitoring a fleet of small devices
You cannot fix what you cannot see, and small devices fail quietly. A little monitoring, set up early, turns surprises into alerts.
What to measure
Start with a handful of signals that catch most problems:
- Heartbeat / reachability — is the device reporting at all?
- Uptime and reboot count — unexpected reboots reveal power or stability issues.
- Temperature and fan state.
- CPU, memory and load, and disk usage.
- Storage health where available, since flash wears out.
- Network — link state, signal, latency, packet loss.
- Service state — is the one thing the device exists to do actually working?
Miners add their own: hashrate, per-board temperatures, hardware-error rates and rejected shares. Platforms like HiveOS collect these for you.
Push or pull?
- Pull (a server scrapes each device, as Prometheus does) is simple when devices are reachable, but hard through NAT.
- Push (devices send to a collector, for example over MQTT or to InfluxDB) works behind NAT and suits intermittent links.
Keep the agent small. On constrained devices use light collectors and a modest sampling interval.
Alert on absence
The most valuable alert is often a missing heartbeat. Add alerts on thresholds (temperature, disk full) and on sudden change (reboot loops), and keep them few enough that people still read them.
Logs
Ship logs to a central place with syslog or a lightweight forwarder so evidence survives a crash. On flash-based devices keep local logs in RAM to limit wear.
Secure the pipe
Encrypt telemetry with TLS, and consider mutual TLS so the collector only accepts known devices. Never put credentials for the monitoring system in an image you distribute.
Show it
A dashboard tool such as Grafana makes trends visible. Container-based platforms such as balenaOS and self-hosting stacks like umbrelOS include their own status views; use them and add your own checks for what matters to you.
Test the whole path by unplugging a device on purpose and confirming the alert arrives.