What it is
A regular automatic check that a service is alive and answering: not just that the process runs, but that the database reads and the page opens. It is what restarts a container and alerts the administrator.
How we use it
In Peregovorka the chat answers /healthz by reading from and writing to the database, and a watchdog restarts a stuck container at most three times an hour. The finance assistant is checked from outside every 5 minutes: containers, database, the "bot is alive" mark and the dashboard address.
Where it helps a business
- A bot "seems to work" but does not answer, and you hear about it from customers.
- A service hung at night, and nobody is around to restart it.
- One part of the project started before the database and crashed.
How we use it
- Peregovorka: containers answer health checks, and the chat answers
/healthz by reading from and writing to the database. A watchdog restarts a stuck container within a minute, but no more than three times an hour. - Finance assistant: every 5 minutes an external check looks at the containers, the database, the "bot is alive" mark (the bot sets it once a minute) and the dashboard address. A failure sends a Telegram message.
- IT help desk: Docker Compose starts the bot only after MariaDB and Redis have passed their health checks.
- Tech Poly VPN: monitoring checks the node's TCP port every 5 minutes and alerts the administrator only after several failures in a row.
Common problems
- A "process is alive" check. The process can be running while the database is not. The check should do what a user does.
- False alarms. A single network blip should not wake anyone at night, hence the alert after several failures in a row.
- Endless restarts. A "no more than three times an hour" limit stops a broken service from restarting in a loop.
When you do not need it
A static page hardly needs a health check: a certificate expiry check is enough.