Alerts: Know When CPU, Disk, or a Service Needs You
The Alerts tab watches a Linux server for you. Switch on ready-made alerts for high CPU, low disk, a stopped service, a site going down, an expiring SSL certificate, or SSH brute force, and CtrlOps messages your Slack, Telegram, or webhook channel when one trips, even while the app is closed.
Most server problems show up well before they turn into outages. CPU pinned for ten minutes, a disk at 85% and climbing, a service that quietly stopped, a certificate with a week left. The trouble is that nobody is looking when it happens, so you hear about it from a user, or from the error that takes the site down. Doing it yourself means a check script per problem, a crontab line per script, and some way to send the message, on every server.
The Alerts tab gives you those checks ready-made. Every alert is listed with a toggle: switch on the ones you want to hear about, set a limit if the alert has one, pick where the message goes, and it is live. Each alert runs on the server on a schedule and messages your Slack, Telegram, or webhook channel when it trips, even while CtrlOps is closed. For a check that is not in the list, write it as a job in Cron Jobs and alert on that.
What you can do in the Alerts tab
- Switch on 13 ready-made alerts covering CPU, memory, swap, out-of-memory kills, disk space, inodes, services, reboots, websites, SSL certificates, SSH brute force, and your own cron jobs.
- Set the limit for an alert, for example CPU at 85%, and how many bad checks in a row it takes before you are told.
- Read a preview of the message before you turn the alert on.
- Send alerts to Slack, Telegram, or a webhook, using channels you set up once on the Home screen.
- Add the same alert more than once, for another service, site, or disk, or at another limit.
- Pause an alert with its toggle and keep its limit and channel for later.
- Narrow the list to what is on, paused, or not set up yet, or search it by name.
- Hear about problems while CtrlOps is closed, because the checks run on the server itself.
Why an Alerts tab instead of writing your own checks
| Doing it by hand | Doing it in CtrlOps |
|---|---|
Write a script that reads /proc/loadavg or parses top | Switch on High CPU usage |
| Decide on a threshold and how long it must hold before it counts | Set CPU limit and Bad checks in a row |
| Add a crontab line for the script on every server | The alert is installed as a cron job for you |
Wire up curl to a Slack webhook or the Telegram bot API | Pick a channel you saved once |
| Send yourself a test to see how the message reads | Read What the message looks like before you turn it on |
Grep auth.log for failed logins now and then | Switch on SSH brute force |
| Remember which checks run on which server | Every alert on the server is listed, with what is on, paused, or not set up |
| Comment the line out to silence a check | Flip the toggle to pause it |
Before you start
| Requirement | Why you need it | What to do if it is missing |
|---|---|---|
| An active server connection | Alerts are installed on the server you are connected to | Open the server from the Home page first |
| Root or sudo on the server | Alerts run as root, so a check can see every process, filesystem, and login | Connect as root, or as a user with sudo |
| Cron on the server | Each alert runs as a cron job | Nearly every Linux distribution ships it. If it is not running, start the cron service |
| An alert channel | It is where the message goes | Set one up once on the Alert Channels tab of the Home screen |
Alerts are per server. Switching on High CPU usage here watches this server only. To watch several servers, open each one and switch its alerts on there.
Open the Alerts tab
Connect to your server
Open CtrlOps and click the server you want to watch. Wait for the header to show Connected.
Click the Alerts tab
In the left sidebar, click Alerts. It sits below Audit Report, just above Cron Jobs.
See what is already set up
Alerts are grouped by category, each card with its toggle, how often it checks, and its status. On a server where nothing is set up yet, every card reads NOT SET UP and the chip next to the title says Nothing on yet.
The banner under the title sums up how the tab works: each alert runs on this server on a schedule and messages your channels when it trips, even while CtrlOps is closed. Refresh at the top right reloads the list, and Add alert next to it adds an alert, including a second copy of one you already use.
Turn on an alert
Flip the toggle
Click the toggle on the alert card
Flip the toggle on the alert you want to hear about, for example High CPU usage. A Turn on alert panel opens with two steps, Configure and Channels, and a line on what the alert watches for.
Set the limit
Configure asks for whatever this alert needs. For High CPU usage that is two fields:
| Field | Default | What it means |
|---|---|---|
| CPU limit | 85% | The CPU usage that counts as a bad check |
| Bad checks in a row before alerting | 2 | How many checks in a row must be over the limit before you are told |
Under the fields, a note shows how often the check runs and what your settings add up to, for example Checks every 5 min. Alerts once it stays above the limit for about 10 minutes. What the message looks like previews the message itself, with this server's name in it. Click Continue.
Pick where the alert goes
In the Channels step, choose which of your saved Slack, Telegram, or webhook channels should get the message, then turn the alert on. The card now shows your limit, a one-line summary, and the channel.
Once an alert is set up, its card says what you chose: the limit next to the name, a summary such as CPU above 85% for 2 checks, an icon for the channel it messages, and how often it checks. Until the first check runs, its status reads FIRST CHECK SOON.
Three alerts have nothing to set, because they watch for an event rather than a limit: Out-of-memory kill, Failed systemd services, and Server rebooted. They show On once they first run.
Turn an alert off
Flip the same toggle off. The alert is paused, not deleted: the card keeps its limit and its channel, reads PAUSED, and the Paused chip above the list counts it. Flip the toggle back on when you want to hear about it again.
Pause an alert during planned work instead of learning to ignore it. A long migration or a load test trips High CPU usage on purpose, so pause it, do the work, and switch it back on.
Add an alert more than once
Most alerts can be added again, for another service, site, or disk, or at another limit. Two CPU alerts, one at 75% and one at 90%, give you an early warning and an urgent one. Two Service stopped alerts watch two services.
Click Add alert
Click Add alert (top right of the Alerts tab)
The Add alert panel lists every alert by category, with a search box at the top and how often each one checks.
Pick the alert
Pick the alert to add again, for example High CPU usage. It opens the same Turn on alert panel as the toggle.
Set the new limit and its channels
Set the second limit, service, site, or disk, choose its channels, and turn it on. It runs alongside the first one.
On a card that is already set up, Add another is the shortcut for adding that same alert again.
The alert catalog
Thirteen alerts in six groups. Checks is how often the check runs on the server, and the limits shown are the defaults.
System resources
| Alert | Alerts when | Checks |
|---|---|---|
| High CPU usage | CPU stays above 85% for about 10 minutes. Includes the top processes. | Every 5 min |
| High memory usage | Memory stays above 90% for about 10 minutes. Includes the biggest processes. | Every 5 min |
| Swap usage high | Swap stays above 80%, the early sign that RAM has run out. | Every 5 min |
| Out-of-memory kill | The kernel killed a process because memory ran out. | Every 5 min |
Storage
| Alert | Alerts when | Checks |
|---|---|---|
| Disk space low | A filesystem goes above 85%. Includes the largest folders. | Every 30 min |
| Inodes running out | Inodes go above 90%. New files fail even while disk space is free. | Hourly |
Services
| Alert | Alerts when | Checks |
|---|---|---|
| Service stopped | A systemd service you name stops running. Can restart it for you. | Every 2 min |
| Failed systemd services | Any systemd unit enters the failed state. | Every 15 min |
| Server rebooted | The server restarted, planned or not. One alert per boot. | Every 5 min |
Website and SSL
| Alert | Alerts when | Checks |
|---|---|---|
| Website down | A URL stops returning 200, after 2 retries. | Every 2 min |
| SSL certificate expiring | A certificate is 14, 7, 3, and 1 days from expiry, with a reminder at each. | Daily at 09:00 |
Run Website down from a different server than the one hosting the site. If that server goes down, a check running on it goes down too.
Security
| Alert | Alerts when | Checks |
|---|---|---|
| SSH brute force | 20 or more failed SSH logins between checks. Includes the top source IPs. | Every 10 min |
Cron jobs
| Alert | Alerts when | Checks |
|---|---|---|
| Cron job fails | One of your own cron jobs fails, times out, or prints a pattern you pick. | Every run of the job |
Read the alerts list
Each alert is a card, and the card is meant to answer "am I covered?" without opening anything.
| On the card | What it tells you |
|---|---|
| Toggle | Whether the alert is running. Off on an alert you set up means paused |
| Limit | The limit you set, for example 85%, next to the name |
| Summary | What the alert watches, or once it is set up, your settings in one line such as CPU above 85% for 2 checks |
| Channel | An icon for where the message goes |
| Interval | How often the check runs, for example Every 5 min |
| Status | ON, PAUSED, or NOT SET UP, and FIRST CHECK SOON between switching an alert on and its first check |
| Add another | Adds the same alert again, at another limit or for another target |
Above the list, filter chips narrow things down:
| Chip | Shows |
|---|---|
| All | Every alert, set up or not |
| On | Alerts that are running |
| Paused | Alerts you set up and then switched off |
| Not set up | Alerts you have not turned on yet |
The chip next to the title counts how many alerts are on, and reads Nothing on yet until the first one is. The search box filters by name, and the toggle at the top right switches between card and table views.
Bad checks in a row and the recovery message
A single high reading is usually a spike rather than a problem: a deploy, a backup, a busy cron job. Bad checks in a row before alerting sets how many checks in a row must cross the limit before you are told. With High CPU usage checking every 5 minutes and the default of 2, a spike that clears before the next check stays quiet, and CPU pinned for about 10 minutes reaches you. Raise the number to hear less, lower it to hear sooner.
High CPU usage also sends one recovery message once CPU is back under the limit, so you know the problem is over without having to go and check.
Where alerts go
Alerts go to a channel: Slack, Telegram, or a webhook. Channels live on the Alert Channels tab of the Home screen, not inside an alert, and the same channel is reused by every alert and every cron job on every server. That tab has a Send test button, so confirm delivery there before you rely on a channel.
Message tokens and the rest of the channel settings are covered under Pick an alert channel in the Cron Jobs docs.
How alerts run on your server
Each alert you turn on is installed on the server as a cron job, which is what the Runs as cron jobs note in the banner refers to. Three things follow from that.
- Alerts arrive with CtrlOps closed. The server runs the check and sends the message itself, so nothing depends on the desktop app being open.
- They are listed in the Cron Jobs tab. Each alert shows up there as a job with its run history and output, which is where to look to see what a check found.
- They run as root. The runs as root note next to the title means the checks run with root privileges, so they can see every process, every filesystem, and the SSH login log.
To stop an alert, use its toggle in the Alerts tab rather than changing its job in Cron Jobs. The toggle pauses it and keeps its limit and channel with it.
Need a check that is not in the list?
Build it in Cron Jobs. Write the check as a cron job, a shell command or a saved script that exits with an error, or prints a word you choose, when something is wrong. For example, a job that fails when the newest backup is more than 26 hours old:
find /var/backups -name '*.tar.gz' -mmin -1560 | grep -q . || exit 1Then switch on Cron job fails for that job here. It fires when one of your own cron jobs fails, times out, or prints a pattern you pick. The same triggers are described under Add alerts to a cron job.
What Alerts does not do
- It keeps no metric history. An alert tells you a limit was crossed, not what CPU looked like last Tuesday. For live CPU, memory, and disk, open Infra Details.
- It does not escalate. There is no on-call rotation, acknowledgement, or paging schedule. The message goes to the channels you picked.
- It watches the server it runs on. A check cannot report its own server going offline, which is why Website down belongs on a different server, and why Server rebooted tells you once the server is back up.
Tips
On a new server, start with High CPU usage, Disk space low, and Service stopped for your app's main service. Those three catch most of what takes a small server down.
Switch on SSH brute force on any server with SSH open to the internet. The top source IPs in the message tell you who is knocking, and the SSH security guide covers what to do about it.
Send the urgent alerts somewhere you will actually see them. A 90% CPU alert in a busy channel is easy to miss, and a Telegram message to your phone is not.
Troubleshooting
Four usual causes.
- The limit has not held long enough. Check how many bad checks in a row the alert needs and how often it checks.
- The channel is misconfigured. Use Send test on the Alert Channels tab of the Home screen to confirm delivery on its own.
- The alert is paused. The card reads PAUSED, and the Paused chip lists it.
- Cron is not running on the server. Alerts run as cron jobs, so nothing fires without it.
Raise the limit, or raise Bad checks in a row before alerting so that a short spike stays quiet. If the noise comes from planned work, pause the alert for the duration and switch it back on afterwards.
The check was running on the server that went offline, so it went down with it. Turn Website down on from a different server, one that does not host the site.
Out-of-memory kill, Failed systemd services, and Server rebooted show On once they first run. Give it one check interval, then click Refresh.
That is expected. Every alert runs on the server as a cron job, which is what lets it fire with CtrlOps closed. Turn alerts on and off from the Alerts tab, and use the job's run history in Cron Jobs to see what a check found.
Frequently Asked Questions
How do I get an alert when CPU usage is high on my Linux server?
Open the server in CtrlOps, go to the Alerts tab, and switch on High CPU usage. Set the CPU limit, 85% by default, and how many bad checks in a row must pass before it alerts, 2 by default. Then pick the Slack, Telegram, or webhook channel to send it to. The check runs every 5 minutes, so with the defaults you hear about it once CPU has stayed above the limit for about 10 minutes.
Do CtrlOps alerts still fire when the app is closed?
Yes. Each alert is installed on the server as a cron job, so the server runs the check and sends the message itself. Close CtrlOps, or close your laptop, and the alert still reaches you.
Do I need to install a monitoring agent on the server?
No. There is no agent or daemon to install. Each alert is a check that cron runs on a schedule, set up over the SSH connection CtrlOps already has, and it is listed in the Cron Jobs tab like any other job.
Can I have two CPU alerts at different limits?
Yes. Click Add alert, pick High CPU usage again, and set the second limit, for example one alert at 75% and another at 90%. The same works for a second service, a second website, or another disk. Most alerts can be added more than once.
What happens when I turn an alert off?
It is paused, not deleted. The card keeps its limit and its channel and reads PAUSED, and the Paused chip above the list counts it. Flip the toggle again when you want to hear about it again.
Why does an alert wait for several bad checks before it fires?
One high reading is usually a spike rather than a problem, such as a deploy or a backup doing its work. Bad checks in a row before alerting sets how many checks in a row must cross the limit before you are told. With High CPU usage checking every 5 minutes and the default of 2, a short spike stays quiet and CPU pinned for about 10 minutes reaches you.
Which channels can alerts be sent to?
Slack, Telegram, or a webhook. Channels are saved once on the Alert Channels tab of the Home screen and reused by every alert and every cron job on every server. That tab has a Send test button, so you can confirm delivery before you rely on a channel.
Why do my alerts show up in the Cron Jobs tab?
Because that is how they run. Each alert you turn on is a cron job on the server, which is what lets it fire with CtrlOps closed, and the Cron Jobs tab lists it with its run history like any other job. Turn alerts on and off, and add new ones, from the Alerts tab.
Can I get an alert for something that is not in the list?
Yes, through Cron Jobs. Write the check as a cron job, a shell command or a saved script that exits with an error or prints a word you choose when something is wrong. Then switch on the Cron job fails alert for it, which fires when one of your own cron jobs fails, times out, or prints a pattern you pick.
Should I run Website down on the same server as the site?
No. Run it from a different server. A check running on the server that hosts the site goes down with it, so it can never tell you that server is offline.
Are alerts set up per server?
Yes. An alert runs on the server you turned it on for and watches that server. To watch several servers, open each one in CtrlOps and switch on its alerts there.
Backups
Schedule automatic Linux server backups to AWS S3 or Dropbox with CtrlOps. No scripts needed, configure, schedule, and restore in minutes.
Cron Jobs
The Cron Jobs tab schedules work on your server without cron syntax. Pick a frequency from a dropdown, point the job at a shell command, a saved script, or a URL, and get a Slack, Telegram, or webhook alert the moment a run fails. Every run is logged with its full output.