CtrlOps
|Docs
Product Modules

Alerts: Know When CPU, Disk, or a Service Needs You

The Alerts tab watches a Linux server for you. Switch on ready-made alerts for high CPU, low disk, a stopped service, a site going down, an expiring SSL certificate, or SSH brute force, and CtrlOps messages your Slack, Telegram, or webhook channel when one trips, even while the app is closed.

Most server problems show up well before they turn into outages. CPU pinned for ten minutes, a disk at 85% and climbing, a service that quietly stopped, a certificate with a week left. The trouble is that nobody is looking when it happens, so you hear about it from a user, or from the error that takes the site down. Doing it yourself means a check script per problem, a crontab line per script, and some way to send the message, on every server.

The Alerts tab gives you those checks ready-made. Every alert is listed with a toggle: switch on the ones you want to hear about, set a limit if the alert has one, pick where the message goes, and it is live. Each alert runs on the server on a schedule and messages your Slack, Telegram, or webhook channel when it trips, even while CtrlOps is closed. For a check that is not in the list, write it as a job in Cron Jobs and alert on that.

What you can do in the Alerts tab

  • Switch on 13 ready-made alerts covering CPU, memory, swap, out-of-memory kills, disk space, inodes, services, reboots, websites, SSL certificates, SSH brute force, and your own cron jobs.
  • Set the limit for an alert, for example CPU at 85%, and how many bad checks in a row it takes before you are told.
  • Read a preview of the message before you turn the alert on.
  • Send alerts to Slack, Telegram, or a webhook, using channels you set up once on the Home screen.
  • Add the same alert more than once, for another service, site, or disk, or at another limit.
  • Pause an alert with its toggle and keep its limit and channel for later.
  • Narrow the list to what is on, paused, or not set up yet, or search it by name.
  • Hear about problems while CtrlOps is closed, because the checks run on the server itself.

Why an Alerts tab instead of writing your own checks

Doing it by handDoing it in CtrlOps
Write a script that reads /proc/loadavg or parses topSwitch on High CPU usage
Decide on a threshold and how long it must hold before it countsSet CPU limit and Bad checks in a row
Add a crontab line for the script on every serverThe alert is installed as a cron job for you
Wire up curl to a Slack webhook or the Telegram bot APIPick a channel you saved once
Send yourself a test to see how the message readsRead What the message looks like before you turn it on
Grep auth.log for failed logins now and thenSwitch on SSH brute force
Remember which checks run on which serverEvery alert on the server is listed, with what is on, paused, or not set up
Comment the line out to silence a checkFlip the toggle to pause it

Before you start

RequirementWhy you need itWhat to do if it is missing
An active server connectionAlerts are installed on the server you are connected toOpen the server from the Home page first
Root or sudo on the serverAlerts run as root, so a check can see every process, filesystem, and loginConnect as root, or as a user with sudo
Cron on the serverEach alert runs as a cron jobNearly every Linux distribution ships it. If it is not running, start the cron service
An alert channelIt is where the message goesSet one up once on the Alert Channels tab of the Home screen

Alerts are per server. Switching on High CPU usage here watches this server only. To watch several servers, open each one and switch its alerts on there.

Open the Alerts tab

Connect to your server

Open CtrlOps and click the server you want to watch. Wait for the header to show Connected.

Click the Alerts tab

In the left sidebar, click Alerts. It sits below Audit Report, just above Cron Jobs.

Alerts sits below Audit Report in the left sidebar and opens on the full list, grouped by category.

See what is already set up

Alerts are grouped by category, each card with its toggle, how often it checks, and its status. On a server where nothing is set up yet, every card reads NOT SET UP and the chip next to the title says Nothing on yet.

The banner under the title sums up how the tab works: each alert runs on this server on a schedule and messages your channels when it trips, even while CtrlOps is closed. Refresh at the top right reloads the list, and Add alert next to it adds an alert, including a second copy of one you already use.

Turn on an alert

Flip the toggle

Click the toggle on the alert card

Flip the toggle on the alert you want to hear about, for example High CPU usage. A Turn on alert panel opens with two steps, Configure and Channels, and a line on what the alert watches for.

Set the limit

Configure asks for whatever this alert needs. For High CPU usage that is two fields:

FieldDefaultWhat it means
CPU limit85%The CPU usage that counts as a bad check
Bad checks in a row before alerting2How many checks in a row must be over the limit before you are told

Under the fields, a note shows how often the check runs and what your settings add up to, for example Checks every 5 min. Alerts once it stays above the limit for about 10 minutes. What the message looks like previews the message itself, with this server's name in it. Click Continue.

Configure shows how often the check runs and exactly what the message will say before you turn it on.

Pick where the alert goes

In the Channels step, choose which of your saved Slack, Telegram, or webhook channels should get the message, then turn the alert on. The card now shows your limit, a one-line summary, and the channel.

Once an alert is set up, its card says what you chose: the limit next to the name, a summary such as CPU above 85% for 2 checks, an icon for the channel it messages, and how often it checks. Until the first check runs, its status reads FIRST CHECK SOON.

Three alerts have nothing to set, because they watch for an event rather than a limit: Out-of-memory kill, Failed systemd services, and Server rebooted. They show On once they first run.

Turn an alert off

Flip the same toggle off. The alert is paused, not deleted: the card keeps its limit and its channel, reads PAUSED, and the Paused chip above the list counts it. Flip the toggle back on when you want to hear about it again.

The same toggle switches an alert on and off. Off pauses it and keeps its limit and channel.

Pause an alert during planned work instead of learning to ignore it. A long migration or a load test trips High CPU usage on purpose, so pause it, do the work, and switch it back on.

Add an alert more than once

Most alerts can be added again, for another service, site, or disk, or at another limit. Two CPU alerts, one at 75% and one at 90%, give you an early warning and an urgent one. Two Service stopped alerts watch two services.

Click Add alert

Click Add alert (top right of the Alerts tab)

The Add alert panel lists every alert by category, with a search box at the top and how often each one checks.

Add alert sits at the top right, next to Refresh.

Pick the alert

Pick the alert to add again, for example High CPU usage. It opens the same Turn on alert panel as the toggle.

Set the new limit and its channels

Set the second limit, service, site, or disk, choose its channels, and turn it on. It runs alongside the first one.

On a card that is already set up, Add another is the shortcut for adding that same alert again.

The alert catalog

Thirteen alerts in six groups. Checks is how often the check runs on the server, and the limits shown are the defaults.

System resources

AlertAlerts whenChecks
High CPU usageCPU stays above 85% for about 10 minutes. Includes the top processes.Every 5 min
High memory usageMemory stays above 90% for about 10 minutes. Includes the biggest processes.Every 5 min
Swap usage highSwap stays above 80%, the early sign that RAM has run out.Every 5 min
Out-of-memory killThe kernel killed a process because memory ran out.Every 5 min

Storage

AlertAlerts whenChecks
Disk space lowA filesystem goes above 85%. Includes the largest folders.Every 30 min
Inodes running outInodes go above 90%. New files fail even while disk space is free.Hourly

Services

AlertAlerts whenChecks
Service stoppedA systemd service you name stops running. Can restart it for you.Every 2 min
Failed systemd servicesAny systemd unit enters the failed state.Every 15 min
Server rebootedThe server restarted, planned or not. One alert per boot.Every 5 min

Website and SSL

AlertAlerts whenChecks
Website downA URL stops returning 200, after 2 retries.Every 2 min
SSL certificate expiringA certificate is 14, 7, 3, and 1 days from expiry, with a reminder at each.Daily at 09:00

Run Website down from a different server than the one hosting the site. If that server goes down, a check running on it goes down too.

Security

AlertAlerts whenChecks
SSH brute force20 or more failed SSH logins between checks. Includes the top source IPs.Every 10 min

Cron jobs

AlertAlerts whenChecks
Cron job failsOne of your own cron jobs fails, times out, or prints a pattern you pick.Every run of the job

Read the alerts list

Each alert is a card, and the card is meant to answer "am I covered?" without opening anything.

On the cardWhat it tells you
ToggleWhether the alert is running. Off on an alert you set up means paused
LimitThe limit you set, for example 85%, next to the name
SummaryWhat the alert watches, or once it is set up, your settings in one line such as CPU above 85% for 2 checks
ChannelAn icon for where the message goes
IntervalHow often the check runs, for example Every 5 min
StatusON, PAUSED, or NOT SET UP, and FIRST CHECK SOON between switching an alert on and its first check
Add anotherAdds the same alert again, at another limit or for another target

Above the list, filter chips narrow things down:

ChipShows
AllEvery alert, set up or not
OnAlerts that are running
PausedAlerts you set up and then switched off
Not set upAlerts you have not turned on yet

The chip next to the title counts how many alerts are on, and reads Nothing on yet until the first one is. The search box filters by name, and the toggle at the top right switches between card and table views.

Bad checks in a row and the recovery message

A single high reading is usually a spike rather than a problem: a deploy, a backup, a busy cron job. Bad checks in a row before alerting sets how many checks in a row must cross the limit before you are told. With High CPU usage checking every 5 minutes and the default of 2, a spike that clears before the next check stays quiet, and CPU pinned for about 10 minutes reaches you. Raise the number to hear less, lower it to hear sooner.

High CPU usage also sends one recovery message once CPU is back under the limit, so you know the problem is over without having to go and check.

Where alerts go

Alerts go to a channel: Slack, Telegram, or a webhook. Channels live on the Alert Channels tab of the Home screen, not inside an alert, and the same channel is reused by every alert and every cron job on every server. That tab has a Send test button, so confirm delivery there before you rely on a channel.

Message tokens and the rest of the channel settings are covered under Pick an alert channel in the Cron Jobs docs.

How alerts run on your server

Each alert you turn on is installed on the server as a cron job, which is what the Runs as cron jobs note in the banner refers to. Three things follow from that.

  • Alerts arrive with CtrlOps closed. The server runs the check and sends the message itself, so nothing depends on the desktop app being open.
  • They are listed in the Cron Jobs tab. Each alert shows up there as a job with its run history and output, which is where to look to see what a check found.
  • They run as root. The runs as root note next to the title means the checks run with root privileges, so they can see every process, every filesystem, and the SSH login log.

To stop an alert, use its toggle in the Alerts tab rather than changing its job in Cron Jobs. The toggle pauses it and keeps its limit and channel with it.

Need a check that is not in the list?

Build it in Cron Jobs. Write the check as a cron job, a shell command or a saved script that exits with an error, or prints a word you choose, when something is wrong. For example, a job that fails when the newest backup is more than 26 hours old:

find /var/backups -name '*.tar.gz' -mmin -1560 | grep -q . || exit 1

Then switch on Cron job fails for that job here. It fires when one of your own cron jobs fails, times out, or prints a pattern you pick. The same triggers are described under Add alerts to a cron job.

What Alerts does not do

  • It keeps no metric history. An alert tells you a limit was crossed, not what CPU looked like last Tuesday. For live CPU, memory, and disk, open Infra Details.
  • It does not escalate. There is no on-call rotation, acknowledgement, or paging schedule. The message goes to the channels you picked.
  • It watches the server it runs on. A check cannot report its own server going offline, which is why Website down belongs on a different server, and why Server rebooted tells you once the server is back up.

Tips

On a new server, start with High CPU usage, Disk space low, and Service stopped for your app's main service. Those three catch most of what takes a small server down.

Switch on SSH brute force on any server with SSH open to the internet. The top source IPs in the message tell you who is knocking, and the SSH security guide covers what to do about it.

Send the urgent alerts somewhere you will actually see them. A 90% CPU alert in a busy channel is easy to miss, and a Telegram message to your phone is not.

Troubleshooting

Four usual causes.

  • The limit has not held long enough. Check how many bad checks in a row the alert needs and how often it checks.
  • The channel is misconfigured. Use Send test on the Alert Channels tab of the Home screen to confirm delivery on its own.
  • The alert is paused. The card reads PAUSED, and the Paused chip lists it.
  • Cron is not running on the server. Alerts run as cron jobs, so nothing fires without it.

Raise the limit, or raise Bad checks in a row before alerting so that a short spike stays quiet. If the noise comes from planned work, pause the alert for the duration and switch it back on afterwards.

The check was running on the server that went offline, so it went down with it. Turn Website down on from a different server, one that does not host the site.

Out-of-memory kill, Failed systemd services, and Server rebooted show On once they first run. Give it one check interval, then click Refresh.

That is expected. Every alert runs on the server as a cron job, which is what lets it fire with CtrlOps closed. Turn alerts on and off from the Alerts tab, and use the job's run history in Cron Jobs to see what a check found.

Frequently Asked Questions

How do I get an alert when CPU usage is high on my Linux server?

Open the server in CtrlOps, go to the Alerts tab, and switch on High CPU usage. Set the CPU limit, 85% by default, and how many bad checks in a row must pass before it alerts, 2 by default. Then pick the Slack, Telegram, or webhook channel to send it to. The check runs every 5 minutes, so with the defaults you hear about it once CPU has stayed above the limit for about 10 minutes.

Do CtrlOps alerts still fire when the app is closed?

Yes. Each alert is installed on the server as a cron job, so the server runs the check and sends the message itself. Close CtrlOps, or close your laptop, and the alert still reaches you.

Do I need to install a monitoring agent on the server?

No. There is no agent or daemon to install. Each alert is a check that cron runs on a schedule, set up over the SSH connection CtrlOps already has, and it is listed in the Cron Jobs tab like any other job.

Can I have two CPU alerts at different limits?

Yes. Click Add alert, pick High CPU usage again, and set the second limit, for example one alert at 75% and another at 90%. The same works for a second service, a second website, or another disk. Most alerts can be added more than once.

What happens when I turn an alert off?

It is paused, not deleted. The card keeps its limit and its channel and reads PAUSED, and the Paused chip above the list counts it. Flip the toggle again when you want to hear about it again.

Why does an alert wait for several bad checks before it fires?

One high reading is usually a spike rather than a problem, such as a deploy or a backup doing its work. Bad checks in a row before alerting sets how many checks in a row must cross the limit before you are told. With High CPU usage checking every 5 minutes and the default of 2, a short spike stays quiet and CPU pinned for about 10 minutes reaches you.

Which channels can alerts be sent to?

Slack, Telegram, or a webhook. Channels are saved once on the Alert Channels tab of the Home screen and reused by every alert and every cron job on every server. That tab has a Send test button, so you can confirm delivery before you rely on a channel.

Why do my alerts show up in the Cron Jobs tab?

Because that is how they run. Each alert you turn on is a cron job on the server, which is what lets it fire with CtrlOps closed, and the Cron Jobs tab lists it with its run history like any other job. Turn alerts on and off, and add new ones, from the Alerts tab.

Can I get an alert for something that is not in the list?

Yes, through Cron Jobs. Write the check as a cron job, a shell command or a saved script that exits with an error or prints a word you choose when something is wrong. Then switch on the Cron job fails alert for it, which fires when one of your own cron jobs fails, times out, or prints a pattern you pick.

Should I run Website down on the same server as the site?

No. Run it from a different server. A check running on the server that hosts the site goes down with it, so it can never tell you that server is offline.

Are alerts set up per server?

Yes. An alert runs on the server you turned it on for and watches that server. To watch several servers, open each one in CtrlOps and switch on its alerts there.