Reducing Alert Noise: Thresholds and Maintenance Modes

A monitoring system that sends fifty emails a day is not protecting anyone. Within a few weeks, the team creates a mail rule, mutes the channel or stops looking, and the one alert that matters is buried among the rest. This is alert fatigue, and it is one of the most common ways monitoring fails.

The fix is not to monitor less. It is to tune your monitoring so every alert deserves attention. This post offers practical steps for doing that.

Start by Looking at What You Get

Export or review a few weeks of alerts and sort them. Find:

The top ten noisiest alerts by count.

Alerts that routinely clear on their own within minutes.

Alerts nobody ever acted on.

Alerts that repeat for the same underlying problem.

Alerts that fire at predictable times, such as during backups or updates.

This list is your tuning to-do list. Most of the noise usually comes from a small number of sources.

Apply a Simple Test to Every Alert

For each alert, ask:

Does someone need to act on this? If not, it is a report, not an alert.

Does it need action now? If it can wait until the next business day, do not wake anyone up.

Is it clear what to do? An alert with no obvious response needs a runbook or should be removed.

Is it accurate? If it is wrong more often than right, fix it or turn it off.

Tune Your Thresholds

Use realistic limits

A server at 85 percent CPU for ten seconds is not a problem. Use values that reflect actual impact, and tune them to how your systems behave. Review thresholds when equipment or workloads change.

Require the condition to persist

Instead of alerting the instant a number crosses a line, alert only if it stays across the line for a set time, such as ten or fifteen minutes. This eliminates momentary spikes and brief network blips.

Use severity levels

Separate urgent issues from informational ones:

Critical: service down or resident-impacting. Notify on-call immediately.

Warning: trending toward a problem. Send to a queue or daily summary.

Informational: record for reports, no notification.

Alert on symptoms as well as causes

An alert that users cannot reach the EHR is often more useful than ten alerts about components behind it. Keep the cause-level data available for troubleshooting but notify on what matters.

Use Maintenance Modes

When you plan work, such as patching or hardware changes, monitoring will see systems reboot and services stop. Without preparation, it floods the team with false alarms.

Most monitoring tools allow a maintenance mode, sometimes called downtime or a silence window. Use it to suppress alerts for specific systems during a planned period. Practical guidelines:

Set an end time so alerts resume automatically and do not stay muted by mistake.

Limit the scope to the affected systems, not everything.

Record the reason and who set it.

Tie it to your published maintenance calendar.

Check afterward that systems came back healthy.

Group and Deduplicate

If a switch fails, every device behind it will appear to fail too. Configure dependencies, so alerts for downstream devices are suppressed when their upstream device is down. Group related alerts into one incident instead of dozens of messages, and avoid sending repeat notifications for the same unresolved problem more often than needed.

Send Alerts to the Right People

Route alerts by system and owner. The network team does not need printer warnings, and the on-call technician does not need a weekly disk report. Use separate channels for urgent and nonurgent messages, and make urgent ones hard to miss.

Review Regularly

Schedule a short monthly review. Look at alert counts, the noisiest sources and any missed incidents. When an outage occurs without an alert, add one. When an alert gets ignored, change it or remove it. Treat every alert as something you have to justify.

Measure Progress

Track alerts per week and the share that led to action. A falling count with the same or better detection is the goal. Ask the on-call team whether the alerts feel trustworthy.

Help With Monitoring

UnityCare IT's managed services include monitoring tuned for healthcare environments, with thresholds, maintenance modes and routing set up to reduce noise. If your team has stopped trusting its alerts, we can help you rebuild that trust.

Related service

An outsourced IT department with proactive maintenance and one number to call.

Related articles

Keep reading

Contact UnityCare Technologies

Call or text: 405-285-3845

New customers: start@unitycareit.com

Existing customers: support@unitycareit.com

Address: UnityCare Technologies, 2524 N Broadway Ste 554, PMB 947974, Edmond, Oklahoma 73034-4172