Photo by Nataliya Vaitkevich from Pexels
Every production platform running LLM APIs faces a paradox: too many alerts burn out your team and mask real incidents, but too few leave you blind to actual outages. This tension, between alert fatigue and missing critical degradation, is one of the hardest problems to solve in AI infrastructure monitoring.
The cost of getting it wrong is high. Miss a real outage because you've been trained to ignore noisy alerts, and users discover the problem before you do. Over-alert and your on-call team stops responding to pages, turning monitoring into security theater.
TL;DR
- Alert fatigue occurs when poorly tuned thresholds generate noise that drowns out real incidents.
- Most teams alert on aggregate latency; regional and provider-specific degradation often hides in that average.
- A tier-based alerting strategy separates signal from noise: urgent (page-worthy) vs informational (weekly summary).
- Baseline comparison and regional variance are critical; a 200 ms spike in one region may be normal, in another it's an outage.
- Observinio's 21-region probes and degradation alerts help distinguish real incidents from transient noise by comparing against historical baselines.
The Alert Fatigue Trap
Alert fatigue is real. Studies across infrastructure monitoring show that teams responding to dozens of alerts per day have a 40–60% miss rate on critical events. When your monitoring system cries wolf hourly, people learn to ignore it. The classic culprit: poorly tuned static thresholds.
Consider a common scenario: your OpenAI API endpoint returns TTFB (time-to-first-byte) between 150–250 ms under normal load. Your SRE sets a threshold at 300 ms, a reasonable safety margin. But network jitter, endpoint congestion, and regional variance mean you hit 310 ms every few hours during peak usage. Within a week, your team has dismissed dozens of alerts. Then, one Tuesday at 11 PM, latency genuinely spikes to 800 ms. The alert fires. No one pages in. By the time the on-call engineer notices in the morning, your chat feature has been slow for six hours.
"A conveyor motor in Plant A running at 60% load behaves differently from an identical motor in Plant B running at 90% load.">, The Best AI Alert Fatigue Reducer for Plants
The same principle applies to LLM API latency. A 300 ms response from OpenAI's us-east endpoint under 80% utilization is normal. The same 300 ms response from their eu-west endpoint under 20% utilization is a red flag, something is wrong. Most monitoring tools treat all 300 ms responses equally and miss this context.
Why Aggregate Metrics Hide Real Problems
Most teams monitor a single "average latency" metric across all regions and request types. This is fast to set up and looks good in dashboards. It's also nearly useless for catching real outages.
Here's why: imagine your OpenRouter requests average 180 ms globally. This average is the sum of:
- us-west requests: 120 ms (5,000 per minute)
- eu-central requests: 240 ms (2,000 per minute)
- ap-southeast requests: 300 ms (500 per minute)
Regional and provider-specific degradation is the norm for globally distributed services. Treating all traffic as one number guarantees you'll miss half your incidents.
Understanding the Missing-Incident Problem
The flip side of alert fatigue is the missed incident. This happens in three ways:
- Blind spots in coverage. You're not monitoring OpenRouter's eu-central endpoint, so you don't know it's degraded. Your users in Germany do.
- Incorrect baselines. You set a hard threshold (e.g., 500 ms) without accounting for normal variance. On a slow Thursday, latency sits at 480 ms for six hours, it's not an outage, just natural fluctuation, but you're not alerting, so you don't notice when it actually jumps to 900 ms two hours later.
- Silent degradation. Latency drifts gradually from 180 ms to 250 ms to 320 ms over three hours. Each jump is within your threshold buffer, but the trend is a clear outage signal. Most teams don't trend-alert; they only alert on absolute thresholds.
The Tier-Based Alerting Strategy
The solution is not to alert on everything or nothing, it's to tier your alerts by severity and audience.
Alert Tiers Explained
Tier 1: Page-Worthy (Urgent) These are genuine incidents requiring immediate human response. Criteria:- Latency spike ≥ 2× baseline in any region for ≥ 5 minutes.
- Complete timeout or 5xx errors on ≥ 10% of requests.
- Degradation affecting ≥ 2 regions simultaneously.
- Latency increase 1.5–2× baseline in one region for ≥ 10 minutes.
- Gradual trend of increasing p95 latency over 1 hour.
- Single-region timeout rate > 2% but < 10%.
- Weekly latency summary by region and provider.
- Monthly comparison: this week vs. last week.
- Anomalies that resolved within 2 minutes (transient jitter).
Your progress is saved automatically in your browser.
Building Baselines to Reduce False Positives
The key to reducing alert fatigue without missing incidents is dynamic baselining. Instead of alerting when latency > 500 ms, alert when latency is > (90th percentile of the last 7 days + 50 ms).
How to Build a Baseline
- Collect 7–14 days of clean data from each region and provider. Use Observinio's daily probes across your 21 regions to gather this.
- Calculate rolling percentiles:
- p50 (median)
- p75 (typical "good" performance)
- p95 (tail latency under normal load)
- p99 (worst-case but still normal)
- Set alert thresholds relative to p95:
- Tier 1 alert: p95 × 1.8 (nearly 2× worst-case-normal)
- Tier 2 alert: p95 × 1.3 (notable but not critical)
- Update baselines weekly. If your provider genuinely improves performance, your baseline should reflect it. If it degrades, you'll detect the trend.
- Alert on anomaly score, not absolute values. An anomaly score of (current_latency − baseline) / baseline tells you whether behavior has changed, regardless of absolute numbers.
Example: OpenAI us-east Region
- Baseline p95: 180 ms (from last 7 days)
- Tier 1 threshold: 180 × 1.8 = 324 ms
- Tier 2 threshold: 180 × 1.3 = 234 ms
- Current latency: 250 ms
- Anomaly score: (250 − 180) / 180 = 0.39 (39% above baseline)
- Action: Tier 2 alert fired. Check if regional load is high or if there's a provider issue.
Regional Variance: The Hidden Signal
LLM API latency is not uniform across the globe. A 50 ms response from us-west might indicate a problem; the same from ap-southeast during peak hours is normal.
Observinio monitors from 21 regions, providing this perspective automatically. When you compare TTFB and TTFT across regions, patterns emerge:
- If latency spikes in one region only → likely a regional provider issue or your routing logic.
- If latency spikes everywhere → provider-wide degradation or your own stack.
- If latency slowly drifts in one region over days → provider capacity planning or infrastructure changes.
Regional Alert Assignment Checklist
Practical Implementation: Build Your Alert Stack
Alert Stack Architecture
Collect baseline data → Calculate regional thresholds → Tier alerts by severity → Implement dynamic comparison → Monitor and iterate monthly.
Step-by-Step Setup
Step 1: Establish your baseline data source.- Use Observinio's daily synthetic probes across OpenRouter and OpenAI endpoints.
- Collect 7–14 days of data before setting thresholds.
- Export TTFB and TTFT by region.
- For each region, compute p75, p95, and p99 latency.
- Document normal variance (e.g., "us-west TTFB p95 = 150 ms, acceptable range 130–180 ms").
- Note any daily or hourly patterns (peak hours are slower).
- Tier 1: p95 × 1.8 for each region.
- Tier 2: p95 × 1.3 for each region.
- Document why these thresholds matter to your business (e.g., "Tier 1 in us-west fires only if behavior is truly abnormal; 95% of users are in us-west").
- Set up alerts in your monitoring tool (Datadog, New Relic, Prometheus, etc.) to compare current latency against rolling 7-day baseline.
- Use region-specific alert channels (e.g., page only if us-west is affected; email for ap-southeast).
- Include context in the alert: current latency, baseline, anomaly score, affected region(s).
- Enable weekly email digests from Observinio showing regional latency trends, provider comparisons, and anomalies.
- Use this data in your weekly review to spot gradual degradation (Tier 3 signal).
- Adjust baselines if you spot long-term provider changes.
- Simulate a 2× latency spike in one region.
- Confirm Tier 1 alert fires.
- Simulate a 1.5× spike.
- Confirm Tier 2 alert fires, but does not page.
- Measure time from alert to on-call response. Aim for < 5 minutes for Tier 1.
- Track which alerts were dismissed as false positives.
- If > 20% of Tier 1 alerts are false positives, increase threshold or refine baseline calculation.
- If > 40% of Tier 2 alerts are false positives, demote them to Tier 3 (weekly digest only).
Real-World Example: Catching a Regional Outage
A mid-market AI chat platform runs on OpenRouter with users in North America, Europe, and Asia. Their baseline latency is:
- us-west: p95 = 160 ms
- eu-central: p95 = 200 ms
- ap-southeast: p95 = 280 ms
- us-west: 155 ms (normal)
- eu-central: 450 ms (+125% vs. baseline)
- ap-southeast: 285 ms (normal)
If the team had not been using dynamic regional baselines and Tier 1 alerting, they would have noticed the incident only when support tickets arrived 30 minutes later.
FAQ
Frequently Asked Questions
Keep Your Team Alert Without Alert Burnout
Alert fatigue and missed incidents are not two separate problems, they're two sides of the same coin: poor signal-to-noise ratio. The fix is not to alert more or less, but to alert smarter: tier by severity, baseline against historical data, and break down by region and provider.
Observinio's 21-region probes and weekly summaries give you the data foundation you need to build this kind of alerting. Start by collecting two weeks of baseline data, set regional thresholds using dynamic baselines, and tier your alerts. Your on-call team will thank you, and your users will feel the difference when you catch incidents minutes instead of hours after they start.
Ready to reduce alert noise without missing the next outage? Check out Observinio's degradation alerts and status page integration to see how other ML teams are solving this exact problem.
Additional Resources
- The Best AI Alert Fatigue Reducer for Plants - Alert fatigue costs plants real downtime. See how an AI alert fatigue reducer filters noise, ranks by impact, and delivers alerts your team ...
- Avoid Alert Fatigue: Prioritize Critical Alerts - Alert fatigue is caused by poor prioritization, not tools, and too many alerts reduce visibility. Excessive alerts without context turn into ...
- What Is Alert Fatigue? | IBM - Alert fatigue is a state of mental and operational exhaustion caused by an overwhelming number of alerts—many of which are low-priority, ...
