Photo by Nataliya Vaitkevich from Pexels

Every production platform running LLM APIs faces a paradox: too many alerts burn out your team and mask real incidents, but too few leave you blind to actual outages. This tension, between alert fatigue and missing critical degradation, is one of the hardest problems to solve in AI infrastructure monitoring.

The cost of getting it wrong is high. Miss a real outage because you've been trained to ignore noisy alerts, and users discover the problem before you do. Over-alert and your on-call team stops responding to pages, turning monitoring into security theater.

Key takeaway: Key Takeaway: Dynamic regional baselines and tier-based alerting reduce false positives by 40–60% while catching real incidents within 5 minutes. Aggregate metrics hide regional degradation; always drill down by region and provider.

TL;DR

  • Alert fatigue occurs when poorly tuned thresholds generate noise that drowns out real incidents.
  • Most teams alert on aggregate latency; regional and provider-specific degradation often hides in that average.
  • A tier-based alerting strategy separates signal from noise: urgent (page-worthy) vs informational (weekly summary).
  • Baseline comparison and regional variance are critical; a 200 ms spike in one region may be normal, in another it's an outage.
  • Observinio's 21-region probes and degradation alerts help distinguish real incidents from transient noise by comparing against historical baselines.

The Alert Fatigue Trap

server room
Photo by panumas nikhomkhai from Pexels
0%
Miss rate on critical events when receiving dozens of alerts daily

Alert fatigue is real. Studies across infrastructure monitoring show that teams responding to dozens of alerts per day have a 40–60% miss rate on critical events. When your monitoring system cries wolf hourly, people learn to ignore it. The classic culprit: poorly tuned static thresholds.

Consider a common scenario: your OpenAI API endpoint returns TTFB (time-to-first-byte) between 150–250 ms under normal load. Your SRE sets a threshold at 300 ms, a reasonable safety margin. But network jitter, endpoint congestion, and regional variance mean you hit 310 ms every few hours during peak usage. Within a week, your team has dismissed dozens of alerts. Then, one Tuesday at 11 PM, latency genuinely spikes to 800 ms. The alert fires. No one pages in. By the time the on-call engineer notices in the morning, your chat feature has been slow for six hours.

"A conveyor motor in Plant A running at 60% load behaves differently from an identical motor in Plant B running at 90% load."
>, The Best AI Alert Fatigue Reducer for Plants

The same principle applies to LLM API latency. A 300 ms response from OpenAI's us-east endpoint under 80% utilization is normal. The same 300 ms response from their eu-west endpoint under 20% utilization is a red flag, something is wrong. Most monitoring tools treat all 300 ms responses equally and miss this context.

Why Aggregate Metrics Hide Real Problems

Most teams monitor a single "average latency" metric across all regions and request types. This is fast to set up and looks good in dashboards. It's also nearly useless for catching real outages.

Here's why: imagine your OpenRouter requests average 180 ms globally. This average is the sum of:

  • us-west requests: 120 ms (5,000 per minute)
  • eu-central requests: 240 ms (2,000 per minute)
  • ap-southeast requests: 300 ms (500 per minute)
Now, eu-central degrades to 600 ms, but it only represents 25% of traffic. The global average moves to 228 ms, a 48 ms increase. If your threshold is 250 ms, you don't alert. Your European users are in pain, but the metric looks "okay." By the time you discover it (via a support ticket), it's been ongoing for 30 minutes.

Regional and provider-specific degradation is the norm for globally distributed services. Treating all traffic as one number guarantees you'll miss half your incidents.

Understanding the Missing-Incident Problem

network cables
Photo by Brett Sayles from Pexels

The flip side of alert fatigue is the missed incident. This happens in three ways:

  1. Blind spots in coverage. You're not monitoring OpenRouter's eu-central endpoint, so you don't know it's degraded. Your users in Germany do.
  1. Incorrect baselines. You set a hard threshold (e.g., 500 ms) without accounting for normal variance. On a slow Thursday, latency sits at 480 ms for six hours, it's not an outage, just natural fluctuation, but you're not alerting, so you don't notice when it actually jumps to 900 ms two hours later.
  1. Silent degradation. Latency drifts gradually from 180 ms to 250 ms to 320 ms over three hours. Each jump is within your threshold buffer, but the trend is a clear outage signal. Most teams don't trend-alert; they only alert on absolute thresholds.
For ML platform engineers and SREs, these gaps are costly. A regional latency spike goes unnoticed for 45 minutes. Your chat feature becomes sluggish. Users switch to a competitor's integration. You find out from the churn data, not from monitoring.

The Tier-Based Alerting Strategy

The solution is not to alert on everything or nothing, it's to tier your alerts by severity and audience.

Alert Tiers Explained

Tier 1: Page-Worthy (Urgent) These are genuine incidents requiring immediate human response. Criteria:
  • Latency spike ≥ 2× baseline in any region for ≥ 5 minutes.
  • Complete timeout or 5xx errors on ≥ 10% of requests.
  • Degradation affecting ≥ 2 regions simultaneously.
Action: Page on-call engineer immediately. Target: < 5-minute MTTD (mean time to detect). Tier 2: Investigate Soon (Warning) Anomalies worth looking at before they escalate. Criteria:
  • Latency increase 1.5–2× baseline in one region for ≥ 10 minutes.
  • Gradual trend of increasing p95 latency over 1 hour.
  • Single-region timeout rate > 2% but < 10%.
Action: Email alert to the team. Include in daily review. Target: < 30-minute response. Tier 3: Informational (Trending) Data for long-term decisions; not urgent. Criteria:
  • Weekly latency summary by region and provider.
  • Monthly comparison: this week vs. last week.
  • Anomalies that resolved within 2 minutes (transient jitter).
Action: Weekly digest. Include in postmortem discussions. Target: review once a week.

Your progress is saved automatically in your browser.

Building Baselines to Reduce False Positives

data center
Photo by panumas nikhomkhai from Pexels
Teams using dynamic baselines catching real incidents
0%

The key to reducing alert fatigue without missing incidents is dynamic baselining. Instead of alerting when latency > 500 ms, alert when latency is > (90th percentile of the last 7 days + 50 ms).

How to Build a Baseline

  1. Collect 7–14 days of clean data from each region and provider. Use Observinio's daily probes across your 21 regions to gather this.
  • Calculate rolling percentiles:
    • p50 (median)
    • p75 (typical "good" performance)
    • p95 (tail latency under normal load)
    • p99 (worst-case but still normal)
  • Set alert thresholds relative to p95:
    • Tier 1 alert: p95 × 1.8 (nearly 2× worst-case-normal)
    • Tier 2 alert: p95 × 1.3 (notable but not critical)
  1. Update baselines weekly. If your provider genuinely improves performance, your baseline should reflect it. If it degrades, you'll detect the trend.
  1. Alert on anomaly score, not absolute values. An anomaly score of (current_latency − baseline) / baseline tells you whether behavior has changed, regardless of absolute numbers.

Example: OpenAI us-east Region

  • Baseline p95: 180 ms (from last 7 days)
  • Tier 1 threshold: 180 × 1.8 = 324 ms
  • Tier 2 threshold: 180 × 1.3 = 234 ms
  • Current latency: 250 ms
  • Anomaly score: (250 − 180) / 180 = 0.39 (39% above baseline)
  • Action: Tier 2 alert fired. Check if regional load is high or if there's a provider issue.
This approach automatically adapts to what "normal" is for your traffic and region, eliminating the noise from static thresholds.

Regional Variance: The Hidden Signal

LLM API latency is not uniform across the globe. A 50 ms response from us-west might indicate a problem; the same from ap-southeast during peak hours is normal.

Observinio monitors from 21 regions, providing this perspective automatically. When you compare TTFB and TTFT across regions, patterns emerge:

  • If latency spikes in one region only → likely a regional provider issue or your routing logic.
  • If latency spikes everywhere → provider-wide degradation or your own stack.
  • If latency slowly drifts in one region over days → provider capacity planning or infrastructure changes.

Regional Alert Assignment Checklist

Practical Implementation: Build Your Alert Stack

Alert Stack Architecture

Collect baseline data → Calculate regional thresholds → Tier alerts by severity → Implement dynamic comparison → Monitor and iterate monthly.

Alert Fatigue vs Missing Real AI Outages process
Figure 1: Alert Fatigue vs Missing Real AI Outages at a glance.

Step-by-Step Setup

Step 1: Establish your baseline data source.
  • Use Observinio's daily synthetic probes across OpenRouter and OpenAI endpoints.
  • Collect 7–14 days of data before setting thresholds.
  • Export TTFB and TTFT by region.
Step 2: Calculate regional baselines.
  • For each region, compute p75, p95, and p99 latency.
  • Document normal variance (e.g., "us-west TTFB p95 = 150 ms, acceptable range 130–180 ms").
  • Note any daily or hourly patterns (peak hours are slower).
Step 3: Define Tier 1 and Tier 2 thresholds.
  • Tier 1: p95 × 1.8 for each region.
  • Tier 2: p95 × 1.3 for each region.
  • Document why these thresholds matter to your business (e.g., "Tier 1 in us-west fires only if behavior is truly abnormal; 95% of users are in us-west").
Step 4: Configure alerting rules.
  • Set up alerts in your monitoring tool (Datadog, New Relic, Prometheus, etc.) to compare current latency against rolling 7-day baseline.
  • Use region-specific alert channels (e.g., page only if us-west is affected; email for ap-southeast).
  • Include context in the alert: current latency, baseline, anomaly score, affected region(s).
Step 5: Implement Observinio weekly summaries.
  • Enable weekly email digests from Observinio showing regional latency trends, provider comparisons, and anomalies.
  • Use this data in your weekly review to spot gradual degradation (Tier 3 signal).
  • Adjust baselines if you spot long-term provider changes.
Step 6: Test your alert response.
  • Simulate a 2× latency spike in one region.
  • Confirm Tier 1 alert fires.
  • Simulate a 1.5× spike.
  • Confirm Tier 2 alert fires, but does not page.
  • Measure time from alert to on-call response. Aim for < 5 minutes for Tier 1.
Step 7: Audit false positives monthly.
  • Track which alerts were dismissed as false positives.
  • If > 20% of Tier 1 alerts are false positives, increase threshold or refine baseline calculation.
  • If > 40% of Tier 2 alerts are false positives, demote them to Tier 3 (weekly digest only).

Real-World Example: Catching a Regional Outage

A mid-market AI chat platform runs on OpenRouter with users in North America, Europe, and Asia. Their baseline latency is:

  • us-west: p95 = 160 ms
  • eu-central: p95 = 200 ms
  • ap-southeast: p95 = 280 ms
One Tuesday, Observinio's daily probes detect:
  • us-west: 155 ms (normal)
  • eu-central: 450 ms (+125% vs. baseline)
  • ap-southeast: 285 ms (normal)
Tier 1 alert fires for eu-central because 450 ms > (200 × 1.8 = 360 ms). The on-call engineer is paged. Within 2 minutes, she reviews Observinio's status page and sees that OpenRouter's eu-central endpoint is experiencing a capacity issue (visible in the degradation alert and regional drill-down). She decides to route European traffic to eu-west temporarily. MTTD: 2 minutes. User impact: minimal because chat requests complete within acceptable latency even on the alternate endpoint.

If the team had not been using dynamic regional baselines and Tier 1 alerting, they would have noticed the incident only when support tickets arrived 30 minutes later.

FAQ

Frequently Asked Questions

Update baselines weekly if your traffic patterns are stable, or more frequently if you're scaling rapidly or testing new endpoints. Observinio's weekly summaries provide the data you need; integrate them into your Monday morning review. If you spot a sustained change in average latency (provider improvement or degradation), recalculate thresholds immediately.
High variance is a signal to investigate first. High variance often means you're mixing requests of different types (e.g., long-context completions with short queries) or regions with different load profiles. If variance is genuinely due to normal load fluctuation, increase your threshold to p95 × 2.0 (not 1.8), but document why and review monthly. If it's due to misrouting or misconfiguration, fix the root cause first, then baseline again.
Both. TTFB (time-to-first-byte) is your detection latency, how quickly the API responds. TTFT (time-to-first-token) includes the first token generation time and is more user-visible. Alert on TTFB for API responsiveness and on TTFT for user experience. Observinio tracks both, so use the regional breakdown to set alerts on the metric that matters most for your use case.
Observinio's degradation alerts are excellent for detecting regional and provider-level issues because they compare against 21-region baselines. You should use them as your primary alert for external API latency. However, complement them with your own monitoring of your own infrastructure (routing, caching, application latency). Together, they give you full visibility.
Missing a 30-minute outage in one region typically results in 5–15% of traffic being slow, translating to 1–3% churn in chat features and support tickets. Over-alerting by more than 20% of alerts being false positives reduces on-call response time by ~40% (people stop trusting pages). The cost of false positives compounds; the cost of missed incidents is immediate. Optimize for catching real incidents first, then tune down false positives.

Keep Your Team Alert Without Alert Burnout

Alert fatigue and missed incidents are not two separate problems, they're two sides of the same coin: poor signal-to-noise ratio. The fix is not to alert more or less, but to alert smarter: tier by severity, baseline against historical data, and break down by region and provider.

Observinio's 21-region probes and weekly summaries give you the data foundation you need to build this kind of alerting. Start by collecting two weeks of baseline data, set regional thresholds using dynamic baselines, and tier your alerts. Your on-call team will thank you, and your users will feel the difference when you catch incidents minutes instead of hours after they start.

Ready to reduce alert noise without missing the next outage? Check out Observinio's degradation alerts and status page integration to see how other ML teams are solving this exact problem.

Additional Resources