Photo by Brett Sayles from Pexels

Setting up synthetic probes for OpenRouter is the practical first step toward understanding real-world latency across regions. Unlike error rates or throughput metrics, latency directly shapes user experience, especially for chat and reasoning workloads where time-to-first-byte (TTFB) and time-to-full-text (TTFT) matter most. This worksheet walks you through configuring OpenRouter probes in Observinio so you can measure response times from 21 global regions, catch regional slowdowns before they affect users, and make provider decisions backed by data.

0regions
Global monitoring coverage
Configuration completion rate
0%
Key takeaway: Synthetic probes eliminate guesswork from provider selection by measuring real latency from actual geographic endpoints. Even a single region outage can be caught minutes before users report it, but only if your probes are configured with realistic thresholds matched to your actual production models and baseline performance.

TL;DR

  • OpenRouter probes measure TTFB and TTFT from 21 regions simultaneously; set them up via Observinio's configuration UI or API.
  • Route traffic through 2–4 representative models (e.g., Claude 3.5 Sonnet, Grok 2) to avoid skewed averages from outlier slow models.
  • Configure probes to run every 5–15 minutes depending on risk tolerance; set TTFB alerts at 1.5–2× your baseline and TTFT alerts per model.
  • Use weekly summaries and degradation alerts to catch regional variance before support tickets arrive.
  • Document your baseline numbers (by region and model) so alerts have context; reference this worksheet when investigating incidents.
Key takeaway: Synthetic probes eliminate guesswork from provider selection by measuring real latency from actual geographic endpoints. Even a single region outage can be caught minutes before users report it, but only if your probes are configured with realistic thresholds matched to your actual production models and baseline performance.

Why synthetic probes matter for OpenRouter

server room
Photo by panumas nikhomkhai from Pexels

OpenRouter abstracts away the complexity of managing multiple LLM providers, Claude, GPT-4, Grok, Llama, and others all route through a single endpoint. But abstraction has a cost: visibility. When your production chat feature slows down, you need to know whether the problem is OpenRouter's routing layer, regional network variance, a specific model's inference time, or your own stack.

Synthetic probes solve this by sending identical test requests from fixed locations around the world, every few minutes, independent of your production traffic. This gives you:

  1. Regional variance visibility: See exactly where latency spikes occur and whether it correlates with geography, time of day, or model selection.
  2. Baseline comparison: Establish what "normal" looks like for each region and model so alerts trigger only on real degradation, not noise.
  3. Incident triage speed: When a user reports slowness, you already have regional probe data to rule out provider-side issues or confirm them in seconds.
  4. Provider decisions backed by data: Compare OpenRouter latency against direct OpenAI or Anthropic endpoints using the same 21 regions, not vendor benchmarks.

Step 1: Gather OpenRouter account details and API keys

Before configuring probes, collect the information you'll need. Log into your OpenRouter account and navigate to the keys section; you'll need your API key and organization ID if you have one. Note the models you depend on in production, typically 2–4 models cover most use cases (e.g., a fast model like Grok 2 for simple queries, Claude 3.5 Sonnet for complex reasoning, and a fallback).

Document these details in a simple table:

FieldValue
OpenRouter API Keysk-or-… (keep secret; use env var in Observinio)
Organization ID(if applicable)
Primary modelanthropic/claude-3.5-sonnet
Secondary modelxai/grok-2
Tertiary model(optional fallback)
Expected baseline TTFB (ms)400–600 (varies by region)
Expected baseline TTFT (ms)1500–3000 (varies by region and model)
OpenRouter's /models endpoint lists all available models and their provider details. If you are unsure which models your traffic uses, check your billing dashboard or request logs over the past week.

Step 2: Define your probe regions and frequency

network cables
Photo by Brett Sayles from Pexels

Observinio probes run from 21 regions worldwide: North America (US East, US West, Canada), Europe (Ireland, Frankfurt, London), Asia-Pacific (Singapore, Tokyo, Sydney, Mumbai), and others. You do not need to enable all 21 regions if your users are concentrated in a few areas, but regional diversity helps catch latency anomalies early.

Recommended region strategy by user base:

  • Global SaaS: Enable all 21 regions; region-specific slowdowns often precede global incidents.
  • North America-focused: Enable US East, US West, Canada, and one EU region (Dublin) to catch transatlantic routing issues.
  • Europe-focused: Enable Dublin, Frankfurt, and London; add Singapore or Sydney if you have APAC users.
  • Startup MVP: Start with 5–7 regions (US East, US West, Dublin, Singapore, Tokyo, Sydney, Canada) and expand as you grow.
Probe frequency recommendations:
  • 5-minute interval: Recommended for production-critical services or if you have past incident history. Generates ~290 data points per model per day per region.
  • 10-minute interval: Good default for most teams; ~144 data points per model per day per region, lower cost.
  • 15-minute interval: Acceptable for non-critical monitoring or cost-sensitive early-stage teams; catches most regional trends but misses spike-and-recover events.
Observinio pricing typically scales with probe count. A 10-minute interval across 10 regions and 3 models = 30 probes per check cycle; monthly cost depends on your plan tier. Estimate your data volume before finalizing frequency.

Step 3: Configure the probe payload and alert thresholds

data center
Photo by Brett Sayles from Pexels

Each probe sends a small completion request to OpenRouter. A good test payload is lightweight (to avoid model-specific variance) but realistic enough to measure true latency. Here's a template:

{
  "model": "anthropic/claude-3.5-sonnet",
  "messages": [
    {
      "role": "user",
      "content": "Respond with a single word."
    }
  ],
  "max_tokens": 10,
  "temperature": 0.7
}

Keep max_tokens low (10–20) so TTFT stays predictable and costs negligible. Rotate between your 2–4 primary models so each probe cycle tests different models fairly.

Alert threshold configuration:

"Documentation IndexFetch the complete documentation index at: /docs/llms.txtUse this file to discover all available pages before exploring further."
>, Presets

After running probes for 3–7 days, you will have enough baseline data to set meaningful alerts. Use Observinio's baseline comparison feature to calculate percentile thresholds:

  • TTFB alert: Set at 1.5–2× the 50th percentile for each region. E.g., if US East baseline TTFB is 400 ms, alert at 600–800 ms.
  • TTFT alert: Set at 1.2–1.5× the 50th percentile for each region and model combination. E.g., if Claude 3.5 Sonnet averages 2000 ms in Ireland, alert at 2400–3000 ms.
  • Availability alert: Trigger if any region fails 3+ consecutive probes (indicates provider or regional network issue).
  • Regional outlier alert: Flag if a single region's TTFB is 50% higher than the global median (suggests regional degradation).
These thresholds avoid alert fatigue while catching real problems. Adjust upward or downward based on your SLO tolerance and user impact tolerance.

Step 4: Deploy and monitor using the configuration worksheet

OpenRouter probe configuration worksheet process
Figure 1: OpenRouter probe configuration worksheet at a glance.

Use this checklist to ensure your probes are correctly deployed:

Your progress is saved automatically in your browser.

Post-deployment monitoring (first 72 hours):

  1. Run a manual test probe from each region to verify connectivity.
  2. Check Observinio dashboard for data arrival; expect 100% success rate on first runs.
  3. Review the raw latency distribution (histogram view) for each region; look for outliers or multi-modal patterns.
  4. Trigger a test alert to verify notification channels work.
  5. Document the baseline TTFB and TTFT for each region and model in your worksheet; use these numbers in future incident triage.
Ongoing maintenance:
  • Weekly review: Check Observinio's weekly summary email for regional trends and model performance variance. Compare this week's median latency to last week's; flag any consistent degradation.
  • Monthly baseline refresh: Every 30 days, recalculate baseline percentiles using the latest 30 days of data. Adjust alert thresholds if OpenRouter performance improves or degrades sustainably.
  • Model rotation: When you adopt a new model in production, add it to your probe rotation and establish baselines before traffic ramps.
  • Incident correlation: After each incident, cross-reference support tickets with Observinio probe data. Document whether the incident was regional, model-specific, or global so future incidents are faster to triage.

Common probe configuration pitfalls to avoid

Over-alerting with tight thresholds: Setting TTFB alerts at the 75th percentile instead of 90th generates 10–15 false alarms per month, causing alert fatigue. Start loose and tighten after 2 weeks of baseline data.

Using a single model in probes: If you only probe Claude 3.5 Sonnet but production uses a mix of Claude, Grok, and Llama, your baseline will not reflect real production latency. Rotate models.

Ignoring regional variance: US East latency is naturally lower than Tokyo latency due to OpenRouter's infrastructure distribution. Set region-specific thresholds, not global ones.

Forgetting to document baselines: When an incident occurs, you need historical context to judge severity. Document baseline TTFB and TTFT for each region and model in your worksheet so you can say "this is 3× normal" vs "this is slightly high".

Probing too infrequently: 30-minute or hourly probes miss spike-and-recover events that last 10–15 minutes. Minimum recommended frequency is 10 minutes for production services.

FAQ

Frequently Asked Questions

TTFB (time-to-first-byte) measures the time from request send to the first token of the response. TTFT (time-to-full-text) measures time from request send to the last token. TTFB reflects OpenRouter's routing and model queue depth; TTFT reflects model inference speed and token count. Monitor both: high TTFB indicates routing delays; high TTFT indicates slow model inference or long outputs.
Probe 2–4 representative models that you actually use in production. Probing 20+ models generates excessive data and costs without insight. Pick a fast model (Grok 2, Llama 405B), a high-quality model (Claude 3.5 Sonnet), and optionally a fallback. Rotate them in your probe payload so each model gets equal attention.
Run probes for at least 3–7 days before setting thresholds. Observinio will show you the distribution (min, p50, p90, p99). For TTFB, start at p90 × 1.2 (20% above the 90th percentile). For TTFT, use p75 × 1.5 (50% above the 75th percentile). Adjust down after a week if you find the threshold is too loose.
Yes. Configure separate probe groups for OpenRouter and direct endpoints (e.g., api.openai.com for GPT-4, api.anthropic.com for Claude) using the same models, regions, and frequency. Compare TTFB and TTFT side-by-side in Observinio's comparison view. This is the best way to decide whether OpenRouter's abstraction overhead is worth it for your use case.
Review your baseline and alert thresholds monthly. After each major incident, document what the probe data showed and whether it helped triage speed. Rotate your probed models every quarter when you adopt new models or deprecate old ones. If your user base shifts geographically (e.g., you expand to APAC), add regions to your probe set.

Next steps: From probes to action

Synthetic probes are only useful if you act on their data. Once your OpenRouter probes are stable, integrate them into your incident response workflow: add Observinio alerts to your on-call Slack channel, include probe latency in weekly engineering standups, and use regional probe data to inform model selection and routing decisions.

When degradation alerts fire, check Observinio's dashboard for regional variance, model-specific slowdowns, and correlation with provider incidents (check OpenRouter's status page at the same timestamp). Document your findings in the incident postmortem so future on-call responders have a playbook.

Observinio's weekly summaries and degradation alerts are designed to surface trends before users complain. Use them to build confidence in your provider choice and catch regional latency regressions early. Start with the baseline configuration in this worksheet, refine thresholds based on your data, and revisit every quarter as your LLM API traffic scales.

Quick Reference: Baseline Thresholds by Region

TTFB targets (p50 baseline): US regions 350–450ms, EU 400–500ms, APAC 500–700ms

Alert trigger (1.5–2× baseline): US 600–800ms, EU 700–900ms, APAC 900–1200ms

TTFT targets (Claude 3.5 Sonnet, p50): 1800–2200ms globally; adjust ±300ms per model

Update these thresholds after your first two weeks of probe data to match your actual infrastructure performance.

Additional Resources