Photo by Jessica Lewis 🦋 thepaintedsquare from Pexels

0regions
Synthetic probe coverage

When you ship an LLM feature to production, latency assumptions crack. A completion that ran in 800 ms during testing now hits 2.5 seconds from your Tokyo data center. Your marketing email goes out, users in Europe start complaining about timeouts, and your incident channel lights up, but you still don't know if the problem is OpenRouter, OpenAI, your routing logic, or regional network degradation.

Synthetic probes and production traffic both carry signal. The question is not which one to use, but how to read them together to make faster, safer decisions about provider selection, regional routing, and degradation response. This worksheet walks you through the measurement discipline that separates reactive fire-fighting from proactive latency management.

Benchmark setup coverage
0%
Use both synthetic probes and production traffic data together: synthetic catches systemic provider degradation early, while production metrics reveal whether latency issues live in your application stack or upstream. Measuring TTFB separately from token generation rate, segmenting results by region, and alerting on synthetic probe lag before production percentiles spike gives you a 15–30 minute window to diagnose and mitigate latency issues before users experience them.

TL;DR

  • Synthetic probes (daily fixed requests from 21 regions) show baseline provider health and catch regional outliers before users notice; production traffic shows real user latency distributions under load, including your stack's overhead.
  • Both matter: synthetic catches systemic degradation early; production latency reveals whether your app's routing, caching, or token handling is the bottleneck.
  • Measure TTFB (time to first byte) separately from total completion time; they answer different questions about where latency lives.
  • Alert on synthetic probe lag, not just on production slowness; if probes degrade but users don't report issues yet, you have a narrow window to investigate.
  • Collect regional variance data; latency often splits by geography, a provider fast in us-east may be slow in eu-west, and raw averages hide that critical detail.

Why both matter: the signal in each

server room
Photo by panumas nikhomkhai from Pexels

Synthetic and production traffic are not rivals, they answer different questions about your LLM latency.

Synthetic probes (fixed, repeatable requests sent on a regular schedule) are your baseline alarm system. Because they are identical across runs, changes in their latency point directly at the provider or network, not your application. If your synthetic TTFB from Singapore to OpenAI increases by 400 ms overnight, you know something shifted upstream. Probes also run whether users are active or not, so they catch provider degradation during off-peak hours before the morning rush hits.

Production traffic is messier but crucial. Real users send variable payloads, make requests during actual load, and experience latency that includes your app's queuing, token counting, cache misses, and timeout logic. Production metrics reveal whether latency is a provider problem or a stack problem. If synthetic latency is stable but production TTFB percentiles are creeping up, your app's request handling, not the provider, needs attention.

"Most teams benefit from using both, with each serving a different role in evaluation."
>, Synthetic vs Real

The practical outcome: synthetic data tells you when to investigate; production data tells you what to fix.


Core metrics: TTFB, token rate, and regional variance

network diagram
Photo by Google DeepMind from Pexels

Before collecting any data, agree on what you are measuring.

Time to first byte (TTFB)

TTFB is the delay from sending a request to receiving the first token of the response. It isolates provider queueing, API processing, and network round-trip time. For user experience, TTFB is often more important than total completion time, a 100 ms TTFB feels responsive even if the full response takes 5 seconds.

Why it matters: A provider can be fast at generating tokens but slow at starting them. OpenAI and OpenRouter may differ sharply in TTFB while generating similar tokens-per-second once the stream begins.

Token generation rate

After the first byte arrives, how fast do tokens flow? Measure tokens per second. A slow token rate (e.g., 8 tokens/sec instead of the expected 15) suggests either provider saturation or a network bottleneck mid-stream.

Regional variance

Latency is not global. Test from at least 5 regions if you serve users worldwide. A provider fast in us-east may lag in ap-southeast. Observinio runs synthetic probes from 21 regions; even if you build a simpler test, segment results by geography.

Percentiles, not just averages

Report p50, p95, and p99 latency. An average of 800 ms hides that 10% of requests hit 3 seconds. Percentile breakdowns show whether slowness is rare outliers or systematic.


Setting up your benchmark worksheet

Benchmark worksheet: synthetic vs production traffic process
Figure 1: Benchmark worksheet: synthetic vs production traffic at a glance.

A benchmark worksheet is a repeatable template, spreadsheet, script, or dashboard, that collects and compares latency across conditions. Here is a step-by-step outline:

1. Define your test cases

List the scenarios you need to understand:

  • Baseline: A small, consistent payload (e.g., "Explain quantum entanglement in 200 words") sent to each provider from each region.
  • Scaling: The same prompt, but with variations in input length (short, medium, long) to see how payload size affects TTFB.
  • Peak load: If possible, send concurrent requests to see whether queueing degrades latency under pressure.
  • Regional edge case: A prompt in a non-English language or from a region with known latency issues.

2. Collect synthetic data weekly

Every Monday (or daily, if you automate it), send your baseline prompts to OpenRouter and OpenAI direct endpoints from representative regions. Record:

  • Provider
  • Region
  • Timestamp
  • TTFB (ms)
  • Token generation rate (tokens/sec)
  • Total completion time (ms)
  • Error rate (if any)
Keep the payload identical across runs so comparisons are valid.

3. Extract production percentiles daily

Pull p50, p95, and p99 TTFB from your application logs or APM tool, segmented by:

  • Provider
  • Region (if your app logs origin)
  • Time of day (peak vs off-peak)

4. Compare weekly

Each week, plot synthetic vs production latency side by side. Look for gaps:

  • If synthetic is stable but production p95 rises, your app's request handling is the issue.
  • If synthetic degrades but production lags behind, you have a buffer before users feel it.
  • If both degrade together, the provider is likely the culprit.

Practical artifact: benchmark template checklist

Your progress is saved automatically in your browser.


Reading the signals: what each pattern means

data analysis
Photo by Thirdman from Pexels

Once you have a few weeks of data, patterns emerge. Here is how to interpret them:

Pattern 1: Synthetic stable, production rising

Signal: Your app or routing is the bottleneck, not the provider.

Action: Check for:
  • Increased request volume (queueing in your orchestration layer).
  • Cache misses (every request hitting the API instead of cached results).
  • Retry loops (failed requests being re-sent, doubling latency).
  • Token counting or validation overhead.

Pattern 2: Synthetic and production both degrade

Signal: The provider or network backbone is degrading.

Action: Check provider status page, run mtr or traceroute to the provider's IP, and consider failover to a secondary provider if your SLA is at risk.

Pattern 3: Synthetic high, production lower

Signal: Rare; usually indicates synthetic test setup is over-conservative (e.g., too-large payload, or tests run during provider maintenance windows).

Action: Revise test payload to match real user requests more closely.

Pattern 4: Regional divergence (one region spiking)

Signal: Geographic or routing issue. That region may be hitting a congested network path or a regional provider endpoint under load.

Action: Failover traffic from that region to another endpoint, or investigate BGP routing with your CDN provider.


Integration with Observinio alerts and status pages

Observinio's synthetic probes from 21 regions and daily email summaries make this workflow automatic. Instead of building your own test harness, you can:

  1. Set degradation alerts on TTFB thresholds (e.g., "alert if OpenRouter TTFB exceeds 1.2 seconds from any region for 5 consecutive probes").
  2. Review weekly summaries to spot trends without manually checking dashboards.
  3. Publish a status page showing real-time and historical latency by region, so stakeholders and customers see the same data you do.
When Observinio detects synthetic degradation, you get email alerts with regional detail. That early signal gives your platform team time to investigate before production traffic percentiles spike, often a 15–30 minute window in which to diagnose and route around the issue.

Latency Monitoring in Action

Teams using structured synthetic and production benchmarks report catching provider degradation 15–30 minutes earlier than reactive monitoring alone, reducing incident duration and customer impact significantly. The investment in setting up baseline probes pays for itself within weeks through faster detection and more informed provider or routing decisions.

FAQ

Frequently Asked Questions

For chat or real-time applications, aim for TTFB under 500 ms from major regions. Under 1 second is acceptable for most use cases. If your TTFB regularly exceeds 1.5 seconds, users will perceive lag, and your churn will rise. Benchmark against your competitors' latency if you have visibility into it.
Daily is the minimum to catch provider degradation before users file support tickets. If you operate a mission-critical service, run probes every hour or every 15 minutes. Observinio runs probes daily by default; that cadence is sufficient for most teams to detect regressions within hours.
Use consistent synthetic payloads for baseline benchmarks, then separately test with production-like prompts (varied length, language, complexity) in a staging environment. Synthetic payloads isolate provider latency; production-like tests reveal end-to-end stack behavior. Do both.
Compare synthetic TTFB (provider only) against production TTFB (provider + your stack). If synthetic is 600 ms and production p95 is 1.2 seconds, your app is adding ~600 ms. Drill into request logs to find where: token counting, queuing, retries, or token streaming delay.
Expect 20–40% variance between the fastest and slowest regions for the same provider, depending on geography and routing. A provider's us-east endpoint might run 400 ms TTFB while ap-southeast runs 700 ms. That is normal. If variance suddenly jumps (e.g., ap-southeast jumps from 700 ms to 2 seconds), investigate network routing or regional outages.
Most providers rate-limit free tiers or stop free trial after a month. Use a small paid account dedicated to probes; the cost is negligible compared to the insight. Observinio handles probe costs and regional distribution out of the box.

Get started with structured latency monitoring

Latency benchmarking is not a one-time exercise, it is a discipline. By collecting synthetic data weekly and comparing it against production percentiles, you move from reactive incident response to proactive degradation detection. The 15–30 minute head start that early alerts provide is often the difference between a quick fix and a customer-visible outage.

If you ship LLM features to multiple regions, start this week: pick a provider, run a baseline synthetic probe from three regions, and set a weekly reminder to compare results. Once the pattern is visible, you will find routing, failover, and provider decisions become data-driven instead of guesswork.

Observinio's 21-region probe network and daily summaries are designed to automate this exact workflow. Set up alerts on TTFB thresholds, publish a status page your team monitors, and let structured latency data inform your next provider or region decision. Try Observinio's degradation alerts to catch latency issues before users do.

Additional Resources