Measurement, not promises

Service Reliability and Performance Metrics

Use comparable request samples to understand success rate, errors, first-token latency, total duration, and the boundaries between gateway and upstream time.

Signals worth comparing

Success and error rate

Compare status classes and categorized errors over a defined window.

TTFT percentiles

Use p50, p90, and p95 first-token latency instead of a single request.

Total duration

Separate first-event delay from the time required to finish the stream.

A comparable sample

  • Use the same model, endpoint, streaming mode, and similar input size.
  • Collect at least 10 requests in the same time window.
  • Record timezone, status, TTFT, total duration, and request ID.
  • Do not infer gateway-wide health from one slow upstream response.

Routing and upstream subscription labels

Some core routes can prioritize upstream accounts labeled Plus or Pro. These labels describe upstream subscription tiers only. They do not make Quick AI Coding an OpenAI service and do not guarantee latency, capacity, availability, model access, or API feature parity.

Historical metrics and SLA boundary

Past measurements are not a guarantee

Historical monitoring supports diagnosis and trend analysis. Unless a separate written agreement states otherwise, public measurements and account labels do not create an uptime, latency, capacity, or model-availability SLA.