Signals worth comparing
Success and error rate
Compare status classes and categorized errors over a defined window.
TTFT percentiles
Use p50, p90, and p95 first-token latency instead of a single request.
Total duration
Separate first-event delay from the time required to finish the stream.
A comparable sample
- Use the same model, endpoint, streaming mode, and similar input size.
- Collect at least 10 requests in the same time window.
- Record timezone, status, TTFT, total duration, and request ID.
- Do not infer gateway-wide health from one slow upstream response.
Routing and upstream subscription labels
Some core routes can prioritize upstream accounts labeled Plus or Pro. These labels describe upstream subscription tiers only. They do not make Quick AI Coding an OpenAI service and do not guarantee latency, capacity, availability, model access, or API feature parity.
Historical metrics and SLA boundary
Historical monitoring supports diagnosis and trend analysis. Unless a separate written agreement states otherwise, public measurements and account labels do not create an uptime, latency, capacity, or model-availability SLA.