Azure OpenAI – Higher Latency for gpt-5-mini and gpt-5-nano compared to gpt-5

tobias.steinmetz 6 Zuverlässigkeitspunkte
2026-04-01T06:22:47.4766667+00:00

Hello,

we are using Azure OpenAI in the Sweden Central region and have noticed a consistent latency difference between our deployed models.

Our observation:

gpt-5-chat is significantly faster than gpt-5-mini, gpt-5-nano, gpt-5.1-chat and gpt-5.2-chat

This is unexpected, as smaller models are generally assumed to have lower latency

All models are deployed in the same region (Sweden Central)

All calls go through the same middleware (no architectural differences)

We have already increased the rate limits (TPM) to the maximum available quota – the latency difference persists

Our deployment configuration:

Resource: Azure OpenAI (Sweden Central)

Deployment type: Standard (Pay-as-you-go)

gpt-5-chat: 1.000.000 TPM

gpt-5-mini: 1.002.000 TPM

gpt-5-nano: 9.922.000 TPM

Our questions:

Why is gpt-5-chat faster than gpt-5-mini and gpt-5-nano despite being a larger model?

Is this related to infrastructure prioritization or hardware allocation for different model sizes?

Is there any configuration or deployment option that could improve the latency of gpt-5-mini and gpt-5-nano?

Thank you in advance for your help!

Azure OpenAI in Foundry-Modellen
Azure OpenAI in Foundry-Modellen

Ein Azure-Dienst, der Zugriff auf die GPT-3-Modelle von OpenAI ermöglicht und Unternehmensfunktionen bietet


1 Antwort

Sortieren nach: Neueste
  1. Karnam Venkata Rajeswari 5,340 Zuverlässigkeitspunkte Externe Microsoft-Mitarbeiter Moderator
    2026-04-01T12:13:55.8033333+00:00

    Hello tobias.steinmetz,

    Welcome to Microsoft Q&A .Thank you for reaching out.

    This behavior can occur and has been observed across regions and models. Smaller models are primarily optimized for cost efficiency and high‑throughput scenarios, not necessarily the lowest per‑request latency. As a result, it is possible for a larger, chat‑optimized model to respond faster in certain workloads. This is considered expected behavior for Standard deployments and does not indicate a service issue.

    Model size alone does not determine response latency. In Azure OpenAI, latency is influenced by multiple factors such as internal model architecture, inference pipelines, backend scheduling, and regional load. Some smaller models, including mini and nano variants, may perform additional internal processing steps before generating the first token, which can increase time‑to‑first‑token even though the overall model is smaller.

    Chat‑optimized models like gpt‑5‑chat are tuned for conversational workloads and streaming scenarios, which can result in faster perceived responses for typical chat requests.

    Standard Azure OpenAI deployments operate on shared, multi‑tenant infrastructure. There is no exposed configuration to control or select specific hardware, nor are latency guarantees provided under this deployment type. Capacity is dynamically managed per region and per model, and observed latency differences typically result from backend scheduling behavior, workload characteristics, and regional demand rather than explicit prioritization of one model over another.

    Increasing Tokens Per Minute (TPM) helps prevent throttling and request rejection, but it does not directly guarantee lower per‑request latency. This explains why latency differences can persist even after quota increases.

    Please consider the following configuration and deployment options to improve latency and see if these help:

    1. Optimizing request shape
      • Reducing prompt size and output token limits where possible.
      • Consider avoiding unnecessarily large system or context prompts.
    2. Enabling streaming responses
      • Streaming lowers perceived latency by returning tokens as they are generated, improving time‑to‑first‑token.
    3. Separating workloads
      • Deploy different workloads (for example, free‑form chat versus structured or function‑style requests) to separate deployments to improve batching and consistency.
    4. Monitoring latency distribution
      • Track P50, P95, and P99 latency metrics using Azure Monitor to identify peak‑load periods and patterns.
    5. Consider Provisioned Throughput Units (PTU)
      • PTU deployments provide dedicated capacity and more predictable latency, and are the recommended option for latency‑sensitive or production‑critical workloads.

    References:

    Thank you 

    Please 'Upvote'(Thumbs-up) and 'Accept' as answer if the response was helpful. This will be benefitting other community members who face the same issue.

    War diese Antwort hilfreich?

    0 Kommentare Keine Kommentare

Ihre Antwort

Antworten können von Fragestellenden als „Angenommen“ und von Moderierenden als „Empfohlen“ gekennzeichnet werden, wodurch Benutzende wissen, dass diese Antwort das Problem des Fragestellenden gelöst hat.