Azure OpenAI Responses API (background=true): progressive latency degradation under volume with no 429

Marta Sandri 0 Punti di reputazione
2026-07-03T15:06:17.32+00:00

Hello everyone,

I'm using the Azure OpenAI Responses API with background=true for long-running completions on gpt-5.1. I submit the job and poll GET /responses/{id} with exponential backoff (5s → 30s), against the v1 endpoint (/openai/v1). My client aborts a job after a 10-minute timeout.

This works fine initially, but after a certain volume of requests latency increases progressively until jobs hit that 10-minute timeout and fail.

What I've observed:

  • The degradation seems to build up over time/volume rather than appearing immediately.
  • No 429 and no error in the payload, requests just take longer and longer until they hit my client-side timeout.

Questions:

  1. Has anyone seen this pattern in background mode, progressive latency degradation with no explicit error signal?
  2. What resolved it? Multi-region failover, Provisioned Throughput Units (PTU), or something else?
  3. Is this capacity/quota-related (implicit throttling on a Standard/PAYG deployment) or something specific to background mode's job scheduling ("soft queueing" server-side)?

Any pointers on what to look at to distinguish gateway/capacity throttling from background-queue latency would be very helpful.

Thanks!

Azure OpenAI in Modelli di Fonderia
0 commenti Nessun commento

Risposta

Le risposte possono essere contrassegnate come "Accettata" dall'autore della domanda e "Consigliata" dai moderatori, in modo da consentire agli utenti di sapere che la risposta ha risolto il problema dell'autore.