Servizio di Azure che fornisce l'accesso ai modelli GPT-3 di OpenAI con funzionalità aziendali.
Azure OpenAI Responses API (background=true): progressive latency degradation under volume with no 429
Hello everyone,
I'm using the Azure OpenAI Responses API with background=true for long-running completions on gpt-5.1. I submit the job and poll GET /responses/{id} with exponential backoff (5s → 30s), against the v1 endpoint (/openai/v1). My client aborts a job after a 10-minute timeout.
This works fine initially, but after a certain volume of requests latency increases progressively until jobs hit that 10-minute timeout and fail.
What I've observed:
- The degradation seems to build up over time/volume rather than appearing immediately.
- No
429and no error in the payload, requests just take longer and longer until they hit my client-side timeout.
Questions:
- Has anyone seen this pattern in background mode, progressive latency degradation with no explicit error signal?
- What resolved it? Multi-region failover, Provisioned Throughput Units (PTU), or something else?
- Is this capacity/quota-related (implicit throttling on a Standard/PAYG deployment) or something specific to background mode's job scheduling ("soft queueing" server-side)?
Any pointers on what to look at to distinguish gateway/capacity throttling from background-queue latency would be very helpful.
Thanks!