An Azure service that integrates speech processing into apps and services.
@Siarhei Lukashenka - Based on the current Azure Speech Service documentation, Microsoft does not provide a customer-configurable setting to increase the number of pre-warmed/standby instances or reserve a larger baseline capacity for a hosted Custom TTS endpoint. Azure Speech manages scaling automatically, and customers cannot directly configure the minimum number of backend instances kept ready to serve requests.
It is also important to distinguish quota from service capacity. Quota defines the throughput limit available to the resource, whereas autoscaling and backend capacity determine how quickly the service can accommodate a sudden increase in demand. The Speech Service documentation explains that when traffic increases rapidly, throttling can occur while the service is scaling to meet the additional demand; in that situation, simply increasing the quota does not resolve the underlying scale-out delay.
For sudden traffic spikes, Microsoft recommends gradually ramping the workload where possible and implementing retry logic with appropriate backoff rather than sending a large instantaneous increase in traffic. These measures can reduce the likelihood of transient failures while the service is scaling.
For your specific question, however, there is currently no documented mechanism to request or configure a larger always-warm pool for a hosted Speech endpoint.
If you continue to see ResourceExhausted: No tokens available for passthrough (WebSocket Error 1013) while remaining within your expected quota and traffic limits, I recommend opening an Azure Support request. The Speech engineering team can then review the backend telemetry for your particular endpoint, region, voice, and traffic pattern to determine whether backend capacity or scale-out behavior is contributing to the failures and whether any service-side mitigation is available.
If possible, please provide the region, voice/custom endpoint, approximate request/concurrency level during the burst, and UTC timestamps of several failed requests. Also, if the same requests succeed after a short retry delay, that would be useful information when investigating whether the failures are transient.
References:
- Azure Speech quotas and limits: Azure Speech quotas and limits
- Azure Speech Service documentation: Azure Speech Service documentation
Hope this help!