An Azure service that provides a hybrid, multi-cloud management platform for APIs.
The reported behavior indicates that Azure API Management is buffering the Server-Sent Events (SSE) response from the Azure OpenAI backend instead of forwarding the streamed content to the client as it is generated. For SSE scenarios, the forward-request policy should be configured with buffer-response="false" so that events are relayed immediately to the client. Microsoft documentation also notes that response inspection, validation, caching, and request/response body diagnostics can affect streaming behavior and should be reviewed when troubleshooting SSE-related issues.
Since buffer-response="false" was already configured and the token limit and LLM metric policies were removed, the next step is to verify whether the streaming behavior is observed only through API Management or also when calling the Azure OpenAI endpoint directly. Additionally, review any diagnostics or response-processing configurations that may be inspecting or logging response bodies, as these can impact streaming behavior.
If the issue persists, please reopen the thread and provide the APIM SKU, APIM region, Azure OpenAI region, and an APIM trace for a sample request so that further analysis can be performed.
If the assistance was helpful, kindly take a moment to click on Accept Answer and click on Yes. It will be helpful for other community members.
Thank you.