A catalog of AI models in Microsoft Foundry that you can discover, compare, and deploy using Azure’s built‑in tools for evaluation, fine‑tuning, and inference
Hi @Chung Quang ,
The issue stems from an internal orchestration and token-buffering boundary in the Azure OpenAI / Microsoft Foundry Responses API during tool-call transitions, rather than an intentional model truncation limit. In the Responses API, intermediate tool-dispatch evaluation intercepts single-segment chain-of-thought at the internal 512 (+4 token protocol overhead) chunk boundary, returning status: "completed" `because the turn successfully produced a tool call rather than hitting an output limit.
Answers to your questions
1. Is there an undocumented reasoning token limit on the Responses API?
There is no intended policy limit, but there is an execution disparity between the endpoints:
- Chat Completions API: Runs the entire chain-of-thought phase unconstrained before emitting the tool call.
- Responses API: Uses an agentic runtime where tool selection is actively evaluated during reasoning. Under tool definitions, once the runtime detects tool-call readiness or hits the internal segment buffer (in your case 516 tokens: 512 reasoning tokens + 4 framing tokens), the orchestrator prematurely exits reasoning to execute/emit the tool call. Because the turn produced a valid tool call, no incomplete_details or truncation warnings are raised.
2. Does 516 reflect actual reasoning_tokens or only the summary length?
It reflects the actual internal reasoning_tokens generated up to the transition point and billed in usage.output_tokens_details.reasoning_tokens. It does not represent a compressed summary length. The model genuinely generated only 516 reasoning tokens before being forced into tool emission.
3. Bug resolution and mitigation
Your may use the following solutions:
- Temporary Workaround is routing via Chat Completions: Because the Chat Completions endpoint (/v1/chat/completions) handles unsegmented deep reasoning with reasoning.effort: xhigh properly without truncating before tool dispatch, use it for workflows requiring extensive pre-tool reasoning.
- If staying on the Responses API is required then decouple Reasoning and Tool Invocations (Two-Turn Pattern): :
- Turn 1: Send the prompt with tools: [] or tool_choice: "none" to let reasoning.effort: xhigh reason through the complete output window.
- Turn 2: Append the previous turn's output/context and provide the tool definitions to emit the call.
- Escalate via Azure Support with Request IDs: Because the screenshot contains reproducible APIM request IDs (apim-request-id) in East US 2, open a ticket via the Azure Portal to escalate directly to the Azure OpenAI product group:
- Go to Azure Portal > Help + support > Create a support request.
- Service: Azure OpenAI Service / Microsoft Foundry.
- Problem type: Model Inference & API Execution / Responses API.
- Paste the 5 provided apim-request-id headers and timestamps to accelerate engineering investigation.
For more details:
- Microsoft Learn - Azure OpenAI Service Reasoning Models Documentation
- Use the Azure OpenAI Responses API
Please let me know if you still face any issues.