GPT-6-Astra Responses API reasoning truncated at exactly 516 tokens — Chat Completions unaffected

Chung Quang 0 Reputation points
2026-10-10T12:12:37.7966667+00:00

Model: gpt-6-astra (gpt-6-astra-2026-09-03)

Region: East US 2

Endpoint: Responses API

Issue:

When using the Responses API, single-segment reasoning is consistently capped at exactly 516 tokens. The model then immediately exits reasoning and calls a tool. The response returns status=completed with no truncation warning or incomplete_details.

This does NOT happen on the Chat Completions API with the same model and parameters.

Request parameters:

  • reasoning.effort: xhigh
  • max_output_tokens: 65536
  • No explicit reasoning token limit set

Reproduction (all UTC 2026-09-20, each showing exactly 516 reasoning tokens):

  1. 19:03:24 — apim-request-id: 3751d57b-b192-4ff3-b433-80b58970ea7c
  2. 18:59:29 — apim-request-id: 9613d0b2-8903-4e7c-8a9f-e34c26071bc4
  3. 19:01:20 — apim-request-id: 3d9f6c92-3b51-466c-a14c-dc98a1c01b69
  4. 19:02:31 — apim-request-id: 86edaee2-b3e9-4b27-97f3-bc87368d78bd
  5. 19:04:24 — apim-request-id: 4b346982-dedb-496f-bdf2-31cfd2e9ffe3

Questions:

  1. Is there an undocumented reasoning token limit on the Responses API?
  2. Does 516 reflect actual reasoning_tokens (from usage.output_tokens_details) or only the summary length?
  3. If this is a bug, when can we expect a fix? If intentional, can it be adjusted?
Foundry Models
Foundry Models

A catalog of AI models in Microsoft Foundry that you can discover, compare, and deploy using Azure’s built‑in tools for evaluation, fine‑tuning, and inference

0 comments No comments

1 answer

Sort by: Most helpful
  1. Harshavardhan Bajoria 0 Reputation points
    2026-10-10T18:05:54.8933333+00:00

    Hi @Chung Quang ,
    The issue stems from an internal orchestration and token-buffering boundary in the Azure OpenAI / Microsoft Foundry Responses API during tool-call transitions, rather than an intentional model truncation limit. In the Responses API, intermediate tool-dispatch evaluation intercepts single-segment chain-of-thought at the internal 512 (+4 token protocol overhead) chunk boundary, returning status: "completed" `because the turn successfully produced a tool call rather than hitting an output limit.


    Answers to your questions

    1. Is there an undocumented reasoning token limit on the Responses API?

    There is no intended policy limit, but there is an execution disparity between the endpoints:

    • Chat Completions API: Runs the entire chain-of-thought phase unconstrained before emitting the tool call.
    • Responses API: Uses an agentic runtime where tool selection is actively evaluated during reasoning. Under tool definitions, once the runtime detects tool-call readiness or hits the internal segment buffer (in your case 516 tokens: 512 reasoning tokens + 4 framing tokens), the orchestrator prematurely exits reasoning to execute/emit the tool call. Because the turn produced a valid tool call, no incomplete_details or truncation warnings are raised.

    2. Does 516 reflect actual reasoning_tokens or only the summary length?

    It reflects the actual internal reasoning_tokens generated up to the transition point and billed in usage.output_tokens_details.reasoning_tokens. It does not represent a compressed summary length. The model genuinely generated only 516 reasoning tokens before being forced into tool emission.

    3. Bug resolution and mitigation

    Your may use the following solutions:

    1. Temporary Workaround is routing via Chat Completions: Because the Chat Completions endpoint (/v1/chat/completions) handles unsegmented deep reasoning with reasoning.effort: xhigh properly without truncating before tool dispatch, use it for workflows requiring extensive pre-tool reasoning.
    2. If staying on the Responses API is required then decouple Reasoning and Tool Invocations (Two-Turn Pattern): :
      • Turn 1: Send the prompt with tools: [] or tool_choice: "none" to let reasoning.effort: xhigh reason through the complete output window.
      • Turn 2: Append the previous turn's output/context and provide the tool definitions to emit the call.
    3. Escalate via Azure Support with Request IDs: Because the screenshot contains reproducible APIM request IDs (apim-request-id) in East US 2, open a ticket via the Azure Portal to escalate directly to the Azure OpenAI product group:
      • Go to Azure Portal > Help + support > Create a support request.
      • Service: Azure OpenAI Service / Microsoft Foundry.
      • Problem type: Model Inference & API Execution / Responses API.
      • Paste the 5 provided apim-request-id headers and timestamps to accelerate engineering investigation.

    For more details:

    Please let me know if you still face any issues.

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.