A catalog of AI models in Microsoft Foundry that you can discover, compare, and deploy using Azure’s built‑in tools for evaluation, fine‑tuning, and inference
GPT-6-Astra Responses API reasoning truncated at exactly 516 tokens — Chat Completions unaffected
Chung Quang
0
Reputation points
Model: gpt-6-astra (gpt-6-astra-2026-09-03)
Region: East US 2
Endpoint: Responses API
Issue:
When using the Responses API, single-segment reasoning is consistently capped at exactly 516 tokens. The model then immediately exits reasoning and calls a tool. The response returns status=completed with no truncation warning or incomplete_details.
This does NOT happen on the Chat Completions API with the same model and parameters.
Request parameters:
- reasoning.effort: xhigh
- max_output_tokens: 65536
- No explicit reasoning token limit set
Reproduction (all UTC 2026-09-20, each showing exactly 516 reasoning tokens):
- 19:03:24 — apim-request-id: 3751d57b-b192-4ff3-b433-80b58970ea7c
- 18:59:29 — apim-request-id: 9613d0b2-8903-4e7c-8a9f-e34c26071bc4
- 19:01:20 — apim-request-id: 3d9f6c92-3b51-466c-a14c-dc98a1c01b69
- 19:02:31 — apim-request-id: 86edaee2-b3e9-4b27-97f3-bc87368d78bd
- 19:04:24 — apim-request-id: 4b346982-dedb-496f-bdf2-31cfd2e9ffe3
Questions:
- Is there an undocumented reasoning token limit on the Responses API?
- Does 516 reflect actual reasoning_tokens (from usage.output_tokens_details) or only the summary length?
- If this is a bug, when can we expect a fix? If intentional, can it be adjusted?
Foundry Models
Foundry Models
Sign in to answer