GPT-6-Astra Responses API reasoning truncated at exactly 516 tokens — Chat Completions unaffected

Chung Quang 0 Reputation points
2026-10-10T12:12:37.7966667+00:00

Model: gpt-6-astra (gpt-6-astra-2026-09-03)

Region: East US 2

Endpoint: Responses API

Issue:

When using the Responses API, single-segment reasoning is consistently capped at exactly 516 tokens. The model then immediately exits reasoning and calls a tool. The response returns status=completed with no truncation warning or incomplete_details.

This does NOT happen on the Chat Completions API with the same model and parameters.

Request parameters:

  • reasoning.effort: xhigh
  • max_output_tokens: 65536
  • No explicit reasoning token limit set

Reproduction (all UTC 2026-09-20, each showing exactly 516 reasoning tokens):

  1. 19:03:24 — apim-request-id: 3751d57b-b192-4ff3-b433-80b58970ea7c
  2. 18:59:29 — apim-request-id: 9613d0b2-8903-4e7c-8a9f-e34c26071bc4
  3. 19:01:20 — apim-request-id: 3d9f6c92-3b51-466c-a14c-dc98a1c01b69
  4. 19:02:31 — apim-request-id: 86edaee2-b3e9-4b27-97f3-bc87368d78bd
  5. 19:04:24 — apim-request-id: 4b346982-dedb-496f-bdf2-31cfd2e9ffe3

Questions:

  1. Is there an undocumented reasoning token limit on the Responses API?
  2. Does 516 reflect actual reasoning_tokens (from usage.output_tokens_details) or only the summary length?
  3. If this is a bug, when can we expect a fix? If intentional, can it be adjusted?
Foundry Models
Foundry Models

A catalog of AI models in Microsoft Foundry that you can discover, compare, and deploy using Azure’s built‑in tools for evaluation, fine‑tuning, and inference

0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.