Managed Compute Qwen3-Coder: /openai/v1/responses returns zero token usage while /chat/completions reports correct usage

Patrick AGIN 0 Points de réputation
2026-09-02T21:37:55.7533333+00:00

Hello,

We are observing incorrect token usage reporting from the OpenAI-compatible

Responses API for a Qwen3-Coder model deployed using Azure AI Foundry

Managed Compute.

The inference itself succeeds and returns the expected generated text, but

all usage counters in the Responses API result are zero.

The same deployment reports correct token usage through Chat Completions.

Environment


Managed Compute deployment:

qwen3-coder-30b-a3b-fp8

Model asset:

azureml://registries/azure-huggingface/models/qwen--qwen3-coder-30b-a3b-instruct-fp8/versions/3

Deployment template:

azureml://registries/azure-huggingface/deploymenttemplates/qwen--qwen3-coder-30b-a3b-instruct-fp8--256k-nvidia-h100/labels/latest

Accelerator:

1 x H100 80 GB

Runtime fingerprint returned by Chat Completions:

vllm-0.22.0-ep-e7db14b8

Reproduction 1: Chat Completions


curl -sS -X POST \

"https://<REDACTED>.cognitiveservices.azure.com/openai/deployments/qwen3-coder-30b-a3b-fp8/chat/completions?api-version=2025-04-01-preview" \

-H "Content-Type: application/json" \

-H "api-key: <REDACTED>" \

-d '{

"messages": [

  {

    "role": "user",

    "content": "Reply only: OK"

  }

],

"max_tokens": 32,

"temperature": 0,

"stream": false
```  }'

The response is successful and reports correct usage:

{

  "model": "qwen--qwen3-coder-30b-a3b-instruct-fp8",

  "choices": [

```scala
{

  "message": {

    "role": "assistant",

    "content": "OK"

  }

}
```  ],

  "system_fingerprint": "vllm-0.22.0-ep-e7db14b8",

  "usage": {

```yaml
"prompt_tokens": 12,

"completion_tokens": 2,

"total_tokens": 14
```  }

}

Reproduction 2: Responses API

-----------------------------

curl -sS -X POST \

  "https://<REDACTED>.cognitiveservices.azure.com/openai/v1/responses" \

  -H "Content-Type: application/json" \

  -H "api-key: <REDACTED>" \

  -d '{

```python
"model": "qwen3-coder-30b-a3b-fp8",

"input": "Reply only: OK",

"max_output_tokens": 32,

"temperature": 0,

"stream": false
```  }'

The response is also successful and contains the expected text "OK".

However, usage is:

{

  "input_tokens": 0,

  "output_tokens": 0,

  "total_tokens": 0,

  "input_tokens_details": {

"cached_tokens": 0


  "output_tokens_details": {

```yaml
"reasoning_tokens": 0
```  }

}

Streaming behavior

------------------

The same discrepancy occurs with streaming:

- Chat Completions with stream=true and

  stream_options.include_usage=true reports 12 input tokens,

  2 output tokens and 14 total tokens.

- The final response.completed event from /openai/v1/responses reports

  input_tokens=0, output_tokens=0 and total_tokens=0.

Expected behavior

-----------------

The Responses API should report the actual input and output token counts,

consistent with the Chat Completions API for the same deployment, prompt

and generated output.

The zero values originate from the Foundry

Managed Compute serving path or its packaged vLLM runtime.

We found vLLM PR #22667, which fixed zero Responses API usage for the

GPT-OSS Harmony context. However, this deployment uses Qwen3-Coder and

appears to follow the non-Harmony Responses path, so that fix may not

apply here.

Questions

---------

1. Can you reproduce the zero usage values for this Managed Compute

   deployment and model?

2. Is /openai/v1/responses processed by a Foundry compatibility layer,

   or passed directly to the Microsoft-curated vLLM runtime?

3. Is there a newer deployment template or runtime image that fixes

   Responses API token usage for Qwen models?

1. If the issue is in upstream vLLM, can you confirm this and provide

   the relevant upstream issue or route it to the appropriate team?
   
Thank you.
Azure OpenAI dans les modèles Foundry
0 commentaires Aucun commentaire

1 réponse

Trier par : Les plus récents
  1. Karnam Venkata Rajeswari 5,340 Points de réputation Personnel externe Microsoft Modérateur
    2026-09-03T00:03:07.0333333+00:00

    Hello Patrick AGIN

    Welcome to Microsoft Q&A .Thank you for reaching out to us.

    Based on the behavior observed, the issue appears to be related to how token usage information is surfaced through the Responses API rather than a problem with model execution itself. Since the deployment successfully generates the expected output and the equivalent Chat Completions requests report non-zero token usage

    Managed Compute uses platform-managed runtimes and exposes OpenAI-compatible endpoints for supported models.

    Deployment templates may also include runtime-specific configurations and serving optimizations. However, the internal implementation details of request processing, usage accounting, and response construction are not publicly documented.

    Because of this, it is not currently possible to determine whether the discrepancy originates within the underlying runtime implementation or another component of the Managed Compute serving stack without further engineering investigation.

    Since the deployment is already using the latest deployment template label, there is currently no known upgrade path expected to resolve this behavior.

    To help narrow down the source of the discrepancy, the following validation steps are recommended:

    1. Comparing Responses API usage values with Azure Monitor metrics:
      • Input Tokens
      • Output Tokens
      • Total Tokens
      If Azure Monitor reports non-zero token values while the Responses API payload continues to return zeros, this would indicate that token accounting is occurring internally and that the discrepancy is limited to the values exposed through the API response.
    2. Performing a comparison using another Managed Compute model and verify usage reporting between:
      • /openai/v1/responses
      • /chat/completions
      This can help determine whether the behavior is specific to Qwen3-Coder or affects a broader set of models deployed on Managed Compute.

    As a temporary workaround, Chat Completions can be used when accurate per-request token usage reporting is required, since usage reporting is currently functioning correctly through that API path. For workloads that require the Responses API, Azure Monitor metrics can be used as an alternative source to validate token consumption

    The following references might be helpful , please check them out

    Please let us know if the response was helpful

     

    Thank you

    Cette réponse a-t-elle été utile ?


Votre réponse

Les réponses peuvent être marquées comme « Acceptées » par l’auteur de la question et « Recommandées » par les modérateurs, ce qui aide les utilisateurs à savoir que la réponse a résolu le problème de l’auteur.