Azure AI Foundry Responses API returns 401 with Entra ID despite resource-scoped inference role

Maurice Fluelling 0 Reputation points
2026-10-10T02:20:30.8233333+00:00

I am testing a non-production Azure AI Foundry deployment using U.S. Data Zone Standard. Authentication uses Microsoft Entra ID; no API keys or desktop-embedded credentials are permitted.

Token acquisition succeeds. The evaluation identity was temporarily assigned Cognitive Services OpenAI User at the exact Foundry resource scope. After waiting more than five minutes for propagation, a single synthetic request to the Responses API returned HTTP 401 PermissionDenied before inference. No model tokens were consumed.

Could Microsoft clarify:

  1. The currently supported Responses API endpoint and Entra token audience.
  2. The least-privilege RBAC role and required assignment scope.
  3. Whether the role must include Microsoft.CognitiveServices/accounts/OpenAI/responses/write or a newer equivalent.
  4. Whether any Foundry project, provider, deployment, or network configuration is additionally required.
  5. Whether a current gpt-5.4-mini U.S. Data Zone Standard deployment supports this authentication path.

The temporary role was removed after testing. No customer or production data was transmitted.

Microsoft Foundry
Microsoft Foundry

A unified Azure platform for creating and managing AI models, agents, and applications with built‑in enterprise security, monitoring, and governance


1 answer

Sort by: Most helpful
  1. AI answer

    2026-10-10T02:21:29.1733333+00:00

    For an Azure OpenAI gpt-5.4-mini deployment, call https://<resource>.openai.azure.com/openai/v1/responses with a token for https://ai.azure.com/.default, and assign Cognitive Services OpenAI User to the calling identity at that Azure OpenAI/Foundry resource scope.

    Correct endpoint and token audience

    Use:

    POST https://<resource>.openai.azure.com/openai/v1/responses
    Authorization: Bearer <token>
    Content-Type: application/json
    

    Acquire the token with this OAuth scope:

    https://ai.azure.com/.default
    

    The current Responses API documentation uses the Azure OpenAI v1 endpoint—not the project endpoint under *.services.ai.azure.com/api/projects/...—for direct model inference. The request’s model value must be the deployment name, which may differ from the underlying model name.^1^

    A minimal Python test is:

    from azure.identity import DefaultAzureCredential, get_bearer_token_provider
    from openai import OpenAI
    
    token_provider = get_bearer_token_provider(
        DefaultAzureCredential(), "https://ai.azure.com/.default"
    )
    client = OpenAI(
        base_url="https://<resource>.openai.azure.com/openai/v1/",
        api_key=token_provider,
    )
    result = client.responses.create(
        model="<deployment-name>", input="Return exactly: OK"
    )
    print(result.output_text)
    

    This endpoint-and-token pattern is the documented keyless Responses API configuration.^2^

    Least-privilege RBAC for this route

    For direct Azure OpenAI inference, Cognitive Services OpenAI User is the appropriate built-in inference role. Assign it to the actual calling user, service principal, or managed identity at the resource represented by <resource>.openai.azure.com. That role permits Entra-authenticated inference calls.^3^

    The current built-in definition includes:

    Microsoft.CognitiveServices/accounts/OpenAI/responses/*
    

    Therefore, a separate Microsoft.CognitiveServices/accounts/OpenAI/responses/write permission is not the documented operation name, and a custom role is unnecessary when Cognitive Services OpenAI User is used.^4^

    After temporarily restoring the role, verify that it is attached to the same principal whose oid is in the token:

    az role assignment list \
      --assignee <principal-object-id> \
      --scope <foundry-resource-id> \
      --include-inherited \
      --output table
    

    Since the previous test already allowed more than five minutes, propagation alone is not a useful next diagnosis.

    Do not mix the two inference surfaces

    Microsoft’s Foundry Models keyless-authentication guidance also documents a different authorization pattern: Cognitive Services User at the Foundry resource, with Microsoft.CognitiveServices/accounts/MaaS/* for a custom role.^5^

    That applies to the Foundry Models inference surface. It should not be substituted for the Azure OpenAI v1 Responses route above. In particular:

    • Azure OpenAI route: *.openai.azure.com/openai/v1/responses
    • Azure OpenAI role family: accounts/OpenAI/...
    • Foundry Models inference role family: accounts/MaaS/...
    • Foundry project endpoint: *.services.ai.azure.com/api/projects/<project-name>, used for project SDK/agent operations rather than the direct Azure OpenAI model route.^6^

    A route/role-family mismatch is therefore the next item to eliminate.

    Additional prerequisites

    For the direct Responses API call:

    • A deployed Azure OpenAI model is required.
    • Microsoft Entra ID is supported and recommended.
    • The resource must have a custom subdomain for token-based authentication.^7^
    • No additional Foundry project or project managed identity is documented as necessary for direct model inference.
    • Use the deployment name in model; a wrong deployment name normally leads to a not-found response rather than establishing an RBAC failure.^1^
    • If public network access is disabled or Private Link is configured, the caller must use whatever network path the resource configuration permits; the supplied evidence does not establish whether this resource has such restrictions.

    gpt-5.4-mini U.S. Data Zone Standard support

    gpt-5.4-mini, version 2026-03-17, is listed as supported by the Responses API.^1^ It is also listed for Data Zone Standard deployment in the documented U.S. regions: centralus, eastus, eastus2, northcentralus, southcentralus, westus, and westus3.^8^

    Thus, U.S. Data Zone Standard does not itself preclude this Entra authentication path. If the canonical endpoint, audience, deployment name, principal, and resource-scoped role all match and the request still returns 401 PermissionDenied, capture the response apim-request-id, timestamp, resource region, endpoint host, token aud/oid claims, and role-assignment output for an Azure support case.


    References

    1. Use the Azure OpenAI Responses API
    2. Creates an OpenAI client in Python using Microsoft Entra ID authentication with DefaultAzureCredential to call the Responses API and get a response from a deplo...
    3. Role-based access control for Azure OpenAI in Azure AI Foundry Models
    4. Azure built-in roles for AI + machine learning
    5. Configure keyless authentication with Microsoft Entra ID (programming-language-cli)
    6. Microsoft Foundry SDKs and endpoints (programming-language-csharp)
    7. Authentication and authorization in Microsoft Foundry
    8. Region availability for Foundry Models sold by Azure (standard)
    AI-generated content may be incorrect. Read our transparency notes for more information.

    Was this answer helpful?

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.