A unified Azure platform for creating and managing AI models, agents, and applications with built‑in enterprise security, monitoring, and governance
For an Azure OpenAI gpt-5.4-mini deployment, call https://<resource>.openai.azure.com/openai/v1/responses with a token for https://ai.azure.com/.default, and assign Cognitive Services OpenAI User to the calling identity at that Azure OpenAI/Foundry resource scope.
Correct endpoint and token audience
Use:
POST https://<resource>.openai.azure.com/openai/v1/responses
Authorization: Bearer <token>
Content-Type: application/json
Acquire the token with this OAuth scope:
https://ai.azure.com/.default
The current Responses API documentation uses the Azure OpenAI v1 endpoint—not the project endpoint under *.services.ai.azure.com/api/projects/...—for direct model inference. The request’s model value must be the deployment name, which may differ from the underlying model name.^1^
A minimal Python test is:
from azure.identity import DefaultAzureCredential, get_bearer_token_provider
from openai import OpenAI
token_provider = get_bearer_token_provider(
DefaultAzureCredential(), "https://ai.azure.com/.default"
)
client = OpenAI(
base_url="https://<resource>.openai.azure.com/openai/v1/",
api_key=token_provider,
)
result = client.responses.create(
model="<deployment-name>", input="Return exactly: OK"
)
print(result.output_text)
This endpoint-and-token pattern is the documented keyless Responses API configuration.^2^
Least-privilege RBAC for this route
For direct Azure OpenAI inference, Cognitive Services OpenAI User is the appropriate built-in inference role. Assign it to the actual calling user, service principal, or managed identity at the resource represented by <resource>.openai.azure.com. That role permits Entra-authenticated inference calls.^3^
The current built-in definition includes:
Microsoft.CognitiveServices/accounts/OpenAI/responses/*
Therefore, a separate Microsoft.CognitiveServices/accounts/OpenAI/responses/write permission is not the documented operation name, and a custom role is unnecessary when Cognitive Services OpenAI User is used.^4^
After temporarily restoring the role, verify that it is attached to the same principal whose oid is in the token:
az role assignment list \
--assignee <principal-object-id> \
--scope <foundry-resource-id> \
--include-inherited \
--output table
Since the previous test already allowed more than five minutes, propagation alone is not a useful next diagnosis.
Do not mix the two inference surfaces
Microsoft’s Foundry Models keyless-authentication guidance also documents a different authorization pattern: Cognitive Services User at the Foundry resource, with Microsoft.CognitiveServices/accounts/MaaS/* for a custom role.^5^
That applies to the Foundry Models inference surface. It should not be substituted for the Azure OpenAI v1 Responses route above. In particular:
- Azure OpenAI route:
*.openai.azure.com/openai/v1/responses - Azure OpenAI role family:
accounts/OpenAI/... - Foundry Models inference role family:
accounts/MaaS/... - Foundry project endpoint:
*.services.ai.azure.com/api/projects/<project-name>, used for project SDK/agent operations rather than the direct Azure OpenAI model route.^6^
A route/role-family mismatch is therefore the next item to eliminate.
Additional prerequisites
For the direct Responses API call:
- A deployed Azure OpenAI model is required.
- Microsoft Entra ID is supported and recommended.
- The resource must have a custom subdomain for token-based authentication.^7^
- No additional Foundry project or project managed identity is documented as necessary for direct model inference.
- Use the deployment name in
model; a wrong deployment name normally leads to a not-found response rather than establishing an RBAC failure.^1^ - If public network access is disabled or Private Link is configured, the caller must use whatever network path the resource configuration permits; the supplied evidence does not establish whether this resource has such restrictions.
gpt-5.4-mini U.S. Data Zone Standard support
gpt-5.4-mini, version 2026-03-17, is listed as supported by the Responses API.^1^ It is also listed for Data Zone Standard deployment in the documented U.S. regions: centralus, eastus, eastus2, northcentralus, southcentralus, westus, and westus3.^8^
Thus, U.S. Data Zone Standard does not itself preclude this Entra authentication path. If the canonical endpoint, audience, deployment name, principal, and resource-scoped role all match and the request still returns 401 PermissionDenied, capture the response apim-request-id, timestamp, resource region, endpoint host, token aud/oid claims, and role-assignment output for an Azure support case.
References
- Use the Azure OpenAI Responses API
- Creates an OpenAI client in Python using Microsoft Entra ID authentication with DefaultAzureCredential to call the Responses API and get a response from a deplo...
- Role-based access control for Azure OpenAI in Azure AI Foundry Models
- Azure built-in roles for AI + machine learning
- Configure keyless authentication with Microsoft Entra ID (programming-language-cli)
- Microsoft Foundry SDKs and endpoints (programming-language-csharp)
- Authentication and authorization in Microsoft Foundry
- Region availability for Foundry Models sold by Azure (standard)