Intermittent "violence: high" content filter block on a benign appointment-booking chat – Azure OpenAI Title: Intermittent "violence: high" content filter block on a benign appointment-booking chat – Azure OpenAI

Nishyanth Nandagopal 0 Reputation points
2026-09-10T18:44:12.9966667+00:00

Hi all,

We run customer-facing chatbots on Azure OpenAI, orchestrated with Semantic Kernel (Python). Some of them intermittently get a prompt-side content filter block in the violence category (severity "high"), even though the conversation is an ordinary appointment booking. The same flow usually works when retried later.

ERROR

HTTP 400, code "content_filter", param "prompt",

innererror "ResponsibleAIPolicyViolation".

content_filter_result: violence filtered (high); hate, sexual, self_harm

all "safe"; jailbreak not detected.

CONVERSATION WHEN IT FAILED

User: hi -> greeting

User: asks to book an appointment -> bot asks for name and email

User: provides name and email -> bot asks for preferred time

User: "tomorrow 10 am" -> request blocked

Thanks!

Content Safety in Foundry Control Plane
Content Safety in Foundry Control Plane

An Azure service that enables users to identify content that is potentially offensive, risky, or otherwise undesirable. Previously known as Azure Content Moderator.

0 comments No comments

2 answers

Sort by: Most helpful
  1. Nishyanth Nandagopal 0 Reputation points
    2026-10-01T07:00:44.26+00:00

    Thank you, @Allan Solomon Mejia , for your response. As you suggested, I tried the same request in the playground available in the Foundry portal, using the same deployed model with identical guardrail configurations. This time, I provided only my query as input, without any additional prompts or tool definitions. However, I still encountered the same content filter block.

    Was this answer helpful?

    0 comments No comments

  2. Allan Solomon Mejia 10,225 Reputation points
    2026-09-12T20:00:37.88+00:00

    Hello @Nishyanth Nandagopal

    Based on the error you've posted, this is an input/prompt-side content-filter decision, not a filter applied to the model's generated response.

    If the complete input really contains only the benign appointment-booking conversation shown above, a violence: high classification would be unexpected. Microsoft's current definition of high-severity violence covers severe harmful content such as attack planning and violent/extremist activity, which doesn't correspond to "tomorrow 10 am" or ordinary appointment scheduling.

    One thing to verify before treating this as a service-side false positive, though: capture the exact request payload Semantic Kernel sends to Azure OpenAI when the failure occurs.

    Semantic Kernel may send much more than the final user message. Depending on your implementation, the request can contain system/developer instructions, previous conversation turns, tool/function definitions, tool results, retrieved context, plugin descriptions, and the current user message.

    So "tomorrow 10 am" may only be the final message in a much larger prompt. Since the error says param: prompt, compare the complete serialized request from a failed invocation with a successful invocation. Redact names, email addresses, tokens, and other sensitive data before sharing it.

    Azure OpenAI's content filtering runs independently on prompts and completions and classifies hate, sexual, violence, and self-harm content by severity. Under the default policy, it blocks content classified as medium or high.

    Also, log the Azure response headers for failed requests, particularly the request ID/correlation information, together with a UTC timestamp, Azure region, model/deployment name, model version, API version, Semantic Kernel version, content-filter configuration, and complete content_filter_result

    If the identical serialized request sometimes succeeds and sometimes returns violence: high, that's especially important evidence. At that point, open an Azure Support case under Azure OpenAI/Foundry and provide several successful and failed request IDs so Microsoft can correlate the classifier decisions internally.

    I don't recommend automatically retrying content_filter errors until they succeed. For a customer-facing application, that can mask the underlying classification problem and isn't a reliable safety strategy. Instead, handle the 400 gracefully and capture telemetry for investigation.

    You can also create a custom content-filter configuration and choose the threshold for each harm category. For example, Azure supports configuring the Violence input filter to block High only, rather than Medium + High. However, in your example, the classifier is already reporting High, so changing from the default Medium threshold to High wouldn't solve this particular failure. Turning filters off or using annotate-only for these harm categories requires approval for modified content filtering.

    Therefore, troubleshoot this in this order:

    Capture exact SK request → compare successful vs failed payloads → capture request IDs/timestamps → reproduce outside Semantic Kernel with the same payload → escalate reproducible false positives to Microsoft.

    Testing the same payload directly against the Azure OpenAI endpoint is particularly useful. If the direct request reproduces the classification, you've largely removed Semantic Kernel from the equation. If it doesn't, inspect what SK adds or changes between calls.

    References:

    Default Azure OpenAI guardrail policies

    Harm categories and severity levels

    Configure content filters

    =============================================================================

    Help make this community better for everyone: If this answer helped or resolved your issue, please accept it or upvote it. If not, share more details in a comment so we can continue the discussion and find the right solution. Thank you.

    Was this answer helpful?


Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.