Issue: Tone and Voice Inconsistency on Realtime 1.5 (GA API)

Abdul Rehman 65 Zuverlässigkeitspunkte
2026-04-23T10:52:09.3333333+00:00

We are observing significant voice output inconsistencies on Realtime 1.5 that we did not experience on the previous gpt-realtime model.

Specifically, we see the following behaviors:

  • Gender/tone switching: The voice changes tone mid-turn and, in more severe cases, fluctuates between a female- and male-sounding voice within the same response.
  • Voice identity drift: The model occasionally produces a voice that sounds completely different from the one configured, as if it has defaulted to another internal voice.
  • Variable pacing: Speech speed changes unexpectedly within a single turn, shifting from a normal cadence to a very fast, "stressed-sounding" pace.

Language note: These issues occur primarily in German, which is the primary language used in our voice agents. Similar reports have surfaced in the community for other languages as well, so this does not appear to be language-specific.

The problem is frequent enough to materially impact user experience and, in its current state, blocks us from relying on Realtime 1.5 in production.

Could you advise on:

  1. Any configuration or prompting guidance to prevent or minimize this behavior.
  2. Whether a fix is planned for an upcoming model deployment, and if so, an approximate timeline.

Thank you.

Azure OpenAI in Foundry-Modellen
Azure OpenAI in Foundry-Modellen

Ein Azure-Dienst, der Zugriff auf die GPT-3-Modelle von OpenAI ermöglicht und Unternehmensfunktionen bietet


1 Antwort

Sortieren nach: Am hilfreichsten
  1. Manas R Mohanty 17,270 Zuverlässigkeitspunkte Moderator
    2026-05-05T22:07:32.3766667+00:00

    Hello Abdul Rehman,

    Thank you for your inputs. Here are additional recommendation based on existing cases on OpenAI forums. Strengthen Voice/Tone Instructions in System Prompt Be explicit and detailed about voice identity, tone, pacing, and emotional register in your system message. With Realtime 1.5's improved instruction following (+7%), specific instructions are now adhered to more literally — use this to your advantage. Example additions to your system prompt:

    You MUST maintain a consistent, calm, female voice throughout the entire conversation. Never change your vocal tone, pitch, speed, or gender presentation — even after tool calls. Speak at a steady, moderate pace at all times. Do not speed up or slow down. After returning results from a tool call, resume speaking in the exact same tone and pace as before.
    

    Use Recommended Voices The official guidance is to use marin or cedar voices for best assistant voice quality — these are the realtime-exclusive voices optimized for this model. Handle Tool-Call Tone Drift The tone shift after tool calls is the most commonly reported trigger. Add explicit post-tool-call instructions:

    When you finish a function/tool call and resume speaking, do NOT change your voice, tone, or pacing.
    

    For mission-critical interactions, use topics and tool-driven logic rather than purely free-form generative speech.

    This reduces the model's tendency to improvise vocally.

    Manage Session Length Long or highly stateful conversations can exceed optimal memory retention windows, causing voice drift. Introduce periodic soft resets or topic boundaries. Realtime sessions have a 30-minute maximum — renew before timeout and restore only minimal context. Temperature

    The GA interface has removed temperature as a configurable parameter. The beta interface limits it to 0.6–1.2 (default 0.8). If you're still on the beta interface, lowering temperature may reduce variation but won't eliminate it.

    Fix Timeline

    No specific fix or model revision timeline has been officially announced by either Microsoft or OpenAI for these voice consistency issues.

    But sharing the insights as feedback post minimal trial here to product group Thank you for your inputs here on forum.

    War diese Antwort hilfreich?


Ihre Antwort

Antworten können von Fragestellenden als „Angenommen“ und von Moderierenden als „Empfohlen“ gekennzeichnet werden, wodurch Benutzende wissen, dass diese Antwort das Problem des Fragestellenden gelöst hat.