Controlling F0 and intensity in Azure Speech Studio TTS

Laura Marchesini 20 Reputation points
2026-10-01T09:04:40.9133333+00:00

Hi everyone! I am using Azure Speech Studio – Text to Speech to create some British English syllables containing specific vowels. To have better control over the speech segment, I generated a sentence in which a word containing the vowel of interest is inserted (e.g., “Now I read kid,” where kid is the target). The sentence remains the same, while the target word changes depending on the vowel. I used IPA (International Phonetic Alphabet) in SSML to specify the target pronunciation. Since I need to extract speech segments that are as constant as possible in terms of their acoustic parameters, apart from the vowel itself, I checked the mean intensity and fundamental frequency of the generated syllable. However, I noticed that these values change each time the sentence is generated, even when I use controls such as <prosody pitch="medium"> and <prosody volume="medium">. I also tried specifying a numeric absolute value for these parameters, but for some reason I cannot get it to work as expected.

Does anyone know whether there is a way to precisely control these parameters so that they remain constant across all productions? Thank you in advance for any help!

Azure Speech in Foundry Tools

Answer accepted by question author
Sai Kiran Mudavath 175 Reputation points Microsoft External Staff Moderator
2026-10-01T09:20:10.42+00:00

Hi @Laura Marchesini ,

Thank you for the detailed explanation.

Azure Speech SSML provides controls for pitch, volume, and other prosodic properties, but these values are interpreted as guidance for the selected neural voice rather than as guaranteed acoustic targets. Therefore, setting values such as <prosody pitch="medium"> or a numeric pitch/volume value does not guarantee that the generated syllable will have exactly the same measured F0 or intensity across separate synthesis requests.

For your use case, the recommended approach is to:

  1. Keep the sentence structure, SSML, voice, and target-word context identical across productions as much as possible, since the neural TTS model also considers context and prosody.
  2. Use <phoneme> when you need to control the pronunciation of the target vowel/word.
  3. Use prosody pitch and volume to adjust the overall prosodic characteristics, understanding that these are not fixed measured F0/intensity controls.

If your experimental requirement is to hold the measured F0 and intensity of a syllable constant across all productions, Azure standard neural TTS does not provide a parameter that guarantees that level of acoustic consistency. You would therefore need to synthesize the audio and perform acoustic measurement/post-processing externally to enforce exact F0 and intensity targets.

For pronunciation consistency specifically, Microsoft also documents custom lexicons as a way to consistently control pronunciation across TTS calls, although a lexicon does not provide direct control over F0 or intensity.

References: SSML prosody, SSML pronunciation

Was this answer helpful?

1 person found this answer helpful.
0 comments No comments

1 additional answer

Sort by: Most helpful
  1. AI answer

    2026-10-01T09:05:01.3233333+00:00

    Azure Speech Studio TTS does not provide a way to guarantee identical fundamental frequency or intensity across repeated generations.

    What is supported:

    1. Use SSML for the highest available control. SSML is the recommended way to fine-tune pronunciation, intonation, pitch, speaking rate, volume, pauses, and related prosodic features.
    2. Apply prosody only to the target segment. Instead of setting pitch and volume for the whole sentence, wrap only the target word or syllable in SSML such as <prosody> so the surrounding carrier phrase stays less affected.
    3. Use phoneme markup for pronunciation control. SSML supports pronunciation control with phonemes and custom lexicon, which is the correct approach for keeping the vowel target consistent at the pronunciation level.
    4. Use Audio Content Creation in Speech Studio for iterative tuning. In Speech Studio, Audio Content Creation allows adjustment of pronunciation, break, pitch, rate, intonation, voice style, and more, then previewing and exporting the tuned result.

    Important limitation:

    • The service documentation states that even with state-of-the-art models, text-to-speech can still produce errors such as subtle unnatural intonations or pronunciation issues.
    • The documented controls are for fine-tuning output quality and prosody, not for enforcing acoustically identical productions across runs.
    • pitch="medium" and volume="medium" are relative synthesis controls, not a guarantee of fixed measured F0 or fixed measured intensity in the exported waveform.

    A practical approach for this scenario:

    1. Keep using SSML rather than plain text.
    2. Keep the carrier sentence fixed.
    3. Wrap only the target word in <phoneme> and, if needed, <prosody>.
    4. Use Audio Content Creation to preview and tune the output before export.
    5. Export the tuned audio and reuse the saved/exported result instead of regenerating the same item repeatedly if acoustic consistency is required for analysis.

    Example pattern:

    <speak version="1.0" xmlns="http://www.w3.org/2001/10/synthesis" xml:lang="en-GB">
      <voice name="en-US-AvaMultilingualNeural">
        Now I read 
        <phoneme alphabet="ipa" ph="kɪd">
          <prosody pitch="high" volume="+20.00%">kid</prosody>
        </phoneme>
        .
      </voice>
    </speak>
    

    If the goal is phonetic/acoustic research with constant F0 and intensity across tokens, the documented Azure Speech Studio controls support tuning, but not precise laboratory-style parameter locking across productions.


    AI-generated content may be incorrect. Read our transparency notes for more information.

    Was this answer helpful?

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.