An Azure service that integrates speech processing into apps and services.
Hi @Laura Marchesini ,
Thank you for the detailed explanation.
Azure Speech SSML provides controls for pitch, volume, and other prosodic properties, but these values are interpreted as guidance for the selected neural voice rather than as guaranteed acoustic targets. Therefore, setting values such as <prosody pitch="medium"> or a numeric pitch/volume value does not guarantee that the generated syllable will have exactly the same measured F0 or intensity across separate synthesis requests.
For your use case, the recommended approach is to:
- Keep the sentence structure, SSML, voice, and target-word context identical across productions as much as possible, since the neural TTS model also considers context and prosody.
- Use
<phoneme>when you need to control the pronunciation of the target vowel/word. - Use
prosodypitchandvolumeto adjust the overall prosodic characteristics, understanding that these are not fixed measured F0/intensity controls.
If your experimental requirement is to hold the measured F0 and intensity of a syllable constant across all productions, Azure standard neural TTS does not provide a parameter that guarantees that level of acoustic consistency. You would therefore need to synthesize the audio and perform acoustic measurement/post-processing externally to enforce exact F0 and intensity targets.
For pronunciation consistency specifically, Microsoft also documents custom lexicons as a way to consistently control pronunciation across TTS calls, although a lexicon does not provide direct control over F0 or intensity.
References: SSML prosody, SSML pronunciation