Can pre-generated MAI-Voice-2.1 (public preview) audio be published as static files on a live website?

Alexey Mayer 0 Reputation points
2026-10-07T11:35:31.8433333+00:00

Hello,

I would like to get a written answer on whether audio generated with MAI-Voice-2.1 while it is in public preview may be published on the site(language course in my case), and on one related question about content credentials.

How we use Azure AI Speech today:

  • A paid S0 Speech resource, on a pay-as-you-go subscription.
  • We generate MP3 files offline through the REST API, with tr-TR-AhmetNeural and tr-TR-EmelNeural, from lesson texts we wrote ourselves. Each batch is listened to before it is published.
  • The site serves the files as static content from its own host. Pages never call Azure: no API request is made when a learner plays a clip, no key is in the website, and no learner data is ever sent to the service.

To make it easy to answer, here are the questions as yes/no:

  1. Does the sentence in the Supplemental Terms of Use for Microsoft Azure Previews, "not intended to be used in production or in a live operating environment", cover audio files already generated with MAI-Voice-2.1 and then published as static files? Or does it cover only applications that call the preview service at runtime?
  2. If we publish such files during the preview, are we in breach of any Microsoft terms?
  3. If the answer is "not during preview": may the same files be published unchanged once MAI-Voice-2.1 is generally available, or must they be generated again after GA?

And one question about the voices we already use. It is separate from the preview and concerns files that are already published:

We call the REST endpoint cognitiveservices/v1 (non-streaming) with the prebuilt voices tr-TR-AhmetNeural and tr-TR-EmelNeural and the output format audio-24khz-48kbitrate-mono-mp3. Does the MP3 that comes back contain AI Content Credentials (a C2PA manifest in an ID3 frame) or an invisible watermark? The Content provenance page lists Azure AI Speech text to speech but says availability "varies by configuration, output format, sample rate, and deployment environment", so I'd like to know for this exact configuration. If it does contain them which frames must we keep unchanged so that we comply with the Code of Conduct, usage restriction 20?

Thank you!

Azure Speech in Foundry Tools
0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.