Hello,
I would like to get a written answer on whether audio generated with MAI-Voice-2.1 while it is in public preview may be published on the site(language course in my case), and on one related question about content credentials.
How we use Azure AI Speech today:
- A paid S0 Speech resource, on a pay-as-you-go subscription.
- We generate MP3 files offline through the REST API, with tr-TR-AhmetNeural and tr-TR-EmelNeural, from lesson texts we wrote ourselves. Each batch is listened to before it is published.
- The site serves the files as static content from its own host. Pages never call Azure: no API request is made when a learner plays a clip, no key is in the website, and no learner data is ever sent to the service.
What we would like to do:
Generate dialogue lines for two fictional course characters with the prebuilt voices tr-TR-Elif:MAI-Voice-2.1 and tr-TR-Aydin:MAI-Voice-2.1 (no voice cloning, no custom or personal voice), in the same way: offline, through the same S0 resource, and published as static MP3 files.
Why we are asking:
The MAI-Voice-2.1 documentation says the feature is in public preview and "not recommended for production workloads". The Supplemental Terms of Use for Microsoft Azure Previews say that previews of generative AI features are "not intended to be used in production or in a live operating environment".
Questions:
- May audio generated with MAI-Voice-2.1 during public preview, through a paid S0 resource, be published as static MP3 files on a live public website? Does "not intended to be used in production or in a live operating environment" cover publishing the generated output, or only systems that call the preview service at runtime?
- If it is allowed, are there any conditions, such as specific disclosure wording or attribution?
- Does the Customer Copyright Commitment apply to output generated with a preview model?
- When MAI-Voice-2.1 becomes generally available, do clips generated during the preview need to be generated again?
- Do MP3 files from Azure AI Speech (prebuilt neural voices and MAI-Voice-2.1) carry AI Content Credentials or another provenance mark, for example in an ID3 tag or a watermark in the audio? If they do, what must we keep when we add our own ID3 tag to the file?
A short written answer, or a pointer to the clause that decides each point, would be enough.
To make it easy to answer, here are the questions as yes/no:
- Does the sentence in the Supplemental Terms of Use for Microsoft Azure Previews, "not intended to be used in production or in a live operating environment", cover audio files already generated with MAI-Voice-2.1 and then published as static files? Or does it cover only applications that call the preview service at runtime?
- If we publish such files during the preview, are we in breach of any Microsoft terms? Yes / No.
- If the answer is "not during preview": may the same files be published unchanged once MAI-Voice-2.1 is generally available, or must they be generated again after GA?
And one question about the voices we already use. It is separate from the preview and concerns files that are already published:
- We call the REST endpoint
cognitiveservices/v1 (non-streaming) with the prebuilt voices tr-TR-AhmetNeural and tr-TR-EmelNeural and the output format audio-24khz-48kbitrate-mono-mp3. Does the MP3 that comes back contain AI Content Credentials (a C2PA manifest in an ID3 frame) or an invisible watermark? The Content provenance page lists Azure AI Speech text to speech but says availability "varies by configuration, output format, sample rate, and deployment environment", so I'd like to know for this exact configuration. If it does contain them which frames must we keep unchanged so that we comply with the Code of Conduct, usage restriction 20?To make it easy to answer, here are the questions as yes/no:
- Does the sentence in the Supplemental Terms of Use for Microsoft Azure Previews, "not intended to be used in production or in a live operating environment", cover audio files already generated with MAI-Voice-2.1 and then published as static files? Or does it cover only applications that call the preview service at runtime?
- If we publish such files during the preview, are we in breach of any Microsoft terms? Yes / No.
- If the answer is "not during preview": may the same files be published unchanged once MAI-Voice-2.1 is generally available, or must they be generated again after GA?
And one question about the voices we already use. It is separate from the preview and concerns files that are already published:
- We call the REST endpoint
cognitiveservices/v1 (non-streaming) with the prebuilt voices tr-TR-AhmetNeural and tr-TR-EmelNeural and the output format audio-24khz-48kbitrate-mono-mp3. Does the MP3 that comes back contain AI Content Credentials (a C2PA manifest in an ID3 frame) or an invisible watermark? The Content provenance page lists Azure AI Speech text to speech but says availability "varies by configuration, output format, sample rate, and deployment environment", so I'd like to know for this exact configuration. If it does contain them which frames must we keep unchanged so that we comply with the Code of Conduct, usage restriction 20?
Thank you!