An Azure service that integrates speech processing into apps and services.
As you noted, the documentation does not answer the publishing question in either direction. Could this thread please be escalated to the Azure Speech product team, or to whoever owns the MAI-Voice-2.1 preview terms? I need an answer from Microsoft that I can rely on, and I'm happy to wait for it. I've also opened an Azure support request for the same question.
To make it easy to answer, here are the questions as yes/no:
- Does the sentence in the Supplemental Terms of Use for Microsoft Azure Previews, "not intended to be used in production or in a live operating environment", cover audio files already generated with MAI-Voice-2.1 and then published as static files? Or does it cover only applications that call the preview service at runtime?
- If we publish such files during the preview, are we in breach of any Microsoft terms? Yes / No.
- If the answer is "not during preview": may the same files be published unchanged once MAI-Voice-2.1 is generally available, or must they be generated again after GA?
And one question about the voices we already use. It is separate from the preview and concerns files that are already published:
- We call the REST endpoint
cognitiveservices/v1(non-streaming) with the prebuilt voicestr-TR-AhmetNeuralandtr-TR-EmelNeuraland the output formataudio-24khz-48kbitrate-mono-mp3. Does the MP3 that comes back contain AI Content Credentials (a C2PA manifest in an ID3 frame) or an invisible watermark? The Content provenance page lists Azure AI Speech text to speech but says availability "varies by configuration, output format, sample rate, and deployment environment", so I'd like to know for this exact configuration. If it does contain them which frames must we keep unchanged so that we comply with the Code of Conduct, usage restriction 20?