Azure AI Translator glossary: no inflections or derived forms recognized and casing issues for headings (DE → EN)

Stefan Müller 20 Zuverlässigkeitspunkte
2026-02-04T12:13:20.7333333+00:00

I am using Azure AI Translator with a custom glossary and I am facing two related issues:

Glossary terms are only matched in their exact form The glossary works only for the exact terms defined in it. Plurals or derived forms are not recognized. For example, if my glossary contains “Server”, the translation does not apply to “Servers”, “Servern”, or other inflected/derived forms. Is there a way to make the glossary match morphological variants (plural, declensions, derived words), or is this a known limitation?

Casing problem for German → English translations (especially Word headings) When translating from German to English, glossary terms and headings are often lowercased. In Word documents, this is a problem because:

I want headings/titles to start with uppercase letters

But I want the same terms in normal sentences to stay lowercase (or follow normal sentence casing)

Currently, the glossary seems to enforce one casing everywhere, which breaks the formatting of headings after translation.

Questions:

Is there a way to control casing behavior depending on context (e.g., headings vs. body text)?

Are there recommended best practices for handling capitalization with glossaries in document translation (especially for Word files)?

Is there any roadmap or workaround for better morphological matching and casing control?

Any guidance or official recommendations would be greatly appreciated.

Azure Translator in Foundry Tools
Azure Translator in Foundry Tools

Ein Azure-Dienst zum einfachen Ausführen einer maschinellen Übersetzung mit einem einfachen REST-API-Aufruf.


Antwort, die vom Frageautor angenommen wurde
Anonym
2026-02-04T12:27:40.6066667+00:00

Hi Stefan Müller

Glossary matching for plurals, inflected, and derived forms:

Azure AI Translator glossaries (Document Translation, Dynamic Dictionary, and phrase dictionaries) work as exact, literal find‑and‑replace rules. They do not perform morphological analysis, stemming, lemmatization, or plural/inflection expansion.

This means a glossary entry such as “Server → Server” will match only that exact surface form; related forms like Servers, Servern, Serversysteme, or other derived/compound forms are not matched automatically. This behavior is documented and intentional, not a defect, and applies equally to German→English translations. [learn.microsoft.com], [learn.microsoft.com]

limitation exists:

Glossaries in Azure Translator are designed to be deterministic and predictable. They are applied before or alongside the neural model as forced replacements, not as linguistic rules. Because of this design, glossaries deliberately avoid “smart” linguistic expansion that could introduce ambiguity or inconsistent output. Morphology and fluency are instead handled by the base neural MT model, not by glossary logic. [learn.microsoft.com], [learn.microsoft.com]

Recommended workarounds for morphological variants :

The officially recommended workaround is to explicitly enumerate all required variants in the glossary (for example: Server, Servers, Serversysteme, etc.). In advanced enterprise pipelines, teams often generate these variants using external German NLP tooling and expand the glossary offline before submitting it to Translator. Using Custom Translator with a neural phrase dictionary can improve overall fluency, but glossary enforcement itself still remains exact‑match and case‑sensitive.

Casing problems in Word headings vs body text

Azure Translator applies glossary replacements without awareness of document structure (for example, Word headings vs body paragraphs). The service does not know whether a matched term appears in a heading, title, or normal sentence. As a result, if your glossary defines a term in lowercase, it will also be inserted in lowercase in Word headings, even though the heading style visually expects capitalization. This behavior follows the documented rule that glossary matching is case‑sensitive and literal. [learn.microsoft.com]

context‑aware casing is not supported:

Document Translation preserves layout and formatting styles (like Word heading styles), but text casing is controlled by the glossary entry itself, not by document context. There is currently no feature to apply different glossary casing rules based on structural context (heading vs body text) across Word, PowerPoint, PDF, or HTML documents. This is a known limitation of the service design. [learn.microsoft.com], [docs.azure.cn]

Best practices for handling capitalization with glossaries :

Microsoft‑aligned best practices are:

  1. Case‑match your glossary entries to the expected source casing (for example, include both Server and server entries if needed).
  2. Post‑process translated Word documents to enforce heading capitalization (Title Case) using Word styles or automation, while leaving body text unchanged.
  3. In high‑fidelity publishing workflows, translate headings and body text separately and then reassemble the document. These approaches are commonly used in regulated and technical translation pipelines. [github.com]

Roadmap and future support: As of early 2026, Microsoft has not published a roadmap commitment for morphology‑aware glossary matching or context‑aware casing control in Azure Translator. Current documentation and public statements continue to position glossaries as exact, deterministic rules, with linguistic intelligence handled by the neural MT model instead. [learn.microsoft.com], [docs.azure.cn]

Morphological/plural/derived form matching: Not supported; exact match only. •

Different casing for headings vs body text: Not supported in glossaries.

Recommended approach: Enumerate variants + case‑matched entries + post‑translation Word formatting. • Roadmap: No public commitment to change this behavior.

Do let me know if you have any further queries.

Thankyou!

War diese Antwort hilfreich?


0 zusätzliche Antworten

Sortieren nach: Älteste

Ihre Antwort

Antworten können von Fragestellenden als „Angenommen“ und von Moderierenden als „Empfohlen“ gekennzeichnet werden, wodurch Benutzende wissen, dass diese Antwort das Problem des Fragestellenden gelöst hat.