An Azure service that turns documents into usable data. Previously known as Azure Form Recognizer.
Hello @Liu, Mingxing | Barry | CNTD ,
Welcome to Microsoft Q&A .Thank you for reaching out to us.
The observed behavior is consistent with a common OCR challenge where official seals or stamps overlap printed content, reducing the visible information available for recognition. When character strokes are partially covered, low-contrast, distorted or mixed with graphical elements, OCR accuracy may decrease, resulting in missing or incorrectly recognized text.
Azure AI Document Intelligence processes the information available in the submitted document image; therefore, if an opaque stamp completely covers the underlying characters, the original text information may no longer be available for extraction, and no OCR engine can reliably reconstruct content that is not visible.
Currently, there is no documented Azure AI Document Intelligence setting or prebuilt model option specifically designed to remove seals/stamps or recover fully hidden text. However, practical mitigation approaches can help improve recognition where the underlying text remains partially visible, including validating whether the issue occurs during OCR recognition or field extraction, improving document quality, applying preprocessing techniques, evaluating available OCR capabilities and implementing confidence-based validation workflows.
Please check if the following steps help-
- Isolating whether the issue occurs during OCR recognition or field extraction The first step is to determine whether the missing or incorrect text originates from the OCR layer or from the extraction model. Recommended validation:
- Process the affected document using the Read model.
- Process the same document using the Layout model.
- Process the document using the applicable prebuilt model:
- Invoice
- Receipt
- Contract
- ID
- W-2
- 1098 Tax forms
- Health insurance card
- Compare the extracted results across the different analysis paths.
- If the text is already missing or incorrect in the Read output, the issue is occurring during OCR recognition due to the visual impact of the stamp.
- If the text appears correctly in Read/Layout but is missing from the prebuilt model output, further investigation should focus on field extraction behavior.
- Comparing outputs across models helps identify the specific processing stage affected.
- Validating document quality and stamp characteristics OCR accuracy depends significantly on the quality and clarity of the input document. Recommended checks:
- Use the highest-quality scan or original document available.
- Avoid excessive compression, blur, shadows, background noise, and skew.
- Ensure the affected text remains visually readable around the stamped area.
- Compare stamped and unstamped versions of the same document, if available.
- Review stamp characteristics:
- Color
- Transparency
- Opacity
- Digital versus physical stamp
- Location relative to the affected text
- Evaluating image preprocessing before OCR processing
ince the deployment includes Web and Container scenarios, preprocessing can be introduced before submitting documents to Azure AI Document Intelligence.
Possible preprocessing techniques:
- Noise reduction
- Contrast enhancement
- Background cleanup
- Deskewing and orientation correction
- Image sharpening
- Color filtering
- Stamp suppression techniques
For example, if stamps are consistently red or blue while the document text is black, application-side color filtering may reduce the visual impact of the stamp and improve text visibility.
Important considerations:
- Preprocessing is performed outside Azure AI Document Intelligence.
- Results depend on stamp color, transparency, document quality, and text visibility.
- Preprocessing is generally more effective for semi-transparent stamps where underlying characters are still partially visible.
- Preprocessing cannot reliably recover text that is completely obscured.
- Testing OCR high-resolution processing capability If supported by the selected Document Intelligence API version and processing scenario, testing the ocrHighResolution add-on capability may improve OCR results for certain challenging documents. Potential scenarios where it may help:
- Small text
- Complex layouts
- Documents containing mixed graphical and text elements
- Difficult OCR regions
- Implementing confidence-based validation workflow For business-critical documents such as invoices, contracts, receipts, tax documents, and application forms, confidence-based validation can help identify potentially affected content. Recommended approach:
- Review extracted words and field confidence scores.
- Compare confidence values for text near stamped areas with unaffected areas.
- Identify low-confidence words or fields.
- Route only impacted documents or fields for additional verification.
- Verifying API version and model version Confirming the API version currently used is recommended, especially when testing OCR behavior across different environments. Recommended checks:
- Verify the Document Intelligence API version.
- Confirm that the latest supported version is being evaluated where possible.
- Compare behavior between Web and Container deployments if both environments are used.
- Considering custom models for recurring document layouts
- Improving extraction accuracy for known document structures.
- Recovering information that is no longer visually available.
The following references might be helpful , please check them out
- Read model OCR data extraction - Document Intelligence - Foundry Tools | Microsoft Learn
- Document layout analysis - Document Intelligence - Foundry Tools | Microsoft Learn
- Document Intelligence APIs analyze document response - Foundry Tools | Microsoft Learn
- Document Processing Models - Document Intelligence - Foundry Tools | Microsoft Learn
- Capabilities and limitations of optical character recognition (OCR) - Azure Vision in Foundry Tools - Foundry Tools | Microsoft Learn
- Add-on capabilities - Document Intelligence - Foundry Tools | Microsoft Learn
- Interpret and improve model accuracy and confidence scores - Foundry Tools | Microsoft Learn
- Capabilities and limitations of optical character recognition (OCR) - Azure Vision in Foundry Tools - Foundry Tools | Microsoft Learn
- Custom document models - Document Intelligence - Foundry Tools | Microsoft Learn
Please let us know if the response was helpful
Thank you