Hello @Liu, Mingxing | Barry | CNTD ,
Welcome to Microsoft Q&A .Thank you for reaching out to us.
The observed behavior is consistent with a common OCR challenge where official seals or stamps overlap printed content, reducing the visible information available for recognition. When character strokes are partially covered, low-contrast, distorted or mixed with graphical elements, OCR accuracy may decrease, resulting in missing or incorrectly recognized text.
Azure AI Document Intelligence processes the information available in the submitted document image; therefore, if an opaque stamp completely covers the underlying characters, the original text information may no longer be available for extraction, and no OCR engine can reliably reconstruct content that is not visible.
Currently, there is no documented Azure AI Document Intelligence setting or prebuilt model option specifically designed to remove seals/stamps or recover fully hidden text. However, practical mitigation approaches can help improve recognition where the underlying text remains partially visible, including validating whether the issue occurs during OCR recognition or field extraction, improving document quality, applying preprocessing techniques, evaluating available OCR capabilities and implementing confidence-based validation workflows.
Please check if the following steps help-
- Isolating whether the issue occurs during OCR recognition or field extraction The first step is to determine whether the missing or incorrect text originates from the OCR layer or from the extraction model. Recommended validation:
- Process the affected document using the Read model.
- Process the same document using the Layout model.
- Process the document using the applicable prebuilt model:
- Invoice
- Receipt
- Contract
- ID
- W-2
- 1098 Tax forms
- Health insurance card
- Compare the extracted results across the different analysis paths.
Expected outcome:
- If the text is already missing or incorrect in the Read output, the issue is occurring during OCR recognition due to the visual impact of the stamp.
- If the text appears correctly in Read/Layout but is missing from the prebuilt model output, further investigation should focus on field extraction behavior.
- Comparing outputs across models helps identify the specific processing stage affected.
- Validating document quality and stamp characteristics OCR accuracy depends significantly on the quality and clarity of the input document. Recommended checks:
- Use the highest-quality scan or original document available.
- Avoid excessive compression, blur, shadows, background noise, and skew.
- Ensure the affected text remains visually readable around the stamped area.
- Compare stamped and unstamped versions of the same document, if available.
- Review stamp characteristics:
- Color
- Transparency
- Opacity
- Digital versus physical stamp
- Location relative to the affected text
These improvements may increase OCR reliability when text information is still available. However, they cannot recover characters that are completely covered by an opaque stamp.
- Evaluating image preprocessing before OCR processing
ince the deployment includes Web and Container scenarios, preprocessing can be introduced before submitting documents to Azure AI Document Intelligence.
Possible preprocessing techniques:
- Noise reduction
- Contrast enhancement
- Background cleanup
- Deskewing and orientation correction
- Image sharpening
- Color filtering
- Stamp suppression techniques
For example, if stamps are consistently red or blue while the document text is black, application-side color filtering may reduce the visual impact of the stamp and improve text visibility.
Important considerations:
- Preprocessing is performed outside Azure AI Document Intelligence.
- Results depend on stamp color, transparency, document quality, and text visibility.
- Preprocessing is generally more effective for semi-transparent stamps where underlying characters are still partially visible.
- Preprocessing cannot reliably recover text that is completely obscured.
- Testing OCR high-resolution processing capability If supported by the selected Document Intelligence API version and processing scenario, testing the ocrHighResolution add-on capability may improve OCR results for certain challenging documents. Potential scenarios where it may help:
- Small text
- Complex layouts
- Documents containing mixed graphical and text elements
- Difficult OCR regions
Please note that - ocrHighResolution should not be considered a stamp removal feature. It cannot reconstruct text that is fully hidden by a seal or stamp, but it may improve recognition where parts of the original characters remain visible.
- Implementing confidence-based validation workflow For business-critical documents such as invoices, contracts, receipts, tax documents, and application forms, confidence-based validation can help identify potentially affected content. Recommended approach:
- Review extracted words and field confidence scores.
- Compare confidence values for text near stamped areas with unaffected areas.
- Identify low-confidence words or fields.
- Route only impacted documents or fields for additional verification.
This approach allows automated processing for reliable cases while focusing manual review only on areas affected by OCR degradation.
- Verifying API version and model version Confirming the API version currently used is recommended, especially when testing OCR behavior across different environments. Recommended checks:
- Verify the Document Intelligence API version.
- Confirm that the latest supported version is being evaluated where possible.
- Compare behavior between Web and Container deployments if both environments are used.
- Considering custom models for recurring document layouts
If the affected documents follow consistent layouts and stamps appear in predictable locations, custom models may improve extraction consistency. However, it is important to distinguish between:
- Improving extraction accuracy for known document structures.
- Recovering information that is no longer visually available.
Custom models can improve extraction behavior for recurring layouts, but they cannot reliably recover characters that are completely hidden beneath a stamp.
The following references might be helpful , please check them out
Please let us know if the response was helpful
Thank you