Azure 內容理解技能

Note

Azure AI 搜尋服務 可透過 Azure 入口網站、REST API 及 Azure SDK 取得。 它同時也是 Foundry IQ 的基礎,這是一個管理式知識層,能將企業內容轉化為可重複使用、權限感知的知識庫,供 Microsoft Foundry 入口網站中的代理使用。

Important

標記(預覽)的功能、能力或屬性不受服務等級協議涵蓋,也不建議用於生產工作負載,且在正式上架前可能會有所變動或受限。 Azure AI 搜尋服務 預覽詞適用於所有預覽功能,無論是獨立功能還是一般可用功能的一部分。

Azure內容理解技能使用Azure Foundry Tools的文件分析器 內容理解工具/c2分析非結構化文件及其他內容類型,產生有組織且可搜尋的輸出,整合進自動化工作負載中。 此技能能擷取文字與圖片,包含保存每張圖片在文件中位置的位置元資料。 影像與相關內容的接近性對於 多模態搜尋、 代理檢索及檢 索增強生成 (RAG)特別有用。

Azure內容理解技能綁定於 Foundry 的 計費Microsoft資源。 與其他Azure AI 資源技能(如 文件佈局技能不同,Azure 內容理解技能並不會每天提供每位索引員免費 20 份文件。 此技能的執行以Azure內容理解價格收費。

你可以使用 Azure 內容理解技能來擷取內容和分塊。 你的技能組裡不需要使用文字分割技能。 此技能實作與文件版面技能相同的介面,後者在 Foundry Tools 版面模型中使用 Azure文件智慧當 設為 。 然而,Azure 內容理解技能相較於文件佈局技能有多項優勢:

  • 表格與圖表以 Markdown 格式輸出,讓大型語言模型(LLM)更容易理解。 相較之下,文件排版技能會將表格和圖表輸出為純文字,可能導致資訊遺失。

  • 對於跨頁的資料表,Azure 內容理解技能能將跨頁資料表識別並擷取為單一單元。

  • Azure 內容理解技能允許區塊透過語意單元跨越多個頁面。

  • Azure 內容理解技能比文件佈局技能更具成本效益,因為內容理解 API 成本較低。

  • Azure Content Understanding 能為圖片、圖表、圖表及嵌入圖表產生基於 AI 的描述。 嵌入的圖形描述直接整合到可檢索的折扣內容中。 這些描述可搜尋,能提升 RAG 的接地與多模態檢索品質。

Azure內容理解技能通常可在 2026-04-01 REST API 中取得。 從 開始 2026-05-01-preview,這項技能可選擇性地產生基於 AI 的圖片描述,用於文件嵌入的圖片、圖表及圖表(預覽)。 要啟用描述,您必須在技能組附帶的 Foundry 資源中部署 Azure OpenAI 聊天完成模型。 此 API 版本也新增語意分塊(預覽),一種尊重段落邊界並以標記衡量區塊長度的版面意識選項。 這兩種功能都需要選擇加入。 當省略新參數時,技能的行為與穩定 2026-04-01 版 API 相同。

局限性

Azure 內容理解技能有以下限制:

  • 這項技能不適合需要在內容理解文件分析器中花費超過五分鐘處理的大型文件。 技能會逾時,但消耗仍會套用到與技能組相連的鑄造廠資源。 確保文件優化以維持在處理限制內,以避免不必要的成本。

  • 此技能稱為 Azure 內容理解文件分析器,因此不同文件類型<服務>的所有文件中已記錄的行為皆適用於其輸出。 例如,Word(DOCX)和 PDF 檔案可能因影像處理方式不同而產生不同結果。 若需要在 DOCX 與 PDF 間保持一致的影像行為,請考慮將文件轉換為 PDF 或檢視 多模態搜尋文件 以尋找替代方案。

支援的區域

Azure內容理解技能呼叫 Content Understanding 2025-11-01 REST API。 您的 Foundry 資源必須位於支援區域,該區域詳見 Azure內容理解區域與語言支援。

您的搜尋服務可以位於任何支援的Azure AI 搜尋服務區域。 當你的 Foundry 資源與 Azure AI 搜尋服務 服務不在同一區域時,跨區域網路延遲會影響索引器的效能。

支援的檔案格式

Azure 內容理解技能可識別以下檔案格式:

  • .PDF
  • .JPEG
  • .JPG
  • .PNG
  • .BMP
  • .HEIF
  • .多倫多國際影展
  • .DOCX
  • .XLSX
  • .PPTX
  • .HTML
  • .TXT
  • .醫學博士
  • .RTF
  • .EML

支援的語言

關於印刷文本,請參見 Azure 內容理解區域與語言支援。

@odata.type

Microsoft.Skills.Util.ContentUnderstandingSkill

資料限制

  • 即使分析文件的檔案大小在 200 MB 限制內,如 Azure 內容理解服務的配額與限制 所述,索引仍須遵守搜尋服務層級的 indexer 限制。

  • 影像尺寸必須介於 50 像素 x 50 像素或 10,000 像素 x 10,000 像素之間。

  • 如果你的 PDF 有密碼鎖定,啟動索引器前先解除鎖定。

技能參數

參數依大小寫區分。

參數名稱 允許的值 Description
extractionOptions ["images"], ["images", "locationMetadata"], ["locationMetadata"] 找出從文件中提取的任何額外內容。 定義一個對應輸出內容的列舉陣列。 例如,若 extractionOptions 為 , ["images", "locationMetadata"]輸出包含影像與位置元資料,提供頁面位置及內容擷取地點的視覺資訊。
modelName (預告) 字串,例如 "gpt-4.1"。 Optional. 自 REST API 起 2026-05-01-preview 即可提供。 Azure OpenAI 聊天完成模型的名稱,用於產生嵌入圖片、圖表和圖表的描述。 影像描述獨立於此 extractionOptions ,且可在不擷取影像的情況下啟用。 必須與 modelDeployment一起指定。 有關支援模型的清單,請參見 支援的生成模型。
modelDeployment (預告) 字串。 Optional. 自 REST API 起 2026-05-01-preview 即可提供。 這是 Foundry 資源中 Foundry 技能組中 Azure OpenAI 模型的部署名稱。 必須與 modelName一起指定。
chunkingProperties 請參見下表。 有選項可以封裝如何將文字內容分割。
chunkingProperties 參數 允許的值 Description
method fixedSize (預設)或 semantic (預覽)。 自 REST API 起 2026-05-01-preview 即可提供。 分塊策略。 fixedSize 使用基於字元的視窗區塊。 semantic 使用尊重段落邊界的版面感知區塊,並聰明地處理跨越區塊邊界的大型資料表。
unit characters (其中 fixedSize)或 tokens (預覽版,與 semantic,從 2026-05-01-preview REST API 開始提供)。 控制區塊單元的基數。 目前只 fixedSize + characters 支援與 semantic + tokens 組合。 若unit省略,則由 推斷。method
maximumLength 當 unit 是 characters,介於 300 到 50,000 之間的整數。 當 unit 是 tokens,介於 100 到 8,000 之間的整數。 預設值為 500。 最大區塊長度,以配置 unit的 為單位。
overlapLength Integer. 該值必須小於 的 maximumLength一半。 兩個文字區塊之間的重疊長度。 僅在 method 為 時 fixedSize適用。 必須省略或設定為 0 當 method 時 semantic。

技能輸入

輸入名稱 Description
file_data 應該從中擷取內容的檔案。

file_data輸入必須是一個定義為:

{
  "$type": "file",
  "data": "BASE64 encoded string of the file"
}

或者,也可以定義為:

{
  "$type": "file",
  "url": "URL to download the file",
  "sasToken": "OPTIONAL: SAS token for authentication if the provided URL is for a file in blob storage"
}

檔案參考物件可透過以下方式之一產生:

  • 將你的索引器定義參數設 allowSkillsetToReadFileData 為 true。 這個設定會建立一個 /document/file_data 路徑,作為一個物件,代表從你的 blob 資料來源下載的原始檔案資料。 此參數僅適用於 Azure Blob 儲存體 中的檔案。

    allowSkillsetToReadFileData 讓下載的檔案資料可供技能使用。 它不會增加 blob 索引器的限制 ,也不會增加 資料限制中描述的內容理解服務限制。

  • 擁有自訂技能回傳一個 JSON 物件定義,該定義提供 $type、 data、 和 urlsastoken。 $type參數必須設為 file,且data必須是檔案內容的 64 編碼位元組陣列。 url參數必須是有效的 URL,且有權限在該位置下載檔案。

技能產出

輸出名稱 Description
text_sections 一組文字區塊物件。 每個區塊可以跨越多頁(考慮更多區塊配置)。 文本區塊物件若適用,則包含 locationMetadata 區塊與文件中圖形跨度重疊時的 imagePath 清單。
normalized_images 僅在 extractionOptions 包含 images時適用。 從文件中擷取的圖片集合,包括 locationMetadata 如適用的圖片。

每個 text_sections 元素的欄位如下:

Field 類型 Description
id String 區塊的唯一識別碼。
content String 區塊的降價內容。 當 method 為 semantic時,內容包含 AI 生成的圖形與表格描述,並內嵌為 Markdown。
locationMetadata 物件 頁距與位置資料(pageNumberFrom, pageNumberTo, ordinalPosition, source)。 當 extractionOptions 包含 locationMetadata時 。
imagePath String 以分號分隔的影像路徑清單,包含在區塊中。 當區塊與文件中的圖形跨度重疊時,會顯示。

每個 normalized_images 元素的欄位如下:

Field 類型 Description
id String 影像的唯一識別碼。
data String Base64 編碼的影像資料。
imagePath String 文件中對影像的路徑參考,例如 "figures/0"。
locationMetadata 物件 頁面範圍與位置資料。 當 extractionOptions 包含 locationMetadata時 。

範例

第一個範例使用固定大小區塊,示範如何以固定大小區塊輸出文字內容,並從文件中擷取圖片及位置元資料。 第二個範例從 REST API 開始 2026-05-01-preview ,使用了 AI 生成的影像描述的語意分塊。

範例 1:固定大小分塊,並擷取影像與元資料

{
  "skills": [
    {
      "description": "Analyze a document",
      "@odata.type": "#Microsoft.Skills.Util.ContentUnderstandingSkill",
      "context": "/document",
      "extractionOptions": ["images", "locationMetadata"],
      "chunkingProperties": {     
          "unit": "characters",
          "maximumLength": 1325, 
          "overlapLength": 0
      },
      "inputs": [
        {
          "name": "file_data",
          "source": "/document/file_data"
        }
      ],
      "outputs": [
        { 
          "name": "text_sections", 
          "targetName": "text_sections" 
        }, 
        { 
          "name": "normalized_images", 
          "targetName": "normalized_images" 
        } 
      ]
    }
  ]
}

範例輸出

{
  "text_sections": [
      {
        "id": "1_d4545398-8df1-409f-acbb-f605d851ae85",
        "content": "What is Azure Content Understanding (preview)?09/16/2025Important· Azure Al Content Understanding is available in preview. Public preview releases provide early access to features that are in active development.· Features, approaches, and processes can change or have limited capabilities, before General Availability (GA).. For more information, see Supplemental Terms of Use for Microsoft Azure PreviewsAzure Content Understanding is a Foundry Tool that uses generative AI to process/ingest content of many types (documents, images, videos, and audio) into a user-defined output format.Content Understanding offers a streamlined process to reason over large amounts of unstructured data, accelerating time-to-value by generating an output that can be integrated into automation and analytical workflows.<figure>\n\nInputs\n\nAnalyzers\n\nOutput\n\n0\nSearch\n\nContent Extraction\n\nField Extraction\n\nDocuments\n\nNew\n\nAgents\n\nPreprocessing\n\nEnrichments\n\nReasoning\n\nImage\n\nNormalization\n(resolution,\nformats)\n\nSpeaker\nrecognition\n\nGen Al\nContext\nwindows\n\nPostprocessing\nConfidence\nscores\nGrounding\nNormalization\n\nMulti-file input\nReference data\n\nDatabases\n\nVideo\n\nOrientation /\nde-skew\n\nLayout and\nstructure\n\nPrompt tuning\n\nStructured\noutput\n\nAudio\n\nFace grouping\n\nMarkdown or JSON schema\n\nCopilots\n\nApps\n\n\\+\n\nFaurIC\n\n</figure>",
        "locationMetadata": {
          "pageNumberFrom": 1,
          "pageNumberTo": 1,
          "ordinalPosition": 0,
          "source": "D(1,0.6348,0.3598,7.2258,0.3805,7.223,1.2662,0.632,1.2455);D(1,0.6334,1.3758,1.3896,1.3738,1.39,1.5401,0.6338,1.542);D(1,0.8104,2.0716,1.8137,2.0692,1.8142,2.2669,0.8109,2.2693);D(1,1.0228,2.5023,7.6222,2.5029,7.6221,3.0075,1.0228,3.0069);D(1,1.0216,3.1121,7.3414,3.1057,7.342,3.6101,1.0221,3.6165);D(1,1.0219,3.7145,7.436,3.7048,7.4362,3.9006,1.0222,3.9103);D(1,0.6303,4.3295,7.7875,4.3236,7.7879,4.812,0.6307,4.8179);D(1,0.6304,5.0295,7.8065,5.0303,7.8064,5.7858,0.6303,5.7849);D(1,0.635,5.9572,7.8544,5.9573,7.8562,8.6971,0.6363,8.6968);D(1,0.6381,9.1451,5.2731,9.1476,5.2729,9.4829,0.6379,9.4803)"
        }
      },
      ...
      {
        "id": "2_e0e57fd4-e835-4879-8532-73a415e47b0b",
        "content": "<table>\n<tr>\n<th>Application</th>\n<th>Description</th>\n</tr>\n<tr>\n<td>Post-call analytics</td>\n<td>Businesses and call centers can generate insights from call recordings to track key KPIs, improve product experience, generate business insights, create differentiated customer experiences, and answer queries faster and more accurately.</td>\n</tr>\n<tr>\n<th>Application</th>\n<th>Description</th>\n</tr>\n<tr>\n<td>Media asset management</td>\n<td>Software and media vendors can use Content Understanding to extract richer, targeted information from videos for media asset management solutions.</td>\n</tr>\n<tr>\n<td>Tax automation</td>\n<td>Tax preparation companies can use Content Understanding to generate a unified view of information from various documents and create comprehensive tax returns.</td>\n</tr>\n<tr>\n<td>Chart understanding</td>\n<td>Businesses can enhance chart understanding by automating the analysis and interpretation of various types of charts and diagrams using Content Understanding.</td>\n</tr>\n<tr>\n<td>Mortgage application processing</td>\n<td>Analyze supplementary supporting documentation and mortgage applications to determine whether a prospective home buyer provided all the necessary documentation to secure a mortgage.</td>\n</tr>\n<tr>\n<td>Invoice contract verification</td>\n<td>Review invoices and contr",
        "locationMetadata": {
          "pageNumberFrom": 2,
          "pageNumberTo": 3,
          "ordinalPosition": 3,
          "source": "D(2,0.6438,9.2645,7.8576,9.2649,7.8565,10.5199,0.6434,10.5194);D(3,0.6494,0.3919,7.8649,0.3929,7.8639,4.3254,0.6485,4.3232)"
        }
        ...
      }
    ],
    "normalized_images": [
        { 
            "id": "1_335140f1-9d31-4507-8916-2cde758639cb", 
            "data": "aW1hZ2UgMSBkYXRh", 
            "imagePath": "aHR0cHM6Ly9henNyb2xsaW5nLmJsb2IuY29yZS53aW5kb3dzLm5ldC9tdWx0aW1vZGFsaXR5L0NVLnBkZg2/normalized_images_0.jpg",  
            "locationMetadata": {
              "pageNumberFrom": 1,
              "pageNumberTo": 1,
              "ordinalPosition": 0,
              "source": "D(1,0.635,5.9572,7.8544,5.9573,7.8562,8.6971,0.6363,8.6968)"
            }
        },
        { 
            "id": "3_699d33ac-1a1b-4015-9cbd-eb8bfff2e6b4", 
            "data": "aW1hZ2UgMiBkYXRh", 
            "imagePath": "aHR0cHM6Ly9henNyb2xsaW5nLmJsb2IuY29yZS53aW5kb3dzLm5ldC9tdWx0aW1vZGFsaXR5L0NVLnBkZg2/normalized_images_1.jpg",  
            "locationMetadata": {
              "pageNumberFrom": 3,
              "pageNumberTo": 3,
              "ordinalPosition": 1,
              "source": "D(3,0.6353,5.2142,7.8428,5.218,7.8443,8.4631,0.6363,8.4594)"
            } 
        }
    ] 
}

locationMetadata 基於 Azure Content Understanding 提供的 source 屬性。 關於該元素在檔案中如何被編碼的視覺位置資訊,請參見 文件分析:擷取結構化內容。

imagePath 表示儲存影像的相對路徑。 如果技能集中已設定知識庫檔案投影,該路徑會與知識庫中影像的相對路徑相符。

範例 2:帶有圖片描述的語意分塊(預覽)

此範例從 REST API 開始 2026-05-01-preview 提供,使用語意分塊,並產生 AI 生成的嵌入圖片、圖表與圖表描述。 綁定於技能組的 Foundry 資源必須由 識別 modelName 聊天完成模型,並 modelDeployment部署 。

{
  "skills": [
    {
      "description": "Extract and chunk document content with image descriptions",
      "@odata.type": "#Microsoft.Skills.Util.ContentUnderstandingSkill",
      "context": "/document",
      "modelName": "gpt-4.1",
      "modelDeployment": "myGpt41Deployment",
      "extractionOptions": ["images", "locationMetadata"],
      "chunkingProperties": {
        "method": "semantic",
        "unit": "tokens",
        "maximumLength": 500
      },
      "inputs": [
        {
          "name": "file_data",
          "source": "/document/file_data"
        }
      ],
      "outputs": [
        {
          "name": "text_sections",
          "targetName": "text_sections"
        },
        {
          "name": "normalized_images",
          "targetName": "normalized_images"
        }
      ]
    }
  ]
}

透過語意分塊,每個 text_sections 區塊包含包含 AI 生成的圖形與表格描述的 Markdown 內容。 當區塊與一個或多個圖幅重疊時,區塊物件還包含一個 imagePath 欄位,列出對應的影像路徑:

{
  "id": "1_d4545398-8df1-409f-acbb-f605d851ae85",
  "content": "# Architecture overview\n\nThe following diagram summarizes the ingestion pipeline...\n\n<figure>The diagram shows three stages: Inputs, Analyzers, and Output. Inputs include documents, images, video, and audio. Analyzers perform preprocessing, enrichments, and reasoning. Output is structured Markdown or JSON consumed by search, agents, copilots, and apps.</figure>",
  "locationMetadata": {
    "pageNumberFrom": 1,
    "pageNumberTo": 1,
    "ordinalPosition": 0,
    "source": "D(1,0.6348,0.3598,7.2258,0.3805,7.223,1.2662,0.632,1.2455)"
  },
  "imagePath": "aHR0cHM6Ly9henNyb2xsaW5nLmJsb2IuY29yZS53aW5kb3dzLm5ldC9tdWx0aW1vZGFsaXR5L0NVLnBkZg2/normalized_images_0.jpg"
}