使用 Azure 內容理解技能進行內容分塊與向量化

Note

Azure AI 搜尋服務 可透過 Azure 入口網站、REST API 及 Azure SDK 取得。 它同時也是 Foundry IQ 的基礎,這是一個管理式知識層,能將企業內容轉化為可重複使用、權限感知的知識庫,供 Microsoft Foundry 入口網站中的代理使用。

Important

標記(預覽)的功能、能力或屬性不受服務等級協議涵蓋,也不建議用於生產工作負載,且在正式上架前可能會有所變動或受限。 Azure AI 搜尋服務 預覽條款適用於所有預覽功能,無論是獨立功能還是正式推出功能的一部分。

Important

這些功能支援與其他 Microsoft 服務 及第三方服務的連結。 使用這些服務須遵守其各自的條款,可能導致資料處理或儲存超出 Azure 合規邊界,以及資料流入 Azure 合規邊界。

你有責任管理資料是否會超出組織的合規與地理邊界及相關影響,並確保適當的權限、邊界與核准被提供。

你有責任仔細審查並測試你在特定使用情境中所建置的應用程式,並做出所有適當的決策與客製化。 這包括實施你自己負責任的 AI 緩解措施,例如元提示、內容過濾器或其他安全系統,並確保你的應用程式符合適當的品質、可靠性、安全性與可信度標準。 欲了解更多資訊,請參閱Azure AI 搜尋服務透明度說明。

在本文中,您將學習如何使用 Azure 內容理解技能來:

  • 從文件中擷取文字與圖片
  • 產生尊重段落與章節邊界的語意連貫區塊(預覽)
  • 產生圖表、示意圖及其他內嵌圖片的 AI 描述(預覽)
  • 將每個區塊嵌入向量搜尋,並投影到 Azure AI 搜尋服務 索引中

Azure 內容理解技能會在每份文件中回傳一個或多個區塊。 每個區塊包含 Markdown 格式的內容、位置元資料(頁碼與邊界多邊形),以及可選的擷取影像參考。 當你設 chunkingProperties.method 為 semantic時,區塊會跟隨段落和標題的邊界,而不是固定字元範圍。 當您設定 modelName 和 modelDeployment 時,技能會呼叫 Azure OpenAI 對話補全部署,以產生內嵌影像的描述。 技能接著會將這些描述合併到區塊內容中。

本文使用 範例健康保險計畫 PDF 作為說明。 您可以針對任何支援的資料來源執行相同管線,該資料來源會公開 Content Understanding 支援之格式的檔案。

先決條件

  • 任何支援區域中的 Azure AI 搜尋服務 服務。 搜尋服務本身在此情況下並無區域限制。

  • 位於Azure 內容理解技能支援的區域中的 Microsoft Foundry 資源。 影像描述與分塊作業會在鑄造廠資源區域內處理。

  • 附加至技能集以進行計費的 Microsoft Foundry 資源。 Azure 內容理解技能依照 Azure 內容理解定價 計費。

  • (選用) 相同 Foundry 資源中的 Azure OpenAI 對話補全模型部署 (例如 gpt-4.1),用於產生影像描述。 只有想要基於 AI 的圖片描述時才需要。

  • Azure OpenAI 內嵌模型部署 (例如 text-embedding-3-small),由 Azure OpenAI Embedding 技能用於向量化區塊。

  • 一個Azure Blob 儲存體容器,裡面放你想索引的檔案。 本文使用 Blob 資料來源搭配 allowSkillsetToReadFileData 索引子設定 (用於將檔案內容傳遞至 Content Understanding 技能)。

Overview

本文建構了一個一對多的索引管線。 每個來源文件會產生多個搜尋文件(每個區塊一個):

  1. 索引器會從 Azure Blob 儲存體 讀取每個檔案,並透過 /document/file_data 將二進位內容傳給技能組。

  2. Azure 內容理解技能使用語意分塊(預覽)來產生 text_sections。 當 modelName 和 modelDeployment 已設定時,它也會產生 AI 產生的嵌入圖片描述(預覽),並將其內嵌至每個區塊的 Markdown 內容中。

  3. Azure OpenAI 嵌入技能 每個區塊執行一次,並產生區塊內容的向量。

  4. 索引投影會將每個區塊各寫入一份搜尋文件到目標索引中,並將內容、頁面中繼資料、影像參照及向量對應到各欄位。

  5. (可選)knowledge store 會將 normalized_images 投影至 Azure Blob 儲存體,讓用戶端應用程式可透過 URL 擷取這些已擷取的影像。

準備數據檔

Azure 內容理解技能會處理每份文件的二進位內容,因此原始檔案必須是該技能支援的格式。 關於目前的清單,請參閱 內容理解服務限制。 常見支援格式包括 PDF、DOCX、XLSX、PPTX 以及許多影像格式。

將你的檔案上傳到支援的資料來源。 你可以使用 Azure 入口網站、REST API 或 Azure SDK 來建立資料來源。

以下最小請求會建立整個流程中所使用的資料來源。

POST {endpoint}/datasources?api-version=2026-08-01-preview

{
  "name": "my_blob_datasource",
  "type": "azureblob",
  "credentials": {
    "connectionString": "<your-blob-connection-string>"
  },
  "container": {
    "name": "my-container"
  }
}

建立一對多編製索引的索引

每份搜尋文件對應內容理解技能產生的一個區塊。 該指數需要:

  • 一個關鍵欄位(chunk_id)。
  • 一個父欄位,用來標示該區塊parent_id來自哪個來源文件()。
  • 儲存區塊內容、頁面元資料和圖片參考的欄位。
  • 一個用於區塊嵌入的向量場。

以下的索引定義與你在下一節所建立的技能組相符。

{
  "name": "my_content_understanding_index",
  "fields": [
    {
      "name": "chunk_id",
      "type": "Edm.String",
      "key": true,
      "searchable": true,
      "filterable": false,
      "retrievable": true,
      "stored": true,
      "sortable": true,
      "facetable": false,
      "analyzer": "keyword"
    },
    {
      "name": "parent_id",
      "type": "Edm.String",
      "searchable": false,
      "filterable": true,
      "retrievable": true,
      "stored": true,
      "sortable": false,
      "facetable": false
    },
    {
      "name": "title",
      "type": "Edm.String",
      "searchable": true,
      "filterable": false,
      "retrievable": true,
      "stored": true,
      "sortable": false,
      "facetable": false
    },
    {
      "name": "chunk",
      "type": "Edm.String",
      "searchable": true,
      "filterable": false,
      "retrievable": true,
      "stored": true,
      "sortable": false,
      "facetable": false
    },
    {
      "name": "page_number_from",
      "type": "Edm.Int32",
      "searchable": false,
      "filterable": true,
      "retrievable": true,
      "stored": true,
      "sortable": true,
      "facetable": false
    },
    {
      "name": "page_number_to",
      "type": "Edm.Int32",
      "searchable": false,
      "filterable": true,
      "retrievable": true,
      "stored": true,
      "sortable": true,
      "facetable": false
    },
    {
      "name": "image_path",
      "type": "Edm.String",
      "searchable": false,
      "filterable": false,
      "retrievable": true,
      "stored": true,
      "sortable": false,
      "facetable": false
    },
    {
      "name": "text_vector",
      "type": "Collection(Edm.Single)",
      "searchable": true,
      "retrievable": true,
      "stored": false,
      "dimensions": 1536,
      "vectorSearchProfile": "profile"
    }
  ],
  "vectorSearch": {
    "profiles": [
      {
        "name": "profile",
        "algorithm": "algorithm"
      }
    ],
    "algorithms": [
      {
        "name": "algorithm",
        "kind": "hnsw"
      }
    ]
  }
}

定義語意分塊(預覽)與向量化的技能組

目標索引就緒後,定義產生區塊、向量和投影對應並將其饋送至索引的技能集。

這個技能組合包含兩項技能:

  • Azure 內容理解技能 會將每份文件切分成區塊。 將 chunkingProperties.method 設定為 semantic 可讓技能遵循段落和標題的界限。 設定 modelName 並 modelDeployment 啟用 AI 生成的影像描述(預覽),技能會在向量化前將描述內嵌到區塊內容中。 有關支援的聊天完成模型及其他參數細節,請參見 技能參數。

  • Azure OpenAI 嵌入技能會為每個區塊的內容產生向量。

這個技能組用 indexProjections 來將每個區塊對應到獨立的搜尋文件。 如需詳細資訊,請參閱 定義索引投影。

在你發送請求前,先把 <subdomain> 替換成你的 Azure OpenAI 子網域,<Azure OpenAI api key> 換成嵌入資源金鑰,<Foundry resource key> 換成綁定在技能組上的 Foundry 資源金鑰。

POST {endpoint}/skillsets?api-version=2026-08-01-preview

{
  "name": "my_content_understanding_skillset",
  "description": "Semantic chunking, image descriptions, and vectorization with the Azure Content Understanding skill",
  "skills": [
    {
      "@odata.type": "#Microsoft.Skills.Util.ContentUnderstandingSkill",
      "name": "my_content_understanding_skill",
      "context": "/document",
      "modelName": "gpt-4.1",
      "modelDeployment": "my-gpt-4-1-deployment",
      "chunkingProperties": {
        "method": "semantic",
        "unit": "tokens",
        "maximumLength": 500
      },
      "extractionOptions": ["images", "locationMetadata"],
      "inputs": [
        {
          "name": "file_data",
          "source": "/document/file_data"
        }
      ],
      "outputs": [
        {
          "name": "text_sections",
          "targetName": "text_sections"
        },
        {
          "name": "normalized_images",
          "targetName": "normalized_images"
        }
      ]
    },
    {
      "@odata.type": "#Microsoft.Skills.Text.AzureOpenAIEmbeddingSkill",
      "name": "my_azure_openai_embedding_skill",
      "context": "/document/text_sections/*",
      "inputs": [
        {
          "name": "text",
          "source": "/document/text_sections/*/content"
        }
      ],
      "outputs": [
        {
          "name": "embedding",
          "targetName": "text_vector"
        }
      ],
      "resourceUri": "https://<subdomain>.openai.azure.com",
      "deploymentId": "text-embedding-3-small",
      "modelName": "text-embedding-3-small",
      "apiKey": "<Azure OpenAI api key>"
    }
  ],
  "cognitiveServices": {
    "@odata.type": "#Microsoft.Azure.Search.CognitiveServicesByKey",
    "key": "<Foundry resource key>"
  },
  "indexProjections": {
    "selectors": [
      {
        "targetIndexName": "my_content_understanding_index",
        "parentKeyFieldName": "parent_id",
        "sourceContext": "/document/text_sections/*",
        "mappings": [
          {
            "name": "chunk",
            "source": "/document/text_sections/*/content"
          },
          {
            "name": "text_vector",
            "source": "/document/text_sections/*/text_vector"
          },
          {
            "name": "page_number_from",
            "source": "/document/text_sections/*/locationMetadata/pageNumberFrom"
          },
          {
            "name": "page_number_to",
            "source": "/document/text_sections/*/locationMetadata/pageNumberTo"
          },
          {
            "name": "image_path",
            "source": "/document/text_sections/*/imagePath"
          },
          {
            "name": "title",
            "source": "/document/metadata_storage_name"
          }
        ]
      }
    ],
    "parameters": {
      "projectionMode": "skipIndexingParentDocuments"
    }
  }
}

關於內容理解技能的完整參數參考、支援值及驗證規則,請參見 Azure內容理解技能。

Note

本文使用 API 金鑰來保持範例簡潔。 對於生產環境,建議使用受控識別:

若要了解端對端概觀,請參閱 使用角色連線至 Azure AI 搜尋服務。

設定和執行索引子

建立並執行一個索引器,從你的資料來源讀取,呼叫技能集,並將區塊投影到索引中。 設定 allowSkillsetToReadFileData 為 true 內容理解技能接收檔案內容,並設 parsingMode 為 default。

在這種情況下你不需要 outputFieldMappings 。 技能組中的 indexProjections 區塊已將每個分塊對應至目標索引欄位。

POST {endpoint}/indexers?api-version=2026-08-01-preview

{
  "name": "my_content_understanding_indexer",
  "dataSourceName": "my_blob_datasource",
  "targetIndexName": "my_content_understanding_index",
  "skillsetName": "my_content_understanding_skillset",
  "parameters": {
    "batchSize": 1,
    "configuration": {
      "dataToExtract": "contentAndMetadata",
      "parsingMode": "default",
      "allowSkillsetToReadFileData": true
    }
  },
  "fieldMappings": [],
  "outputFieldMappings": []
}

當索引器執行時,內容理解技能會使用語意分塊(預覽),可選擇性產生基於 AI 的圖片描述(預覽),並每個區塊寫入一份搜尋文件到索引。

檢查索引器狀態

在查詢前,請確認索引器的執行完成:

GET {endpoint}/indexers/my_content_understanding_indexer/status?api-version=2026-08-01-preview

請確認lastResult.status是否為success。 如果 transientFailureitemsProcessed 的值大於 0,執行是部分成功,你仍然可以查詢已填充的區塊。 欲了解更多資訊,請參閱 「監控索引器狀態」。

驗證結果

查詢索引以確認區塊是否包含預期內容,且向量搜尋是否如預期運作。 使用 搜尋檔案總管 或任何能發送 HTTP 請求的工具。

以下請求會執行混合式查詢(在 chunk 上進行關鍵字搜尋,並對 text_vector 執行向量查詢),以確認分塊文字和嵌入皆已填入。

POST /indexes/my_content_understanding_index/docs/search?api-version=2026-08-01-preview
{
  "search": "copay for in-network providers",
  "count": true,
  "searchMode": "all",
  "vectorQueries": [
    {
      "kind": "text",
      "text": "copay for in-network providers",
      "fields": "text_vector"
    }
  ],
  "select": "chunk, title, page_number_from, page_number_to, image_path"
}

成功的回應大致如下(為簡潔而刪減):

{
  "@odata.count": 2,
  "value": [
    {
      "@search.score": 0.0317,
      "chunk": "## Cost sharing\n\nFor in-network providers, the copay is $20 per visit...\n\n![Chart: Copay comparison across plans](figures/3)",
      "title": "Northwind_Standard_Benefits_Details.pdf",
      "page_number_from": 4,
      "page_number_to": 4,
      "image_path": "figures/3"
    },
    {
      "@search.score": 0.0289,
      "chunk": "### Out-of-network providers\n\nWhen you visit a provider that isn't in the Northwind network, the copay is $40 per visit...",
      "title": "Northwind_Standard_Benefits_Details.pdf",
      "page_number_from": 5,
      "page_number_to": 6,
      "image_path": null
    }
  ]
}

回應包括:

  • chunk: 每個區塊的 Markdown 內容。 當你設定 modelName 和 modelDeployment時,AI 生成的圖片描述(預覽)會直接顯示在 Markdown 中。
  • page_number_from 以及 page_number_to:產生該區塊的頁數範圍。
  • image_path:用區塊提取到影像的路徑,或當區塊跨多張影像時,則是用分號分隔的路徑清單。 具體形狀取決於是否設定了知識儲存檔案投影。 若無檔案投影,路徑即為範例所示的簡短形式(figures/3)。 在檔案投影中,路徑是知識庫中影像的相對路徑。 若要讓這些影像提供給用戶端應用程式,請參見 (Optional) Project 影像以供檢索。

(可選)用於檢索的專案影像

image_path索引中儲存的值是指向技能豐富樹的指標,而非可直接檢索的網址。 若要擷取映像,請使用知識存放區將 normalized_images 投影到 Azure Blob 儲存體,然後為每個區塊產生對應的 Blob URL。

這個步驟是選擇性的。 只有當你的客戶端應用程式需要顯示或下載擷取後的圖片時才加入。

將以下屬性加入前一節的技能組合載荷中。 技能集要求使用 api-version=2026-08-01-preview。

"knowledgeStore": {
  "storageConnectionString": "<your-azure-storage-connection-string>",
  "projections": [
    {
      "files": [
        {
          "storageContainer": "extracted-images",
          "source": "/document/normalized_images/*"
        }
      ],
      "tables": [],
      "objects": []
    }
  ]
}

索引器執行後,容器中的 extracted-images 每個 blob 對應一個 normalized_images 元素。 blob URL 的形式為 https://<storage-account>.blob.core.windows.net/<container>/<imagePath>,其中 <imagePath> 會對應到儲存在 image_path 欄位中的值。

如需完整的結構描述(包括其他投影類型(tables 和 objects)及驗證選項),請參閱 Azure AI 搜尋服務 中的知識存放區「投影」。

清理資源

完成後,請刪除索引子、技能集和索引,以停止產生 Content Understanding 和 Azure OpenAI 費用。 Azure Blob 儲存體 裡的原始檔案和 Foundry 資源本身會保留,直到你刪除它們為止。

Troubleshooting

如果索引器失敗或回傳意外結果,請檢查以下常見原因。

技能組驗證在 400 級時失敗

當參數組合不一致時,技能會回傳 400 Skill validation failed 錯誤。 常見原因:

  • 已設定 modelName 但未設定 modelDeployment,或反之亦然。 兩者必須同時設置。
  • method 是 semantic (預覽),且 overlapLength 大於 0。 設定 overlapLength 為 0 或省略。
  • method 和 unit 不是受支援的組合。 將 fixedSize 與 characters 搭配使用,或將 semantic 與 tokens 搭配使用。

對 Foundry 資源授權失敗

如果技能呼叫 Foundry 資源時回傳 401 或 403,請確認:

text_sections 是空的

如果索引文件沒有區塊,請確認:

  • 支援此檔案格式。 關於清單,請參見 支援的檔案格式。
  • Foundry 資源位於受支援的區域。
  • 受密碼保護的 PDF 在索引前會先解鎖。

圖片描述(預覽)缺少

如果區塊中沒有包含內嵌圖片描述,請確認:

  • 技能集中已同時設定 modelName 和 modelDeployment。
  • modelName 中的對話補全模型部署在技能集所參考的相同 Foundry 資源中。
  • 此部署的 TPM 或 RPM 配額足以因應您的文件量。

索引子在大型文件上逾時

Content Understanding 會對每份文件的處理施加逾時限制。 如果大型PDF失敗:

  • 在索引前,先將原始文件拆分成較小的檔案。
  • 縮減 batchSize 為 1 每份文件獨立處理。

關於Azure內容理解技能的完整資料限制,請參見資料限制。