Note
Azure AI 搜尋服務 可透過 Azure 入口網站、REST API 及 Azure SDK 取得。 它同時也是 Foundry IQ 的基礎,這是一個管理式知識層,能將企業內容轉化為可重複使用、權限感知的知識庫,供 Microsoft Foundry 入口網站中的代理使用。
Important
標記(預覽)的功能、能力或屬性不受服務等級協議涵蓋,也不建議用於生產工作負載,且在正式上架前可能會有所變動或受限。 Azure AI 搜尋服務 預覽條款適用於所有預覽功能,無論是獨立功能還是正式推出功能的一部分。
Important
這些功能支援與其他 Microsoft 服務 及第三方服務的連結。 使用這些服務須遵守其各自的條款,可能導致資料處理或儲存超出 Azure 合規邊界,以及資料流入 Azure 合規邊界。
你有責任管理資料是否會超出組織的合規與地理邊界及相關影響,並確保適當的權限、邊界與核准被提供。
你有責任仔細審查並測試你在特定使用情境中所建置的應用程式,並做出所有適當的決策與客製化。 這包括實施你自己負責任的 AI 緩解措施,例如元提示、內容過濾器或其他安全系統,並確保你的應用程式符合適當的品質、可靠性、安全性與可信度標準。 欲了解更多資訊,請參閱Azure AI 搜尋服務透明度說明。
在本文中,您將學習如何使用 Azure 內容理解技能來:
- 從文件中擷取文字與圖片
- 產生尊重段落與章節邊界的語意連貫區塊(預覽)
- 產生圖表、示意圖及其他內嵌圖片的 AI 描述(預覽)
- 將每個區塊嵌入向量搜尋,並投影到 Azure AI 搜尋服務 索引中
Azure 內容理解技能會在每份文件中回傳一個或多個區塊。 每個區塊包含 Markdown 格式的內容、位置元資料(頁碼與邊界多邊形),以及可選的擷取影像參考。 當你設 chunkingProperties.method 為 semantic時,區塊會跟隨段落和標題的邊界,而不是固定字元範圍。 當您設定 modelName 和 modelDeployment 時,技能會呼叫 Azure OpenAI 對話補全部署,以產生內嵌影像的描述。 技能接著會將這些描述合併到區塊內容中。
本文使用 範例健康保險計畫 PDF 作為說明。 您可以針對任何支援的資料來源執行相同管線,該資料來源會公開 Content Understanding 支援之格式的檔案。
先決條件
任何支援區域中的 Azure AI 搜尋服務 服務。 搜尋服務本身在此情況下並無區域限制。
位於Azure 內容理解技能支援的區域中的 Microsoft Foundry 資源。 影像描述與分塊作業會在鑄造廠資源區域內處理。
附加至技能集以進行計費的 Microsoft Foundry 資源。 Azure 內容理解技能依照 Azure 內容理解定價 計費。
(選用) 相同 Foundry 資源中的 Azure OpenAI 對話補全模型部署 (例如
gpt-4.1),用於產生影像描述。 只有想要基於 AI 的圖片描述時才需要。Azure OpenAI 內嵌模型部署 (例如
text-embedding-3-small),由 Azure OpenAI Embedding 技能用於向量化區塊。一個Azure Blob 儲存體容器,裡面放你想索引的檔案。 本文使用 Blob 資料來源搭配
allowSkillsetToReadFileData索引子設定 (用於將檔案內容傳遞至 Content Understanding 技能)。
Overview
本文建構了一個一對多的索引管線。 每個來源文件會產生多個搜尋文件(每個區塊一個):
索引器會從 Azure Blob 儲存體 讀取每個檔案,並透過
/document/file_data將二進位內容傳給技能組。Azure 內容理解技能使用語意分塊(預覽)來產生
text_sections。 當modelName和modelDeployment已設定時,它也會產生 AI 產生的嵌入圖片描述(預覽),並將其內嵌至每個區塊的 Markdown 內容中。Azure OpenAI 嵌入技能 每個區塊執行一次,並產生區塊內容的向量。
索引投影會將每個區塊各寫入一份搜尋文件到目標索引中,並將內容、頁面中繼資料、影像參照及向量對應到各欄位。
(可選)knowledge store 會將
normalized_images投影至 Azure Blob 儲存體,讓用戶端應用程式可透過 URL 擷取這些已擷取的影像。
準備數據檔
Azure 內容理解技能會處理每份文件的二進位內容,因此原始檔案必須是該技能支援的格式。 關於目前的清單,請參閱 內容理解服務限制。 常見支援格式包括 PDF、DOCX、XLSX、PPTX 以及許多影像格式。
將你的檔案上傳到支援的資料來源。 你可以使用 Azure 入口網站、REST API 或 Azure SDK 來建立資料來源。
以下最小請求會建立整個流程中所使用的資料來源。
POST {endpoint}/datasources?api-version=2026-08-01-preview
{
"name": "my_blob_datasource",
"type": "azureblob",
"credentials": {
"connectionString": "<your-blob-connection-string>"
},
"container": {
"name": "my-container"
}
}
建立一對多編製索引的索引
每份搜尋文件對應內容理解技能產生的一個區塊。 該指數需要:
- 一個關鍵欄位(
chunk_id)。 - 一個父欄位,用來標示該區塊
parent_id來自哪個來源文件()。 - 儲存區塊內容、頁面元資料和圖片參考的欄位。
- 一個用於區塊嵌入的向量場。
以下的索引定義與你在下一節所建立的技能組相符。
{
"name": "my_content_understanding_index",
"fields": [
{
"name": "chunk_id",
"type": "Edm.String",
"key": true,
"searchable": true,
"filterable": false,
"retrievable": true,
"stored": true,
"sortable": true,
"facetable": false,
"analyzer": "keyword"
},
{
"name": "parent_id",
"type": "Edm.String",
"searchable": false,
"filterable": true,
"retrievable": true,
"stored": true,
"sortable": false,
"facetable": false
},
{
"name": "title",
"type": "Edm.String",
"searchable": true,
"filterable": false,
"retrievable": true,
"stored": true,
"sortable": false,
"facetable": false
},
{
"name": "chunk",
"type": "Edm.String",
"searchable": true,
"filterable": false,
"retrievable": true,
"stored": true,
"sortable": false,
"facetable": false
},
{
"name": "page_number_from",
"type": "Edm.Int32",
"searchable": false,
"filterable": true,
"retrievable": true,
"stored": true,
"sortable": true,
"facetable": false
},
{
"name": "page_number_to",
"type": "Edm.Int32",
"searchable": false,
"filterable": true,
"retrievable": true,
"stored": true,
"sortable": true,
"facetable": false
},
{
"name": "image_path",
"type": "Edm.String",
"searchable": false,
"filterable": false,
"retrievable": true,
"stored": true,
"sortable": false,
"facetable": false
},
{
"name": "text_vector",
"type": "Collection(Edm.Single)",
"searchable": true,
"retrievable": true,
"stored": false,
"dimensions": 1536,
"vectorSearchProfile": "profile"
}
],
"vectorSearch": {
"profiles": [
{
"name": "profile",
"algorithm": "algorithm"
}
],
"algorithms": [
{
"name": "algorithm",
"kind": "hnsw"
}
]
}
}
定義語意分塊(預覽)與向量化的技能組
目標索引就緒後,定義產生區塊、向量和投影對應並將其饋送至索引的技能集。
這個技能組合包含兩項技能:
Azure 內容理解技能 會將每份文件切分成區塊。 將
chunkingProperties.method設定為semantic可讓技能遵循段落和標題的界限。 設定modelName並modelDeployment啟用 AI 生成的影像描述(預覽),技能會在向量化前將描述內嵌到區塊內容中。 有關支援的聊天完成模型及其他參數細節,請參見 技能參數。Azure OpenAI 嵌入技能會為每個區塊的內容產生向量。
這個技能組用 indexProjections 來將每個區塊對應到獨立的搜尋文件。 如需詳細資訊,請參閱 定義索引投影。
在你發送請求前,先把 <subdomain> 替換成你的 Azure OpenAI 子網域,<Azure OpenAI api key> 換成嵌入資源金鑰,<Foundry resource key> 換成綁定在技能組上的 Foundry 資源金鑰。
POST {endpoint}/skillsets?api-version=2026-08-01-preview
{
"name": "my_content_understanding_skillset",
"description": "Semantic chunking, image descriptions, and vectorization with the Azure Content Understanding skill",
"skills": [
{
"@odata.type": "#Microsoft.Skills.Util.ContentUnderstandingSkill",
"name": "my_content_understanding_skill",
"context": "/document",
"modelName": "gpt-4.1",
"modelDeployment": "my-gpt-4-1-deployment",
"chunkingProperties": {
"method": "semantic",
"unit": "tokens",
"maximumLength": 500
},
"extractionOptions": ["images", "locationMetadata"],
"inputs": [
{
"name": "file_data",
"source": "/document/file_data"
}
],
"outputs": [
{
"name": "text_sections",
"targetName": "text_sections"
},
{
"name": "normalized_images",
"targetName": "normalized_images"
}
]
},
{
"@odata.type": "#Microsoft.Skills.Text.AzureOpenAIEmbeddingSkill",
"name": "my_azure_openai_embedding_skill",
"context": "/document/text_sections/*",
"inputs": [
{
"name": "text",
"source": "/document/text_sections/*/content"
}
],
"outputs": [
{
"name": "embedding",
"targetName": "text_vector"
}
],
"resourceUri": "https://<subdomain>.openai.azure.com",
"deploymentId": "text-embedding-3-small",
"modelName": "text-embedding-3-small",
"apiKey": "<Azure OpenAI api key>"
}
],
"cognitiveServices": {
"@odata.type": "#Microsoft.Azure.Search.CognitiveServicesByKey",
"key": "<Foundry resource key>"
},
"indexProjections": {
"selectors": [
{
"targetIndexName": "my_content_understanding_index",
"parentKeyFieldName": "parent_id",
"sourceContext": "/document/text_sections/*",
"mappings": [
{
"name": "chunk",
"source": "/document/text_sections/*/content"
},
{
"name": "text_vector",
"source": "/document/text_sections/*/text_vector"
},
{
"name": "page_number_from",
"source": "/document/text_sections/*/locationMetadata/pageNumberFrom"
},
{
"name": "page_number_to",
"source": "/document/text_sections/*/locationMetadata/pageNumberTo"
},
{
"name": "image_path",
"source": "/document/text_sections/*/imagePath"
},
{
"name": "title",
"source": "/document/metadata_storage_name"
}
]
}
],
"parameters": {
"projectionMode": "skipIndexingParentDocuments"
}
}
}
關於內容理解技能的完整參數參考、支援值及驗證規則,請參見 Azure內容理解技能。
Note
本文使用 API 金鑰來保持範例簡潔。 對於生產環境,建議使用受控識別:
技能集連線至 Foundry 資源: 若要使用受控識別而非金鑰,將技能集連結至 Foundry 資源,請參閱 將搜尋服務連線至 Azure AI 服務。 當你使用管理身份時,將該
key屬性從技能集的cognitiveServices區塊中省略。技能集至 Azure OpenAI:Azure OpenAI Embedding 技能支援使用受控識別取代
apiKey。Azure Blob 儲存體的索引子: 將連線字串取代為受控識別連線。 請參見 使用管理身份建立與資料來源的連線。
若要了解端對端概觀,請參閱 使用角色連線至 Azure AI 搜尋服務。
設定和執行索引子
建立並執行一個索引器,從你的資料來源讀取,呼叫技能集,並將區塊投影到索引中。 設定 allowSkillsetToReadFileData 為 true 內容理解技能接收檔案內容,並設 parsingMode 為 default。
在這種情況下你不需要 outputFieldMappings 。 技能組中的 indexProjections 區塊已將每個分塊對應至目標索引欄位。
POST {endpoint}/indexers?api-version=2026-08-01-preview
{
"name": "my_content_understanding_indexer",
"dataSourceName": "my_blob_datasource",
"targetIndexName": "my_content_understanding_index",
"skillsetName": "my_content_understanding_skillset",
"parameters": {
"batchSize": 1,
"configuration": {
"dataToExtract": "contentAndMetadata",
"parsingMode": "default",
"allowSkillsetToReadFileData": true
}
},
"fieldMappings": [],
"outputFieldMappings": []
}
當索引器執行時,內容理解技能會使用語意分塊(預覽),可選擇性產生基於 AI 的圖片描述(預覽),並每個區塊寫入一份搜尋文件到索引。
檢查索引器狀態
在查詢前,請確認索引器的執行完成:
GET {endpoint}/indexers/my_content_understanding_indexer/status?api-version=2026-08-01-preview
請確認lastResult.status是否為success。 如果 transientFailureitemsProcessed 的值大於 0,執行是部分成功,你仍然可以查詢已填充的區塊。 欲了解更多資訊,請參閱 「監控索引器狀態」。
驗證結果
查詢索引以確認區塊是否包含預期內容,且向量搜尋是否如預期運作。 使用 搜尋檔案總管 或任何能發送 HTTP 請求的工具。
以下請求會執行混合式查詢(在 chunk 上進行關鍵字搜尋,並對 text_vector 執行向量查詢),以確認分塊文字和嵌入皆已填入。
POST /indexes/my_content_understanding_index/docs/search?api-version=2026-08-01-preview
{
"search": "copay for in-network providers",
"count": true,
"searchMode": "all",
"vectorQueries": [
{
"kind": "text",
"text": "copay for in-network providers",
"fields": "text_vector"
}
],
"select": "chunk, title, page_number_from, page_number_to, image_path"
}
成功的回應大致如下(為簡潔而刪減):
{
"@odata.count": 2,
"value": [
{
"@search.score": 0.0317,
"chunk": "## Cost sharing\n\nFor in-network providers, the copay is $20 per visit...\n\n",
"title": "Northwind_Standard_Benefits_Details.pdf",
"page_number_from": 4,
"page_number_to": 4,
"image_path": "figures/3"
},
{
"@search.score": 0.0289,
"chunk": "### Out-of-network providers\n\nWhen you visit a provider that isn't in the Northwind network, the copay is $40 per visit...",
"title": "Northwind_Standard_Benefits_Details.pdf",
"page_number_from": 5,
"page_number_to": 6,
"image_path": null
}
]
}
回應包括:
-
chunk: 每個區塊的 Markdown 內容。 當你設定modelName和modelDeployment時,AI 生成的圖片描述(預覽)會直接顯示在 Markdown 中。 -
page_number_from以及page_number_to:產生該區塊的頁數範圍。 -
image_path:用區塊提取到影像的路徑,或當區塊跨多張影像時,則是用分號分隔的路徑清單。 具體形狀取決於是否設定了知識儲存檔案投影。 若無檔案投影,路徑即為範例所示的簡短形式(figures/3)。 在檔案投影中,路徑是知識庫中影像的相對路徑。 若要讓這些影像提供給用戶端應用程式,請參見 (Optional) Project 影像以供檢索。
(可選)用於檢索的專案影像
image_path索引中儲存的值是指向技能豐富樹的指標,而非可直接檢索的網址。 若要擷取映像,請使用知識存放區將 normalized_images 投影到 Azure Blob 儲存體,然後為每個區塊產生對應的 Blob URL。
這個步驟是選擇性的。 只有當你的客戶端應用程式需要顯示或下載擷取後的圖片時才加入。
將以下屬性加入前一節的技能組合載荷中。 技能集要求使用 api-version=2026-08-01-preview。
"knowledgeStore": {
"storageConnectionString": "<your-azure-storage-connection-string>",
"projections": [
{
"files": [
{
"storageContainer": "extracted-images",
"source": "/document/normalized_images/*"
}
],
"tables": [],
"objects": []
}
]
}
索引器執行後,容器中的 extracted-images 每個 blob 對應一個 normalized_images 元素。 blob URL 的形式為 https://<storage-account>.blob.core.windows.net/<container>/<imagePath>,其中 <imagePath> 會對應到儲存在 image_path 欄位中的值。
如需完整的結構描述(包括其他投影類型(tables 和 objects)及驗證選項),請參閱 Azure AI 搜尋服務 中的知識存放區「投影」。
清理資源
完成後,請刪除索引子、技能集和索引,以停止產生 Content Understanding 和 Azure OpenAI 費用。 Azure Blob 儲存體 裡的原始檔案和 Foundry 資源本身會保留,直到你刪除它們為止。
Troubleshooting
如果索引器失敗或回傳意外結果,請檢查以下常見原因。
技能組驗證在 400 級時失敗
當參數組合不一致時,技能會回傳 400 Skill validation failed 錯誤。 常見原因:
- 已設定
modelName但未設定modelDeployment,或反之亦然。 兩者必須同時設置。 -
method是semantic(預覽),且overlapLength大於0。 設定overlapLength為0或省略。 -
method和unit不是受支援的組合。 將fixedSize與characters搭配使用,或將semantic與tokens搭配使用。
對 Foundry 資源授權失敗
如果技能呼叫 Foundry 資源時回傳 401 或 403,請確認:
- 技能組中的
cognitiveServices方塊指向正確的鑄造資源。 - 搜尋服務所使用的身份在 Foundry 資源中擁有所需的角色。 如需了解受控識別設定,請參閱 在 Azure AI 搜尋服務 中將計費資源附加至技能集。
text_sections 是空的
如果索引文件沒有區塊,請確認:
- 支援此檔案格式。 關於清單,請參見 支援的檔案格式。
- Foundry 資源位於受支援的區域。
- 受密碼保護的 PDF 在索引前會先解鎖。
圖片描述(預覽)缺少
如果區塊中沒有包含內嵌圖片描述,請確認:
- 技能集中已同時設定
modelName和modelDeployment。 -
modelName中的對話補全模型部署在技能集所參考的相同 Foundry 資源中。 - 此部署的 TPM 或 RPM 配額足以因應您的文件量。
索引子在大型文件上逾時
Content Understanding 會對每份文件的處理施加逾時限制。 如果大型PDF失敗:
- 在索引前,先將原始文件拆分成較小的檔案。
- 縮減
batchSize為1每份文件獨立處理。
關於Azure內容理解技能的完整資料限制,請參見資料限制。