文件排版技能

Note

Azure AI 搜尋服務 可透過 Azure 入口網站、REST API 及 Azure SDK 取得。 它同時也是 Foundry IQ 的基礎,這是一個管理式知識層,能將企業內容轉化為可重複使用、權限感知的知識庫,供 Microsoft Foundry 入口網站中的代理使用。

Document Layout 技能利用 Foundry Tools 中 Azure 文件智慧中的 layout 模型來分析文件、偵測其結構與特性,並以 Markdown 或文字格式產生語法表示。 此技能支援文字與影像擷取,後者包含位置元資料,以保留文件中的影像位置。 影像與相關內容的接近性在檢索增強生成(RAG)及 多模態搜尋 情境中非常有利。

對於每位索引員每日超過20份文件的交易,這項技能要求你將可計費的Microsoft Foundry資源><附加到你的技能組中。 內建技能的執行依現有 的 Foundry Tools 標準價格收費。

本文為文件版面技能的參考文件。 使用資訊請參考 「如何依文件版面分塊與向量化」。

Tip

這項技能常用於有結構和圖片的內容,例如 PDF。 多模態教學示範了兩種不同的資料分塊策略來進行影像語言化。

局限性

此技能有以下限制:

  • 這項技能不適合在 Azure 文件智慧版面模型中處理需要超過五分鐘處理的大型文件。 技能會逾時,但如果 Foundry 資源在技能組上綁定以計費,收費仍然會生效。 確保文件優化以維持在處理限制內,以避免不必要的成本。

  • 由於此技能呼叫Azure文件智慧佈局模型,所有文件類型服務行為,不同檔案類型的行為都會套用於其輸出。 例如,Word(DOCX)與 PDF 檔案可能因影像處理方式不同而產生不同結果。 若需要在 DOCX 與 PDF 間保持一致的影像行為,請考慮將文件轉換為 PDF 或檢視 多模態搜尋文件 以尋找替代方案。

支援的區域

文件排版技能呼叫 Azure文件智慧 REST API 的 v4.0(2024-11-30)。

支援的區域會依模式及技能與 Azure 文件智慧佈局模型的連結而異。 目前,實作版面佈局模型不支援 21 Vianet 區域。

方法 需求
匯入資料 精靈 在以下地區之一建立Azure AI 搜尋服務服務及Azure AI 多服務帳號:美國東部、西歐2區或美國中北部。
程式化,使用 Microsoft Foundry 資源金鑰來計費 在同一區域建立 Azure AI 搜尋服務 服務和 Microsoft Foundry 資源。 該區域必須同時支援Azure AI 搜尋服務及Azure文件智慧。
程式化,使用 Microsoft Entra ID 認證來進行帳單 沒有同區域限制。 在任何有每個服務可用的地區建立Azure AI 搜尋服務服務並Microsoft Foundry 資源。

支援的檔案格式

此技能可識別以下檔案格式:

  • .PDF
  • .JPEG
  • .JPG
  • .PNG
  • .BMP
  • .多倫多國際影展
  • .DOCX
  • .XLSX
  • .PPTX
  • .HTML

支援的語言

關於印刷文字,請參見 Azure 文件智慧版面模型支援語言。

@odata.type

Microsoft.Skills.Util.DocumentIntelligenceLayoutSkill

資料限制

  • 對於 PDF 和 TIFF,最多可處理 2,000 頁(免費訂閱時僅處理前兩頁)。
  • 即使文件分析的檔案大小為 Azure文件智慧付費層級 500 MB Azure文件智慧免費層級 為 4 MB,索引仍受搜尋服務層級索引器限制限制。
  • 影像尺寸必須介於 50 像素 x 50 像素或 10,000 像素 x 10,000 像素之間。
  • 如果你的 PDF 有密碼鎖,啟動索引器前先解除鎖定。

技能參數

參數依大小寫區分。

參數名稱 允許的值 Description
outputMode oneToMany 控制技能所產生輸出的基數。
markdownHeaderDepth h1, h2, , h3h4, h5, ( h6 預設) 只有當 outputFormat 設定為 markdown時才適用。 此參數描述應考慮的最深巢狀層級。 例如,若 ,則markdownHeaderDepth任何較深h3的截面,如 ,皆被捲入 h4。h3
outputFormat markdown (預設), text 控制技能產生的輸出格式。
extractionOptions ["images"], ["images", "locationMetadata"], ["locationMetadata"] 找出從文件中提取的任何額外內容。 定義一個對應輸出內容的列舉陣列。 例如,若 extractionOptions 為 ["images", "locationMetadata"],輸出包含圖片與位置元資料,提供與內容擷取地點相關的頁面位置資訊,例如頁碼或區段。 此參數適用於兩種輸出格式。
chunkingProperties 請參見下表。 只有當 outputFormat 設定為 text時才適用。 這些選項包含如何在重新計算其他元資料的同時,將文字內容分割成區塊。
chunkingProperties 參數 允許的值 Description
unit characters 控制區塊單元的基數。 區塊長度是以字元為單位,而非單字或標記。
maximumLength 一個介於300到50000之間的整數。 以 String.Length 為單位,字元區塊長度最大。
overlapLength 整數小於 的一半。maximumLength 兩個文字區塊之間的重疊長度。

技能輸入

輸入名稱 Description
file_data 內容應該從中擷取的檔案。

「file_data」輸入必須是定義為:

{
  "$type": "file",
  "data": "BASE64 encoded string of the file"
}

或者,也可以定義為:

{
  "$type": "file",
  "url": "URL to download file",
  "sasToken": "OPTIONAL: SAS token for authentication if the URL provided is for a file in blob storage"
}

檔案參考物件可透過以下方式之一產生:

  • 將你的索引器定義中的參數設 allowSkillsetToReadFileData 為 true。 這個設定會建立一個 /document/file_data 路徑,作為一個物件,代表從你的 blob 資料來源下載的原始檔案資料。 這個參數只適用於 Azure Blob 儲存中的檔案。

    allowSkillsetToReadFileData 讓下載的檔案資料可供技能使用。 它不會增加 blob 索引器的限制 ,也不會增加 資料限制中描述的文件智慧限制。

  • 擁有自訂技能回傳一個 JSON 物件定義,該定義提供 $type、 data、 和 urlsastoken。 $type參數必須設為 file,且data必須是檔案內容的 64 編碼位元組陣列。 url參數必須是有效的 URL,且有權限在該位置下載檔案。

技能產出

輸出名稱 Description
markdown_document 只有當 outputFormat 設定為 markdown時才適用。 一組「區段」物件,代表 Markdown 文件中每個獨立區段。
text_sections 只有當 outputFormat 設定為 text時才適用。 一組文字區塊物件,代表頁面範圍內的文字(考慮更多區塊配置), 包含 任何區塊標題。 文本區塊物件(如適用)包含 locationMetadata 。
normalized_images 僅當 outputFormat 設定為 text 且 extractionOptions 包含 images時才適用。 從文件中擷取的圖片集合,包括 locationMetadata 如適用的圖片。

標記降級輸出模式的範例定義

{
  "skills": [
    {
      "description": "Analyze a document",
      "@odata.type": "#Microsoft.Skills.Util.DocumentIntelligenceLayoutSkill",
      "context": "/document",
      "outputMode": "oneToMany", 
      "markdownHeaderDepth": "h3", 
      "inputs": [
        {
          "name": "file_data",
          "source": "/document/file_data"
        }
      ],
      "outputs": [
        {
          "name": "markdown_document", 
          "targetName": "markdown_document" 
        }
      ]
    }
  ]
}

標記降級輸出模式的取樣輸出

{
  "markdown_document": [
    { 
      "content": "Hi this is Jim \r\nHi this is Joe", 
      "sections": { 
        "h1": "Foo", 
        "h2": "Bar", 
        "h3": "" 
      },
      "ordinal_position": 0
    }, 
    { 
      "content": "Hi this is Lance",
      "sections": { 
         "h1": "Foo", 
         "h2": "Bar", 
         "h3": "Boo" 
      },
      "ordinal_position": 1,
    } 
  ] 
}

的 markdownHeaderDepth 值控制「sections」字典中的鍵數。 在範例技能定義中,由於 是 markdownHeaderDepth 「h3」,「sections」字典中有三個鍵:h1、h2、h3。

文字輸出模式及影像與元資料擷取的範例

此範例示範如何將文字內容輸出為固定大小的區塊,並從文件中擷取圖片及位置元資料。

文字輸出模式及影像與元資料擷取的範例定義

{
  "skills": [
    {
      "description": "Analyze a document",
      "@odata.type": "#Microsoft.Skills.Util.DocumentIntelligenceLayoutSkill",
      "context": "/document",
      "outputMode": "oneToMany",
      "outputFormat": "text",
      "extractionOptions": ["images", "locationMetadata"],
      "chunkingProperties": {     
          "unit": "characters",
          "maximumLength": 2000, 
          "overlapLength": 200
      },
      "inputs": [
        {
          "name": "file_data",
          "source": "/document/file_data"
        }
      ],
      "outputs": [
        { 
          "name": "text_sections", 
          "targetName": "text_sections" 
        }, 
        { 
          "name": "normalized_images", 
          "targetName": "normalized_images" 
        } 
      ]
    }
  ]
}

文字輸出模式及影像與元資料擷取的取樣輸出

{
  "text_sections": [
      {
        "id": "1_7e6ef1f0-d2c0-479c-b11c-5d3c0fc88f56",
        "content": "the effects of analyzers using Analyze Text (REST). For more information about analyzers, see Analyzers for text processing.During indexing, an indexer only checks field names and types. There's no validation step that ensures incoming content is correct for the corresponding search field in the index.Create an indexerWhen you're ready to create an indexer on a remote search service, you need a search client. A search client can be the Azure portal, a REST client, or code that instantiates an indexer client. We recommend the Azure portal or REST APIs for early development and proof-of-concept testing.Azure portal1. Sign in to the Azure portal 2, then find your search service.2. On the search service Overview page, choose from two options:· Import data wizard: The wizard is unique in that it creates all of the required elements. Other approaches require a predefined data source and index.All services > Azure Al services | Al Search >demo-search-svc Search serviceSearchAdd indexImport dataImport and vectorize dataOverviewActivity logEssentialsAccess control (IAM)Get startedPropertiesUsageMonitoring· Add indexer: A visual editor for specifying an indexer definition.",
        "locationMetadata": {
          "pageNumber": 1,
          "ordinalPosition": 0,
          "boundingPolygons": "[[{\"x\":1.5548,\"y\":0.4036},{\"x\":6.9691,\"y\":0.4033},{\"x\":6.9691,\"y\":0.8577},{\"x\":1.5548,\"y\":0.8581}],[{\"x\":1.181,\"y\":1.0627},{\"x\":7.1393,\"y\":1.0626},{\"x\":7.1393,\"y\":1.7363},{\"x\":1.181,\"y\":1.7365}],[{\"x\":1.1923,\"y\":2.1466},{\"x\":3.4585,\"y\":2.1496},{\"x\":3.4582,\"y\":2.4251},{\"x\":1.1919,\"y\":2.4221}],[{\"x\":1.1813,\"y\":2.6518},{\"x\":7.2464,\"y\":2.6375},{\"x\":7.2486,\"y\":3.5913},{\"x\":1.1835,\"y\":3.6056}],[{\"x\":1.3349,\"y\":3.9489},{\"x\":2.1237,\"y\":3.9508},{\"x\":2.1233,\"y\":4.1128},{\"x\":1.3346,\"y\":4.111}],[{\"x\":1.5705,\"y\":4.5322},{\"x\":5.801,\"y\":4.5326},{\"x\":5.801,\"y\":4.7311},{\"x\":1.5704,\"y\":4.7307}]]"
        },
        "sections": []
      },
      {
        "id": "2_25134f52-04c3-415a-ab3d-80729bd58e67",
        "content": "All services > Azure Al services | Al Search >demo-search-svc | Indexers Search serviceSearch0«Add indexerRefreshDelete:selected: TagsFilter by name ...:selected: Diagnose and solve problemsSearch managementStatusNameIndexesIndexers*Data sourcesRun the indexerBy default, an indexer runs immediately when you create it on the search service. You can override this behavior by setting disabled to true in the indexer definition. Indexer execution is the moment of truth where you find out if there are problems with connections, field mappings, or skillset construction.There are several ways to run an indexer:· Run on indexer creation or update (default).. Run on demand when there are no changes to the definition, or precede with reset for full indexing. For more information, see Run or reset indexers.· Schedule indexer processing to invoke execution at regular intervals.Scheduled execution is usually implemented when you have a need for incremental indexing so that you can pick up the latest changes. As such, scheduling has a dependency on change detection.Indexers are one of the few subsystems that make overt outbound calls to other Azure resources. In terms of Azure roles, indexers don't have separate identities; a connection from the search engine to another Azure resource is made using the system or user- assigned managed identity of a search service. If the indexer connects to an Azure resource on a virtual network, you should create a shared private link for that connection. For more information about secure connections, see Security in Azure Al Search.Check results",
        "locationMetadata": {
          "pageNumber": 2,
          "ordinalPosition": 1,
          "boundingPolygons": "[[{\"x\":2.2041,\"y\":0.4109},{\"x\":4.3967,\"y\":0.4131},{\"x\":4.3966,\"y\":0.5505},{\"x\":2.204,\"y\":0.5482}],[{\"x\":2.5042,\"y\":0.6422},{\"x\":4.8539,\"y\":0.6506},{\"x\":4.8527,\"y\":0.993},{\"x\":2.5029,\"y\":0.9845}],[{\"x\":2.3705,\"y\":1.1496},{\"x\":2.6859,\"y\":1.15},{\"x\":2.6858,\"y\":1.2612},{\"x\":2.3704,\"y\":1.2608}],[{\"x\":3.7418,\"y\":1.1709},{\"x\":3.8082,\"y\":1.171},{\"x\":3.8081,\"y\":1.2508},{\"x\":3.7417,\"y\":1.2507}],[{\"x\":3.9692,\"y\":1.1445},{\"x\":4.0541,\"y\":1.1445},{\"x\":4.0542,\"y\":1.2621},{\"x\":3.9692,\"y\":1.2622}],[{\"x\":4.5326,\"y\":1.2263},{\"x\":5.1065,\"y\":1.229},{\"x\":5.106,\"y\":1.346},{\"x\":4.5321,\"y\":1.3433}],[{\"x\":5.5508,\"y\":1.2267},{\"x\":5.8992,\"y\":1.2268},{\"x\":5.8991,\"y\":1.3408},{\"x\":5.5508,\"y\":1.3408}]]"
        },
        "sections": []
       }
    ],
    "normalized_images": [ 
        { 
            "id": "1_550e8400-e29b-41d4-a716-446655440000", 
            "data": "SGVsbG8sIFdvcmxkIQ==", 
            "imagePath": "aHR0cHM6Ly9henNyb2xsaW5nLmJsb2IuY29yZS53aW5kb3dzLm5ldC9tdWx0aW1vZGFsaXR5L0NyZWF0ZUluZGV4ZXJwNnA3LnBkZg2/normalized_images_0.jpg",  
            "locationMetadata": {
              "pageNumber": 1,
              "ordinalPosition": 0,
              "boundingPolygons": "[[{\"x\":2.0834,\"y\":6.2245},{\"x\":7.1818,\"y\":6.2244},{\"x\":7.1816,\"y\":7.9375},{\"x\":2.0831,\"y\":7.9377}]]"
            }
        },
        { 
            "id": "2_123e4567-e89b-12d3-a456-426614174000", 
            "data": "U29tZSBtb3JlIGV4YW1wbGUgdGV4dA==", 
            "imagePath": "aHR0cHM6Ly9henNyb2xsaW5nLmJsb2IuY29yZS53aW5kb3dzLm5ldC9tdWx0aW1vZGFsaXR5L0NyZWF0ZUluZGV4ZXJwNnA3LnBkZg2/normalized_images_1.jpg",  
            "locationMetadata": {
              "pageNumber": 2,
              "ordinalPosition": 1,
              "boundingPolygons": "[[{\"x\":2.0784,\"y\":0.3734},{\"x\":7.1837,\"y\":0.3729},{\"x\":7.183,\"y\":2.8611},{\"x\":2.0775,\"y\":2.8615}]]"
            } 
        }
    ] 
}

請注意,上述範例輸出中顯示 “sections” 為空白。 要填滿這些區塊,你需要新增一個技能,並設定為 outputFormat set 來 markdown確保區塊被正確填滿。

此技能使用 Azure 文件智慧來計算位置元資料。 有關頁面與邊界多邊形座標的定義,請參閱 Azure 文件智慧版面模型。

代表 imagePath 儲存影像的相對路徑。 如果技能集中已設定知識庫檔案投影,該路徑會與知識庫中影像的相對路徑相符。