문서 레이아웃 기술

메모

Azure AI 검색 Azure 포털, REST API 및 Azure SDK 통해 사용할 수 있습니다. 또한 엔터프라이즈 콘텐츠를 Microsoft Foundry 포털의 에이전트에 대해 재사용 가능한 사용 권한 인식 기술 자료로 변환하는 관리되는 기술 계층인 Foundry IQ를 뒷받침합니다.

문서 레이아웃 스킬은 Foundry Tools의 문서 인텔리전Azure스의 레이아웃 모델을 사용하여 문서를 분석하고, 그 구조와 특성을 감지하며, 마크다운 또는 텍스트 형식으로 구문 표현을 생성합니다. 이 기술은 텍스트와 이미지 추출을 지원하며, 이미지 추출에는 문서 내 이미지 위치를 보존하는 위치 메타데이터가 포함됩니다. 이미지 근접성은 검색 증강 생성(RAG) 및 다중 모달 검색 시나리오에서 유용합니다.

인덱서당 하루에 20개를 초과하는 거래가 발생할 경우, 이 기술은 청구 가능한 Microsoft Foundry 리소스를 귀하의 스킬셋에 부착해야 합니다. 내장 기술 실행은 기존 파운드리 공구 표준 가격으로 부과됩니다.

이 문서는 문서 레이아웃 기술에 대한 참고 문서입니다. 사용 정보는 문서 레이아웃에 따라 청크와 벡터화 방법을 참조하세요.

Tip

이 기술은 PDF처럼 구조와 이미지가 있는 콘텐츠에 흔히 사용됩니다. 멀티모달 튜토리얼은 두 가지 다른 데이터 청킹 전략을 사용한 이미지 언어화를 시연합니다.

제한점

이 스킬에는 다음과 같은 한계가 있습니다:

  • 이 기술은 Azure 문서 인텔리전스 레이아웃 모델에서 5분 이상 처리가 필요한 대규모 문서에는 적합하지 않습니다. 스킬은 시간 초과되지만, 청구 목적으로 스킬셋에 연결된 파운드리 자원에는 여전히 요금이 적용됩니다. 문서 처리가 한도 내에서 최적화되어 불필요한 비용을 피하도록 하세요.

  • 이 스킬이 Azure 문서 인텔리전스 레이아웃 모델을 호출하기 때문에, 서로 다른 문서 유형에 대한 모든 문서화된 service 동작 데이터가 출력에 적용됩니다. 예를 들어, Word(DOCX)와 PDF 파일은 이미지 처리 방식의 차이로 인해 서로 다른 결과를 낼 수 있습니다. DOCX와 PDF 전반에 걸쳐 일관된 이미지 동작이 필요하다면, 문서를 PDF로 변환하거나 멀티모달 검색 문서를 검토하여 대체 방법을 찾아보는 것을 고려하세요.

지원되는 지역

문서 레이아웃 스킬은 Azure 문서 인텔리전스 REST API의 v4.0 (2024-11-30)을 호출합니다.

지원되는 지역은 모달리티와 스킬이 Azure 문서 인텔리전스 레이아웃 모델과 어떻게 연결되는지에 따라 다릅니다. 현재 구현된 레이아웃 모델은 21Vianet 영역을 지원하지 않습니다.

접근법 요구 사항
데이터 가져오기 마법사 다음 지역 중 하나(동부 미국, 서유럽 2, 또는 북중부 미국)에서 Azure AI 검색 서비스와 Azure AI 다중 서비스 계정을 생성하세요.
Microsoft Foundry 리소스 키를 사용하여 청구하는 프로그래매틱 방식입니다 같은 지역 내에서 Azure AI 검색 서비스와 Microsoft Foundry 리소스를 생성하세요. 이 지역은 Azure AI 검색와 문서 인텔리전스 Azure 모두를 지원해야 합니다.
프로그래밍 방식, 청구에 Microsoft Entra ID 인증 사용 동일 지역 요구사항은 없습니다. 각 서비스가 제공되는 지역에 Azure AI 검색 서비스와 파운드리 자원을 Microsoft하세요.

지원되는 파일 형식

이 기술은 다음 파일 형식을 인식합니다:

  • .Pdf
  • . Jpeg
  • .Jpg
  • .Png
  • .Bmp
  • . Tiff
  • .Docx
  • . Xlsx
  • .Pptx
  • .Html

지원되는 언어

인쇄된 텍스트에 대해서는 Azure 문서 인텔리전스 레이아웃 모델 지원 언어를 참조하세요.

@odata.type

Microsoft.Skills.Util.DocumentIntelligenceLayoutSkill

데이터 제한

  • PDF 및 TIFF의 경우 최대 2,000페이지를 처리할 수 있습니다(무료 계층 구독에서는 처음 두 페이지만 처리됨).
  • 문서 분석 파일 크기가 Azure문서 인텔리전스 유료 계층 500MB이고 Azure문서 인텔리전스 무료(F0) 계층 4MB라 하더라도, 색인은 검색 서비스 계층의 인덱서 제한에 적용됩니다.
  • 이미지 크기는 50 픽셀 x 50 픽셀 또는 10,000 픽셀 x 10,000 픽셀 사이여야 합니다.
  • PDF가 비밀번호로 잠겨 있다면, 인덱서를 실행하기 전에 잠금장치를 해제하세요.

기술 매개 변수

매개변수는 대소문자에 구분됩니다.

매개 변수 이름 허용되는 값 Description
outputMode oneToMany 스킬이 생성하는 출력의 기수를 제어합니다.
markdownHeaderDepth h1, h2, , h3h4, h5, ( h6 기본값) 가 로 설정outputFormat된 경우에만 markdown 적용됩니다. 이 매개변수는 고려해야 할 가장 깊은 중첩 수준을 설명합니다. 예를 들어, 가 이라markdownHeaderDepth면h3, 더 깊은 구간들, 예를 h4들어 는 는 로 말h3려 들어가게 됩니다.
outputFormat markdown (기본값), text 스킬이 생성하는 출력 형식을 제어합니다.
extractionOptions ["images"], , ["images", "locationMetadata"]["locationMetadata"] 문서에서 추출한 추가 내용을 식별하세요. 출력에 포함될 내용에 대응하는 열거 배열을 정의하세요. 예를 들어, 가 이extractionOptions라면["images", "locationMetadata"], 출력에는 페이지 번호나 섹션과 같은 페이지 위치 정보를 제공하는 이미지와 위치 메타데이터가 포함되어 있습니다. 이 매개변수는 두 출력 포맷 모두에 적용됩니다.
chunkingProperties 다음 표를 참고하세요. 가 로 설정outputFormat된 경우에만 text 적용됩니다. 텍스트 콘텐츠를 청크하는 방법과 다른 메타데이터를 재계산하는 방법을 캡슐화하는 옵션들.
chunkingProperties 매개변수 허용되는 값 Description
unit characters 청크 단위의 기수를 제어합니다. 청크 길이는 단어나 토큰이 아닌 문자 단위로 측정됩니다.
maximumLength 300에서 50000 사이의 정수입니다. String.Length로 측정된 문자의 최대 청크 길이입니다.
overlapLength 의 절반 maximumLength보다 작은 정수입니다. 두 텍스트 청크 사이에 제공되는 중복 길이.

기술 입력

입력 이름 Description
file_data 그 콘텐츠가 추출되어야 할 파일입니다.

"file_data" 입력은 다음과 같이 정의된 객체여야 합니다:

{
  "$type": "file",
  "data": "BASE64 encoded string of the file"
}

또는 다음과 같이 정의할 수 있습니다:

{
  "$type": "file",
  "url": "URL to download file",
  "sasToken": "OPTIONAL: SAS token for authentication if the URL provided is for a file in blob storage"
}

파일 참조 객체는 다음과 같은 방법 중 하나로 생성할 수 있습니다:

  • 인덱서 정의에서 매개변수를 allowSkillsetToReadFileData true로 설정하는 것입니다. 이 설정은 블롭 데이터 소스에서 다운로드한 원본 파일 데이터를 나타내는 객체 경로 /document/file_data 가 생성됩니다. 이 매개변수는 Azure Blob 저장소 내 파일에만 적용됩니다.

    allowSkillsetToReadFileData 는 다운로드한 파일 데이터를 기술에 사용할 수 있도록 합니다. 데이터 제한에 설명된 Blob 인덱서 제한 또는 문서 인텔리전스 제한은 증가하지 않습니다.

  • 커스텀 스킬이 JSON 객체 정의를 반환하여 , $typedata 를 url제공합니다sastoken. 매개변수는 $type 로 설정 file되어야 하며, data 파일 내용의 기본 64바이트 배열이어야 합니다. 매개변수는 url 해당 위치에서 파일을 다운로드할 수 있는 유효한 URL이어야 합니다.

기술 성과

출력 이름 Description
markdown_document 가 로 설정outputFormat된 경우에만 markdown 적용됩니다. Markdown 문서 내 각 개별 섹션을 나타내는 "섹션" 객체들의 모음입니다.
text_sections 가 로 설정outputFormat된 경우에만 text 적용됩니다. 텍스트 청크 객체들의 집합으로, 페이지 경계 내의 텍스트를 나타내며(추가로 청킹이 설정된 것을 제외하고), 섹션 헤더 자체를 포함 합니다. 텍스트 청크 객체는 해당되는 경우 포함합니다 locationMetadata .
normalized_images 가 로 outputFormat 설정되어 있고 text 를 포함할 extractionOptions때 images 만 적용됩니다. 문서에서 추출한 이미지 모음으로, locationMetadata 해당되는 경우 포함된다.

마크다운 출력 모드의 샘플 정의

{
  "skills": [
    {
      "description": "Analyze a document",
      "@odata.type": "#Microsoft.Skills.Util.DocumentIntelligenceLayoutSkill",
      "context": "/document",
      "outputMode": "oneToMany", 
      "markdownHeaderDepth": "h3", 
      "inputs": [
        {
          "name": "file_data",
          "source": "/document/file_data"
        }
      ],
      "outputs": [
        {
          "name": "markdown_document", 
          "targetName": "markdown_document" 
        }
      ]
    }
  ]
}

마크다운 출력 모드의 샘플 출력

{
  "markdown_document": [
    { 
      "content": "Hi this is Jim \r\nHi this is Joe", 
      "sections": { 
        "h1": "Foo", 
        "h2": "Bar", 
        "h3": "" 
      },
      "ordinal_position": 0
    }, 
    { 
      "content": "Hi this is Lance",
      "sections": { 
         "h1": "Foo", 
         "h2": "Bar", 
         "h3": "Boo" 
      },
      "ordinal_position": 1,
    } 
  ] 
}

의 markdownHeaderDepth 값은 "섹션" 사전의 키 수를 제어합니다. 예시 스킬 정의에서 ' markdownHeaderDepth h3'이기 때문에, '섹션' 사전에는 세 가지 키가 있습니다: h1, h2, h3.

텍스트 출력 모드와 이미지 및 메타데이터 추출의 예

이 예시는 텍스트 콘텐츠를 고정 크기의 청크로 출력하고 문서에서 이미지와 위치 메타데이터를 추출하는 방법을 보여줍니다.

텍스트 출력 모드 및 이미지 및 메타데이터 추출을 위한 샘플 정의

{
  "skills": [
    {
      "description": "Analyze a document",
      "@odata.type": "#Microsoft.Skills.Util.DocumentIntelligenceLayoutSkill",
      "context": "/document",
      "outputMode": "oneToMany",
      "outputFormat": "text",
      "extractionOptions": ["images", "locationMetadata"],
      "chunkingProperties": {     
          "unit": "characters",
          "maximumLength": 2000, 
          "overlapLength": 200
      },
      "inputs": [
        {
          "name": "file_data",
          "source": "/document/file_data"
        }
      ],
      "outputs": [
        { 
          "name": "text_sections", 
          "targetName": "text_sections" 
        }, 
        { 
          "name": "normalized_images", 
          "targetName": "normalized_images" 
        } 
      ]
    }
  ]
}

텍스트 출력 모드와 이미지 및 메타데이터 추출을 위한 샘플 출력

{
  "text_sections": [
      {
        "id": "1_7e6ef1f0-d2c0-479c-b11c-5d3c0fc88f56",
        "content": "the effects of analyzers using Analyze Text (REST). For more information about analyzers, see Analyzers for text processing.During indexing, an indexer only checks field names and types. There's no validation step that ensures incoming content is correct for the corresponding search field in the index.Create an indexerWhen you're ready to create an indexer on a remote search service, you need a search client. A search client can be the Azure portal, a REST client, or code that instantiates an indexer client. We recommend the Azure portal or REST APIs for early development and proof-of-concept testing.Azure portal1. Sign in to the Azure portal 2, then find your search service.2. On the search service Overview page, choose from two options:· Import data wizard: The wizard is unique in that it creates all of the required elements. Other approaches require a predefined data source and index.All services > Azure Al services | Al Search >demo-search-svc Search serviceSearchAdd indexImport dataImport and vectorize dataOverviewActivity logEssentialsAccess control (IAM)Get startedPropertiesUsageMonitoring· Add indexer: A visual editor for specifying an indexer definition.",
        "locationMetadata": {
          "pageNumber": 1,
          "ordinalPosition": 0,
          "boundingPolygons": "[[{\"x\":1.5548,\"y\":0.4036},{\"x\":6.9691,\"y\":0.4033},{\"x\":6.9691,\"y\":0.8577},{\"x\":1.5548,\"y\":0.8581}],[{\"x\":1.181,\"y\":1.0627},{\"x\":7.1393,\"y\":1.0626},{\"x\":7.1393,\"y\":1.7363},{\"x\":1.181,\"y\":1.7365}],[{\"x\":1.1923,\"y\":2.1466},{\"x\":3.4585,\"y\":2.1496},{\"x\":3.4582,\"y\":2.4251},{\"x\":1.1919,\"y\":2.4221}],[{\"x\":1.1813,\"y\":2.6518},{\"x\":7.2464,\"y\":2.6375},{\"x\":7.2486,\"y\":3.5913},{\"x\":1.1835,\"y\":3.6056}],[{\"x\":1.3349,\"y\":3.9489},{\"x\":2.1237,\"y\":3.9508},{\"x\":2.1233,\"y\":4.1128},{\"x\":1.3346,\"y\":4.111}],[{\"x\":1.5705,\"y\":4.5322},{\"x\":5.801,\"y\":4.5326},{\"x\":5.801,\"y\":4.7311},{\"x\":1.5704,\"y\":4.7307}]]"
        },
        "sections": []
      },
      {
        "id": "2_25134f52-04c3-415a-ab3d-80729bd58e67",
        "content": "All services > Azure Al services | Al Search >demo-search-svc | Indexers Search serviceSearch0«Add indexerRefreshDelete:selected: TagsFilter by name ...:selected: Diagnose and solve problemsSearch managementStatusNameIndexesIndexers*Data sourcesRun the indexerBy default, an indexer runs immediately when you create it on the search service. You can override this behavior by setting disabled to true in the indexer definition. Indexer execution is the moment of truth where you find out if there are problems with connections, field mappings, or skillset construction.There are several ways to run an indexer:· Run on indexer creation or update (default).. Run on demand when there are no changes to the definition, or precede with reset for full indexing. For more information, see Run or reset indexers.· Schedule indexer processing to invoke execution at regular intervals.Scheduled execution is usually implemented when you have a need for incremental indexing so that you can pick up the latest changes. As such, scheduling has a dependency on change detection.Indexers are one of the few subsystems that make overt outbound calls to other Azure resources. In terms of Azure roles, indexers don't have separate identities; a connection from the search engine to another Azure resource is made using the system or user- assigned managed identity of a search service. If the indexer connects to an Azure resource on a virtual network, you should create a shared private link for that connection. For more information about secure connections, see Security in Azure Al Search.Check results",
        "locationMetadata": {
          "pageNumber": 2,
          "ordinalPosition": 1,
          "boundingPolygons": "[[{\"x\":2.2041,\"y\":0.4109},{\"x\":4.3967,\"y\":0.4131},{\"x\":4.3966,\"y\":0.5505},{\"x\":2.204,\"y\":0.5482}],[{\"x\":2.5042,\"y\":0.6422},{\"x\":4.8539,\"y\":0.6506},{\"x\":4.8527,\"y\":0.993},{\"x\":2.5029,\"y\":0.9845}],[{\"x\":2.3705,\"y\":1.1496},{\"x\":2.6859,\"y\":1.15},{\"x\":2.6858,\"y\":1.2612},{\"x\":2.3704,\"y\":1.2608}],[{\"x\":3.7418,\"y\":1.1709},{\"x\":3.8082,\"y\":1.171},{\"x\":3.8081,\"y\":1.2508},{\"x\":3.7417,\"y\":1.2507}],[{\"x\":3.9692,\"y\":1.1445},{\"x\":4.0541,\"y\":1.1445},{\"x\":4.0542,\"y\":1.2621},{\"x\":3.9692,\"y\":1.2622}],[{\"x\":4.5326,\"y\":1.2263},{\"x\":5.1065,\"y\":1.229},{\"x\":5.106,\"y\":1.346},{\"x\":4.5321,\"y\":1.3433}],[{\"x\":5.5508,\"y\":1.2267},{\"x\":5.8992,\"y\":1.2268},{\"x\":5.8991,\"y\":1.3408},{\"x\":5.5508,\"y\":1.3408}]]"
        },
        "sections": []
       }
    ],
    "normalized_images": [ 
        { 
            "id": "1_550e8400-e29b-41d4-a716-446655440000", 
            "data": "SGVsbG8sIFdvcmxkIQ==", 
            "imagePath": "aHR0cHM6Ly9henNyb2xsaW5nLmJsb2IuY29yZS53aW5kb3dzLm5ldC9tdWx0aW1vZGFsaXR5L0NyZWF0ZUluZGV4ZXJwNnA3LnBkZg2/normalized_images_0.jpg",  
            "locationMetadata": {
              "pageNumber": 1,
              "ordinalPosition": 0,
              "boundingPolygons": "[[{\"x\":2.0834,\"y\":6.2245},{\"x\":7.1818,\"y\":6.2244},{\"x\":7.1816,\"y\":7.9375},{\"x\":2.0831,\"y\":7.9377}]]"
            }
        },
        { 
            "id": "2_123e4567-e89b-12d3-a456-426614174000", 
            "data": "U29tZSBtb3JlIGV4YW1wbGUgdGV4dA==", 
            "imagePath": "aHR0cHM6Ly9henNyb2xsaW5nLmJsb2IuY29yZS53aW5kb3dzLm5ldC9tdWx0aW1vZGFsaXR5L0NyZWF0ZUluZGV4ZXJwNnA3LnBkZg2/normalized_images_1.jpg",  
            "locationMetadata": {
              "pageNumber": 2,
              "ordinalPosition": 1,
              "boundingPolygons": "[[{\"x\":2.0784,\"y\":0.3734},{\"x\":7.1837,\"y\":0.3729},{\"x\":7.183,\"y\":2.8611},{\"x\":2.0775,\"y\":2.8615}]]"
            } 
        }
    ] 
}

위 샘플 출력에서는 빈 “sections” 칸으로 표시되어 있습니다. 섹션을 채우려면 섹션이 제대로 채워지도록 설정 outputFormatmarkdown한 추가 스킬을 추가해야 합니다.

이 스킬은 Azure 문서 인텔리전스를 사용하여 위치 메타데이터를 계산합니다. 페이지와 경계 다각형 좌표가 어떻게 정의되는지에 대한 자세한 내용은 Azure 문서 인텔리전스 레이아웃 모델을 참조하십시오.

는 imagePath 저장된 이미지의 상대적 경로를 나타냅니다. 지식 저장소 파일 투영이 스킬셋에 설정되어 있다면, 이 경로는 지식 저장소에 저장된 이미지의 상대적 경로와 일치합니다.