你当前正在访问 Microsoft Azure Global Edition 技术文档网站。 如果需要访问由世纪互联运营的 Microsoft Azure 中国技术文档网站,请访问 https://docs.azure.cn。

使用Azure内容理解技能对内容进行区块和矢量化

注释

Azure AI 搜索可通过Azure门户、REST API 和Azure SDK获取。 它也是 Foundry IQ 的基础;Foundry IQ 是一个托管式知识层,可将企业内容转化为供 Microsoft Foundry 门户中的智能体使用的、可复用且具备权限感知能力的知识库。

Important

标记为“预览”的特性、功能或属性不受服务级别协议 (SLA) 保障,不建议用于生产工作负载,并且在正式发布之前可能会更改或受到限制。 Azure AI 搜索预览条款适用于所有预览功能,无论是独立功能还是正式版功能的一部分。

Important

这些特性和功能支持与其他Microsoft 服务和第三方服务的连接。 使用这些服务受其各自的条款的约束,可能会导致数据处理或存储超出Azure符合性边界,以及流入Azure符合性边界的数据。

您有责任管理您的数据是否会流出您组织的合规和地理边界之外及其任何相关影响,并确保已配置适当的权限、边界和审批。

你负责仔细查看和测试在特定用例上下文中生成的应用程序,并做出所有适当的决策和自定义。 这包括实施自己的负责任的 AI 缓解措施,例如元系统、内容筛选器或其他安全系统,并确保应用程序满足适当的质量、可靠性、安全性和可信度标准。 有关详细信息,请参阅 Azure AI 搜索 透明度说明。

本文介绍如何使用 Azure 内容理解技能:

  • 从文档中提取文本和图像
  • 生成遵循段落和节边界的语义上一致的区块(预览版)
  • 生成图表、示意图和其他内联图像的 AI 描述(预览版)
  • 将每个区块嵌入矢量搜索,并将其投影到Azure AI 搜索索引中

Azure内容理解技能返回每个文档的一个或多个区块。 每个区块都包含 Markdown 格式的内容、位置元数据(页码和边界多边形),以及对提取的图像的可选引用。 当您将 chunkingProperties.method 设置为 semantic 时,分块将按照段落和标题边界进行划分,而不是按固定字符长度划分。 当你设置 modelName 和 modelDeployment 时,该技能会调用 Azure OpenAI 聊天补全部署,以生成内嵌图像的描述。 然后,技能将这些说明合并到区块内容中。

本文使用 示例健康保险计划 PDF作为示例。 你可以对任何 受支持的数据源 运行同一管道,只要这些数据源提供的文件采用 Content Understanding 支持的格式。

先决条件

  • 位于任何受支持区域中的 Azure AI 搜索 服务。 对于此方案,搜索服务本身不受区域限制。

  • 位于Azure 内容理解技能支持的区域中的 Microsoft Foundry 资源。 图像描述和分块处理在 Foundry 资源所在的区域中进行。

  • Microsoft Foundry 资源附加到技能组进行计费。 Azure内容理解技能按 Azure 内容理解定价计费。

  • (可选)在同一 Foundry 资源中部署的 Azure OpenAI 聊天补全模型(如 gpt-4.1),用于生成图像描述。 仅当需要基于 AI 的图像说明时才是必需的。

  • 由 text-embedding-3-small使用、用于将分块向量化的嵌入模型(例如 )的 Azure OpenAI 部署。

  • 包含要编制索引的文件的 Azure Blob 存储 容器。 本文使用带有 allowSkillsetToReadFileData 索引器设置的 blob 数据源(该设置用于将文件内容传递给内容理解技能)。

Overview

本文构建了一个一对多的索引流水线。 每个源文档生成多个搜索文档(每个区块一个):

  1. 索引器从Azure Blob 存储读取每个文件,并通过 /document/file_data 将二进制内容传递给技能组。

  2. Azure内容理解技能使用语义分块(预览版)生成text_sections。 当设置了 modelName 和 modelDeployment 时,它还会为嵌入的图像生成 AI 生成的描述(预览),并将这些描述内联到每个区块的 Markdown 内容中。

  3. Azure OpenAI 嵌入技能每个区块运行一次,并为区块内容生成矢量。

  4. 索引投影会将每个区块写入为目标索引中的一份搜索文档,并将内容、页面元数据、图像引用和向量映射到各个字段。

  5. (可选)知识存储将normalized_images投影到 Azure Blob 存储,以便客户端应用可以通过 URL 检索提取的图像。

准备数据文件

Azure内容理解技能处理每个文档的二进制内容,因此源文件必须采用技能支持的格式。 有关当前列表,请参阅 内容理解服务限制。 常见的支持格式包括 PDF、DOCX、XLSX、PPTX 和许多图像格式。

将文件上传到受支持的数据源。 可以使用Azure门户、REST API 或Azure SDK创建数据源。

以下最小请求创建在本演练中使用的数据源。

POST {endpoint}/datasources?api-version=2026-08-01-preview

{
  "name": "my_blob_datasource",
  "type": "azureblob",
  "credentials": {
    "connectionString": "<your-blob-connection-string>"
  },
  "container": {
    "name": "my-container"
  }
}

为一对多索引编制创建索引

每个搜索文档对应于内容理解技能生成的一个区块。 索引需要:

  • 关键字段(chunk_id)。
  • 一个父字段,用于标识区块来自的源文档(parent_id)。
  • 存储区块内容、页面元数据和图像引用的字段。
  • 用于分块嵌入的向量字段。

以下索引定义与在下一部分中创建的技能集匹配。

{
  "name": "my_content_understanding_index",
  "fields": [
    {
      "name": "chunk_id",
      "type": "Edm.String",
      "key": true,
      "searchable": true,
      "filterable": false,
      "retrievable": true,
      "stored": true,
      "sortable": true,
      "facetable": false,
      "analyzer": "keyword"
    },
    {
      "name": "parent_id",
      "type": "Edm.String",
      "searchable": false,
      "filterable": true,
      "retrievable": true,
      "stored": true,
      "sortable": false,
      "facetable": false
    },
    {
      "name": "title",
      "type": "Edm.String",
      "searchable": true,
      "filterable": false,
      "retrievable": true,
      "stored": true,
      "sortable": false,
      "facetable": false
    },
    {
      "name": "chunk",
      "type": "Edm.String",
      "searchable": true,
      "filterable": false,
      "retrievable": true,
      "stored": true,
      "sortable": false,
      "facetable": false
    },
    {
      "name": "page_number_from",
      "type": "Edm.Int32",
      "searchable": false,
      "filterable": true,
      "retrievable": true,
      "stored": true,
      "sortable": true,
      "facetable": false
    },
    {
      "name": "page_number_to",
      "type": "Edm.Int32",
      "searchable": false,
      "filterable": true,
      "retrievable": true,
      "stored": true,
      "sortable": true,
      "facetable": false
    },
    {
      "name": "image_path",
      "type": "Edm.String",
      "searchable": false,
      "filterable": false,
      "retrievable": true,
      "stored": true,
      "sortable": false,
      "facetable": false
    },
    {
      "name": "text_vector",
      "type": "Collection(Edm.Single)",
      "searchable": true,
      "retrievable": true,
      "stored": false,
      "dimensions": 1536,
      "vectorSearchProfile": "profile"
    }
  ],
  "vectorSearch": {
    "profiles": [
      {
        "name": "profile",
        "algorithm": "algorithm"
      }
    ],
    "algorithms": [
      {
        "name": "algorithm",
        "kind": "hnsw"
      }
    ]
  }
}

定义语义分块(预览版)和矢量化的技能集

有了目标索引后,定义一个技能组,用于生成输入该索引的块、矢量和投影映射。

技能组有两种技能:

  • Azure 内容理解技能将每个文档分块。 将 chunkingProperties.method 设置为 semantic 会使该技能遵循段落和标题边界。 设置 modelName 和 modelDeployment 可启用 AI 生成的图像描述(预览),该技能会在矢量化之前将其内联到块内容中。 有关支持的聊天完成模型和其他参数详细信息的列表,请参阅 技能参数。

  • Azure OpenAI 嵌入技能为每个区块的内容生成矢量。

该技能集使用 indexProjections 将每个内容块映射到单独的搜索文档中。 有关详细信息,请参阅定义索引投影。

发送请求之前,请将 <subdomain> 替换为 Azure OpenAI 子域,<Azure OpenAI api key>替换为嵌入资源密钥,并将 <Foundry resource key> 替换为附加到技能组的 Foundry 资源的密钥。

POST {endpoint}/skillsets?api-version=2026-08-01-preview

{
  "name": "my_content_understanding_skillset",
  "description": "Semantic chunking, image descriptions, and vectorization with the Azure Content Understanding skill",
  "skills": [
    {
      "@odata.type": "#Microsoft.Skills.Util.ContentUnderstandingSkill",
      "name": "my_content_understanding_skill",
      "context": "/document",
      "modelName": "gpt-4.1",
      "modelDeployment": "my-gpt-4-1-deployment",
      "chunkingProperties": {
        "method": "semantic",
        "unit": "tokens",
        "maximumLength": 500
      },
      "extractionOptions": ["images", "locationMetadata"],
      "inputs": [
        {
          "name": "file_data",
          "source": "/document/file_data"
        }
      ],
      "outputs": [
        {
          "name": "text_sections",
          "targetName": "text_sections"
        },
        {
          "name": "normalized_images",
          "targetName": "normalized_images"
        }
      ]
    },
    {
      "@odata.type": "#Microsoft.Skills.Text.AzureOpenAIEmbeddingSkill",
      "name": "my_azure_openai_embedding_skill",
      "context": "/document/text_sections/*",
      "inputs": [
        {
          "name": "text",
          "source": "/document/text_sections/*/content"
        }
      ],
      "outputs": [
        {
          "name": "embedding",
          "targetName": "text_vector"
        }
      ],
      "resourceUri": "https://<subdomain>.openai.azure.com",
      "deploymentId": "text-embedding-3-small",
      "modelName": "text-embedding-3-small",
      "apiKey": "<Azure OpenAI api key>"
    }
  ],
  "cognitiveServices": {
    "@odata.type": "#Microsoft.Azure.Search.CognitiveServicesByKey",
    "key": "<Foundry resource key>"
  },
  "indexProjections": {
    "selectors": [
      {
        "targetIndexName": "my_content_understanding_index",
        "parentKeyFieldName": "parent_id",
        "sourceContext": "/document/text_sections/*",
        "mappings": [
          {
            "name": "chunk",
            "source": "/document/text_sections/*/content"
          },
          {
            "name": "text_vector",
            "source": "/document/text_sections/*/text_vector"
          },
          {
            "name": "page_number_from",
            "source": "/document/text_sections/*/locationMetadata/pageNumberFrom"
          },
          {
            "name": "page_number_to",
            "source": "/document/text_sections/*/locationMetadata/pageNumberTo"
          },
          {
            "name": "image_path",
            "source": "/document/text_sections/*/imagePath"
          },
          {
            "name": "title",
            "source": "/document/metadata_storage_name"
          }
        ]
      }
    ],
    "parameters": {
      "projectionMode": "skipIndexingParentDocuments"
    }
  }
}

有关内容理解技能的完整参数引用、支持的值和验证规则,请参阅Azure内容理解技能。

注释

本文使用 API 密钥来保持示例简洁。 我们建议在生产环境中使用托管标识:

有关端到端的概述,请参阅使用角色连接到 Azure AI 搜索。

配置和运行索引器

创建并运行一个索引器,使其从数据源读取数据、调用技能集,并将分块投射到索引中。 将 allowSkillsetToReadFileData 设置为 true,以便内容理解技能接收文件内容,并将 parsingMode 设置为 default。

在这种情况下,你不需要 outputFieldMappings。 技能组中的 indexProjections 块已将每个块映射到目标索引字段。

POST {endpoint}/indexers?api-version=2026-08-01-preview

{
  "name": "my_content_understanding_indexer",
  "dataSourceName": "my_blob_datasource",
  "targetIndexName": "my_content_understanding_index",
  "skillsetName": "my_content_understanding_skillset",
  "parameters": {
    "batchSize": 1,
    "configuration": {
      "dataToExtract": "contentAndMetadata",
      "parsingMode": "default",
      "allowSkillsetToReadFileData": true
    }
  },
  "fieldMappings": [],
  "outputFieldMappings": []
}

索引器运行时,内容理解技能使用语义分块(预览),(可选)生成基于 AI 的图像说明(预览),并将每个区块的一个搜索文档写入索引。

检查索引器状态

在查询之前,请确认索引器运行已完成:

GET {endpoint}/indexers/my_content_understanding_indexer/status?api-version=2026-08-01-preview

验证是否 lastResult.status 为 success. 如果其为transientFailure,且itemsProcessed高于0,则此次运行会被视为部分成功,您仍然可以查询已填充的区块。 有关详细信息,请参阅 监视器索引器状态。

验证结果

查询索引以验证区块是否包含预期内容,以及矢量搜索是否按预期工作。 使用 搜索资源管理器 或任何发送 HTTP 请求的工具。

以下请求会运行一个混合查询(在 chunk 上执行关键字搜索,并针对 text_vector 执行向量查询),以确认分块文本和嵌入向量这两项都已填充。

POST /indexes/my_content_understanding_index/docs/search?api-version=2026-08-01-preview
{
  "search": "copay for in-network providers",
  "count": true,
  "searchMode": "all",
  "vectorQueries": [
    {
      "kind": "text",
      "text": "copay for in-network providers",
      "fields": "text_vector"
    }
  ],
  "select": "chunk, title, page_number_from, page_number_to, image_path"
}

成功的响应与下面的示例类似(为简洁起见,以下内容有所删节):

{
  "@odata.count": 2,
  "value": [
    {
      "@search.score": 0.0317,
      "chunk": "## Cost sharing\n\nFor in-network providers, the copay is $20 per visit...\n\n![Chart: Copay comparison across plans](figures/3)",
      "title": "Northwind_Standard_Benefits_Details.pdf",
      "page_number_from": 4,
      "page_number_to": 4,
      "image_path": "figures/3"
    },
    {
      "@search.score": 0.0289,
      "chunk": "### Out-of-network providers\n\nWhen you visit a provider that isn't in the Northwind network, the copay is $40 per visit...",
      "title": "Northwind_Standard_Benefits_Details.pdf",
      "page_number_from": 5,
      "page_number_to": 6,
      "image_path": null
    }
  ]
}

响应包括:

  • chunk:每个区块的 Markdown 内容。 配置 modelName 和 modelDeployment 时,AI 生成的图像描述(预览版)会以内联方式显示在 Markdown 中。
  • page_number_from 和 page_number_to:生成区块的页面范围。
  • image_path:随该块提取出的图像路径;如果一个块跨多个图像,则为以分号分隔的路径列表。 确切的形状取决于是否配置了知识存储文件投影。 如果没有文件投影,路径是示例(figures/3)中显示的短格式。 使用文件投影时,路径是知识存储中图像的相对路径。 若要使这些映像可供客户端应用使用,请参阅 (可选)用于检索的项目映像。

(可选)用于检索的项目图片

索引中存储的 image_path 值是指向技能扩充树的指针,并不是可直接获取的 URL。 若要获取图像,请使用知识存储将normalized_images投射到 Azure Blob 存储中,然后为每个数据块生成对应的 Blob URL。

此步骤是可选的。 仅当客户端应用需要显示或下载提取的图像时,才添加它。

将以下属性添加到上一部分中的技能集有效负载。 技能集请求使用 api-version=2026-08-01-preview。

"knowledgeStore": {
  "storageConnectionString": "<your-azure-storage-connection-string>",
  "projections": [
    {
      "files": [
        {
          "storageContainer": "extracted-images",
          "source": "/document/normalized_images/*"
        }
      ],
      "tables": [],
      "objects": []
    }
  ]
}

索引器运行后,容器中的每个 extracted-images Blob 对应于一个 normalized_images 元素。 Blob URL 的格式为 https://<storage-account>.blob.core.windows.net/<container>/<imagePath>,其中 <imagePath> 与存储在 image_path 字段中的值匹配。

有关完整架构(包括其他投影类型(tables 和 objects)以及身份验证选项),请参阅 Azure AI 搜索中的知识存储“投影”。

清理资源

完成后,请删除索引器、技能组和索引,以避免继续产生内容理解和 Azure OpenAI 费用。 Azure Blob 存储中的源文件和 Foundry 资源本身会一直保留,直到删除它们。

故障排除

如果索引器失败或返回意外结果,请检查以下常见原因。

技能集验证失败并返回 400 错误

当参数组合冲突时,技能将返回错误 400 Skill validation failed 。 常见原因:

  • 设置了 modelName 但未设置 modelDeployment,反之亦然。 两者必须同时设置。
  • method 为 semantic (预览版)且 overlapLength 大于 0。 overlapLength设置为0或省略它。
  • method 和 unit 不是受支持的组合。 将 fixedSize 与 characters 一起使用,或将 semantic 与 tokens 一起使用。

针对 Foundry 资源的授权失败

如果在调用 Foundry 资源时技能返回 401 或 403,请验证:

text_sections 为空

如果索引文档没有区块,请验证:

  • 该文件格式受支持。 有关列表,请参阅 支持的文件格式。
  • Foundry 资源位于受支持的区域中。
  • 在编制索引之前,密码保护的 PDF 将解锁。

缺少图像说明(预览版)

如果区块不包含内联图像说明,请验证:

  • modelName 和 modelDeployment 均已在技能组中设置。
  • modelName 中的聊天补全模型部署在与技能组引用的相同 Foundry 资源中。
  • 该部署具备足够的 TPM 或 RPM 配额来支持文档量。

索引器处理大型文档时超时

内容理解强制执行每个文档的处理超时。 如果大型 PDF 失败:

  • 在编制索引之前,将源文档拆分为较小的文件。
  • 减少 batchSize 为 1 使每个文档独立处理。

有关Azure内容理解技能的完整数据限制,请参阅 Data 限制。