你当前正在访问 Microsoft Azure Global Edition 技术文档网站。 如果需要访问由世纪互联运营的 Microsoft Azure 中国技术文档网站,请访问 https://docs.azure.cn。
注释
Azure AI 搜索可通过Azure门户、REST API 和Azure SDK获取。 它还支持 Foundry IQ,该知识层将企业内容转换为 Microsoft Foundry 门户中代理的可重用权限感知知识库。
Important
服务级别协议未涵盖标记为(预览)的功能、功能或属性,不建议用于生产工作负荷,在正式发布之前可能会更改或约束它们。 Azure AI 搜索预览条款适用于所有预览功能,无论是独立功能还是正式版功能的一部分。
Azure内容理解技能绑定在计费Microsoft Foundry资源。 与其他Azure AI资源技能,如文档布局技能不同,Azure内容理解技能并未为每个索引员每天提供20份免费文档。 该技能的执行按Azure内容理解价格计费。
你可以用 Azure 内容理解技能来提取内容和分块。 你的技能组里不需要用文本分割技能。 该技能实现了与文档布局技能相同的界面,后者在 设置为 outputFormat 时使用 Foundry Tools 中的 text。 然而,Azure内容理解技能相较于文档布局技能有若干优势:
表格和图形以Markdown格式输出,便于大型语言模型(LLM)理解。 相比之下,文档布局技能将表格和图表输出为纯文本,可能导致信息丢失。
对于跨多个页面的表,Azure内容理解技能可以将跨页表识别和提取为单个单元。
Azure内容理解技能允许区块通过语义单元跨多个页面。
Azure内容理解技能比文档布局技能更具成本效益,因为内容理解API更便宜。
Azure内容理解可以为图像、图表、图表和嵌入式图表生成基于 AI 的说明。 嵌入的数字说明直接合并到为检索生成的 markdown 内容中。 这些说明是可搜索的,可以改进 RAG 地面和多模式检索质量。
Azure内容理解技能在 2026-04-01 REST API 中正式发布。 从此 2026-05-01-preview开始,技能可以选择为文档嵌入的图像、图表和图表生成基于 AI 的图像说明(预览)。 若要启用说明,必须在附加到技能集的 Foundry 资源中部署 Azure OpenAI 聊天完成模型。 此 API 版本还添加了语义分块(预览),这是一个布局感知选项,它尊重段落边界,并度量令牌中的区块长度。 这两项功能都需要选择加入。 省略新参数时,技能的行为与稳定 2026-04-01 API 版本中的行为相同。
局限性
Azure内容理解技能存在以下局限性:
该技能不适合需要在内容理解文档分析器中处理超过五分钟的大型文档。 技能会超时,但消耗仍然会应用到与该技能组相关的铸造厂资源。 确保文档优化以控制处理限制,以避免不必要的成本。
该技能调用Azure内容理解文档分析器,因此不同文档类型
服务>的所有文档<行为都适用于其输出。 例如,Word(DOCX)和PDF文件可能因图像处理方式不同而产生不同的结果。 如果需要在DOCX和PDF上保持一致的图像行为,可以考虑将文档转换为PDF或查看 多模态搜索文档 以寻找替代方法。
支持的区域
Azure内容理解技能调用内容理解 2025-11-01 REST API。 您的Foundry资源必须位于支持区域,具体描述见Azure内容理解区域及语言支持。
您的搜索服务可以位于任何支持的Azure AI 搜索区域。 当你的Foundry资源和Azure AI 搜索服务不在同一区域时,跨区域的网络延迟会影响索引器的性能。
支持的文件格式
Azure内容理解技能识别以下文件格式:
- .JPEG
- .JPG
- .PNG
- .BMP
- .HEIF
- .TIFF
- .DOCX
- .XLSX
- .PPTX
- .HTML
- .TXT
- .MD
- .RTF
- .EML
支持的语言
关于印刷文本,请参见 Azure内容理解区域与语言支持。
@odata.type
Microsoft.Skills.Util.ContentUnderstandingSkill
数据限制
即使分析文档的文件大小在200 MB限制内,如Azure内容理解服务配额和限制中描述,索引仍受你搜索服务层级的indexer限制约束。
图像尺寸必须在50像素×50像素或10,000像素×10,000像素之间。
如果你的PDF被密码锁住了,运行索引器前先解除密码。
技能参数
参数区分大小写。
| 参数名称 | 允许的值 | 说明 |
|---|---|---|
extractionOptions |
["images"], ["images", "locationMetadata"], ["locationMetadata"] |
识别从文档中提取的任何额外内容。 定义一个对应输出内容的枚举数组。 例如,如果 extractionOptions , ["images", "locationMetadata"]输出包含图片和位置元数据,提供页面位置和与内容提取地点相关的视觉信息。 |
modelName (预览版) |
字符串,如 "gpt-4.1". |
Optional. 从 REST API 开始 2026-05-01-preview 可用。 Azure OpenAI 聊天完成模型的名称,用于生成嵌入图像、图表和图表的说明。 映像说明独立于 extractionOptions 并且无需提取图像即可启用。 必须一 modelDeployment起指定 。 有关支持模型的列表,请参阅 支持的生成模型。 |
modelDeployment (预览版) |
String. | Optional. 从 REST API 开始 2026-05-01-preview 可用。 附加到技能集的 Foundry 资源的 Azure OpenAI 模型的部署名称。 必须一 modelName起指定 。 |
chunkingProperties |
请参见下表。 | 可以选择如何分块文本内容。 |
chunkingProperties 参数 |
允许的值 | 说明 |
|---|---|---|
method |
fixedSize (默认值)或 semantic (预览版)。 从 REST API 开始 2026-05-01-preview 可用。 |
分块策略。
fixedSize 使用基于字符的开窗分块。
semantic 使用布局感知分块,尊重段落边界,并智能处理跨区块边界的大型表。 |
unit |
characters (含 fixedSize)或 tokens (预览版,从 semanticREST API 开始 2026-05-01-preview 提供)。 |
控制区块单元的基数。 仅支持和fixedSize + characters组合。semantic + tokens 如果 unit 省略,则从中 method推断出来。 |
maximumLength |
如果 unit 为 characters,介于 300 和 50,000 之间的整数。 如果 unit 为 tokens100 到 8,000 之间的整数,则为 100 到 8,000。 默认值为 500。 |
在配置的 unit区块长度中测量的最大区块长度。 |
overlapLength |
Integer. 该值必须小于一 maximumLength半。 |
两个文本区块之间的重叠长度。 仅当为 method.fixedSize 必须省略或设置为0何时methodsemantic。 |
技能输入
| 输入名称 | 说明 |
|---|---|
file_data |
内容应提取的文件。 |
file_data输入必须是定义为:
{
"$type": "file",
"data": "BASE64 encoded string of the file"
}
或者,也可以定义为:
{
"$type": "file",
"url": "URL to download the file",
"sasToken": "OPTIONAL: SAS token for authentication if the provided URL is for a file in blob storage"
}
文件引用对象可以通过以下方式之一生成:
将你的索引器定义
allowSkillsetToReadFileData参数设置为true。 这个设置/document/file_data创建了一个路径,该路径代表从你的 blob 数据源下载的原始文件数据。 该参数仅适用于 Azure Blob 存储 中的文件。allowSkillsetToReadFileData使下载的文件数据可供技能使用。 它不会增加 Blob 索引器限制 或 数据限制中所述的内容理解服务限制。拥有自定义技能返回一个JSON对象定义,提供
$type、data、urlsastoken和 。$type参数必须设置为file,并且data必须是文件内容的基础64编码字节数组。url参数必须是有效的URL,并且有权限在该位置下载文件。
技能输出
| 输出名称 | 说明 |
|---|---|
text_sections |
一组文本块对象。 每个区块可以跨越多个页面(考虑了更多的分块配置)。 文本区块对象包括 locationMetadata (如果适用)以及 imagePath 当区块与文档中的数字跨度重叠时的列表。 |
normalized_images |
仅当 extractionOptions 包含 images时才适用。 一组从文档中提取的图片,包括 locationMetadata 如适用的话。 |
每个 text_sections 元素都有以下字段:
| 领域 | 类型 | 说明 |
|---|---|---|
id |
字符串 | 区块的唯一标识符。 |
content |
字符串 | 区块的 Markdown 内容。 届时methodsemantic,内容包括以 Markdown 形式内联的图形和表的 AI 生成的说明。 |
locationMetadata |
对象 | 页面范围和位置数据(pageNumberFrom、、pageNumberToordinalPosition、source)。 包含extractionOptions时locationMetadata显示 。 |
imagePath |
字符串 | 分号分隔的区块中包含的图像的路径列表。 当区块与文档中的数字跨度重叠时出现。 |
每个 normalized_images 元素都有以下字段:
| 领域 | 类型 | 说明 |
|---|---|---|
id |
字符串 | 图像的唯一标识符。 |
data |
字符串 | Base64 编码的图像数据。 |
imagePath |
字符串 | 对文档中图像的路径引用,例如 "figures/0"。 |
locationMetadata |
对象 | 页面范围和位置数据。 包含extractionOptions时locationMetadata显示 。 |
示例
第一个示例使用固定大小的分块,并演示如何输出固定大小的区块中的文本内容,以及文档的位置元数据。 第二个示例(从 2026-05-01-preview REST API 开始提供)通过 AI 生成的图像说明使用语义分块。
示例 1:使用图像和元数据提取固定大小的分块
{
"skills": [
{
"description": "Analyze a document",
"@odata.type": "#Microsoft.Skills.Util.ContentUnderstandingSkill",
"context": "/document",
"extractionOptions": ["images", "locationMetadata"],
"chunkingProperties": {
"unit": "characters",
"maximumLength": 1325,
"overlapLength": 0
},
"inputs": [
{
"name": "file_data",
"source": "/document/file_data"
}
],
"outputs": [
{
"name": "text_sections",
"targetName": "text_sections"
},
{
"name": "normalized_images",
"targetName": "normalized_images"
}
]
}
]
}
示例输出
{
"text_sections": [
{
"id": "1_d4545398-8df1-409f-acbb-f605d851ae85",
"content": "What is Azure Content Understanding (preview)?09/16/2025Important· Azure Al Content Understanding is available in preview. Public preview releases provide early access to features that are in active development.· Features, approaches, and processes can change or have limited capabilities, before General Availability (GA).. For more information, see Supplemental Terms of Use for Microsoft Azure PreviewsAzure Content Understanding is a Foundry Tool that uses generative AI to process/ingest content of many types (documents, images, videos, and audio) into a user-defined output format.Content Understanding offers a streamlined process to reason over large amounts of unstructured data, accelerating time-to-value by generating an output that can be integrated into automation and analytical workflows.<figure>\n\nInputs\n\nAnalyzers\n\nOutput\n\n0\nSearch\n\nContent Extraction\n\nField Extraction\n\nDocuments\n\nNew\n\nAgents\n\nPreprocessing\n\nEnrichments\n\nReasoning\n\nImage\n\nNormalization\n(resolution,\nformats)\n\nSpeaker\nrecognition\n\nGen Al\nContext\nwindows\n\nPostprocessing\nConfidence\nscores\nGrounding\nNormalization\n\nMulti-file input\nReference data\n\nDatabases\n\nVideo\n\nOrientation /\nde-skew\n\nLayout and\nstructure\n\nPrompt tuning\n\nStructured\noutput\n\nAudio\n\nFace grouping\n\nMarkdown or JSON schema\n\nCopilots\n\nApps\n\n\\+\n\nFaurIC\n\n</figure>",
"locationMetadata": {
"pageNumberFrom": 1,
"pageNumberTo": 1,
"ordinalPosition": 0,
"source": "D(1,0.6348,0.3598,7.2258,0.3805,7.223,1.2662,0.632,1.2455);D(1,0.6334,1.3758,1.3896,1.3738,1.39,1.5401,0.6338,1.542);D(1,0.8104,2.0716,1.8137,2.0692,1.8142,2.2669,0.8109,2.2693);D(1,1.0228,2.5023,7.6222,2.5029,7.6221,3.0075,1.0228,3.0069);D(1,1.0216,3.1121,7.3414,3.1057,7.342,3.6101,1.0221,3.6165);D(1,1.0219,3.7145,7.436,3.7048,7.4362,3.9006,1.0222,3.9103);D(1,0.6303,4.3295,7.7875,4.3236,7.7879,4.812,0.6307,4.8179);D(1,0.6304,5.0295,7.8065,5.0303,7.8064,5.7858,0.6303,5.7849);D(1,0.635,5.9572,7.8544,5.9573,7.8562,8.6971,0.6363,8.6968);D(1,0.6381,9.1451,5.2731,9.1476,5.2729,9.4829,0.6379,9.4803)"
}
},
...
{
"id": "2_e0e57fd4-e835-4879-8532-73a415e47b0b",
"content": "<table>\n<tr>\n<th>Application</th>\n<th>Description</th>\n</tr>\n<tr>\n<td>Post-call analytics</td>\n<td>Businesses and call centers can generate insights from call recordings to track key KPIs, improve product experience, generate business insights, create differentiated customer experiences, and answer queries faster and more accurately.</td>\n</tr>\n<tr>\n<th>Application</th>\n<th>Description</th>\n</tr>\n<tr>\n<td>Media asset management</td>\n<td>Software and media vendors can use Content Understanding to extract richer, targeted information from videos for media asset management solutions.</td>\n</tr>\n<tr>\n<td>Tax automation</td>\n<td>Tax preparation companies can use Content Understanding to generate a unified view of information from various documents and create comprehensive tax returns.</td>\n</tr>\n<tr>\n<td>Chart understanding</td>\n<td>Businesses can enhance chart understanding by automating the analysis and interpretation of various types of charts and diagrams using Content Understanding.</td>\n</tr>\n<tr>\n<td>Mortgage application processing</td>\n<td>Analyze supplementary supporting documentation and mortgage applications to determine whether a prospective home buyer provided all the necessary documentation to secure a mortgage.</td>\n</tr>\n<tr>\n<td>Invoice contract verification</td>\n<td>Review invoices and contr",
"locationMetadata": {
"pageNumberFrom": 2,
"pageNumberTo": 3,
"ordinalPosition": 3,
"source": "D(2,0.6438,9.2645,7.8576,9.2649,7.8565,10.5199,0.6434,10.5194);D(3,0.6494,0.3919,7.8649,0.3929,7.8639,4.3254,0.6485,4.3232)"
}
...
}
],
"normalized_images": [
{
"id": "1_335140f1-9d31-4507-8916-2cde758639cb",
"data": "aW1hZ2UgMSBkYXRh",
"imagePath": "aHR0cHM6Ly9henNyb2xsaW5nLmJsb2IuY29yZS53aW5kb3dzLm5ldC9tdWx0aW1vZGFsaXR5L0NVLnBkZg2/normalized_images_0.jpg",
"locationMetadata": {
"pageNumberFrom": 1,
"pageNumberTo": 1,
"ordinalPosition": 0,
"source": "D(1,0.635,5.9572,7.8544,5.9573,7.8562,8.6971,0.6363,8.6968)"
}
},
{
"id": "3_699d33ac-1a1b-4015-9cbd-eb8bfff2e6b4",
"data": "aW1hZ2UgMiBkYXRh",
"imagePath": "aHR0cHM6Ly9henNyb2xsaW5nLmJsb2IuY29yZS53aW5kb3dzLm5ldC9tdWx0aW1vZGFsaXR5L0NVLnBkZg2/normalized_images_1.jpg",
"locationMetadata": {
"pageNumberFrom": 3,
"pageNumberTo": 3,
"ordinalPosition": 1,
"source": "D(3,0.6353,5.2142,7.8428,5.218,7.8443,8.4631,0.6363,8.4594)"
}
}
]
}
locationMetadata基于内容理解提供的source属性Azure。 有关文件中元素的可视化位置如何编码,请参见 文档分析:提取结构化内容。
imagePath 表示存储图像的相对路径。 如果技能集中配置了知识库文件投影,该路径与存储在知识库中的图像的相对路径相匹配。
示例 2:带有图像说明的语义分块(预览版)
此示例从 REST API 开始 2026-05-01-preview 提供,使用语义分块并生成嵌入式图像、图表和关系图的 AI 生成的说明。 附加到技能组的 Foundry 资源必须具有由 modelName 和已部署 modelDeployment的聊天完成模型标识。
{
"skills": [
{
"description": "Extract and chunk document content with image descriptions",
"@odata.type": "#Microsoft.Skills.Util.ContentUnderstandingSkill",
"context": "/document",
"modelName": "gpt-4.1",
"modelDeployment": "myGpt41Deployment",
"extractionOptions": ["images", "locationMetadata"],
"chunkingProperties": {
"method": "semantic",
"unit": "tokens",
"maximumLength": 500
},
"inputs": [
{
"name": "file_data",
"source": "/document/file_data"
}
],
"outputs": [
{
"name": "text_sections",
"targetName": "text_sections"
},
{
"name": "normalized_images",
"targetName": "normalized_images"
}
]
}
]
}
借助语义分块,每个区块 text_sections 都包含 Markdown 内容,其中包括它涵盖的任何图形和表的 AI 生成的说明。 当区块与一个或多个图形跨度重叠时,区块对象还包括列出 imagePath 相应图像路径的字段:
{
"id": "1_d4545398-8df1-409f-acbb-f605d851ae85",
"content": "# Architecture overview\n\nThe following diagram summarizes the ingestion pipeline...\n\n<figure>The diagram shows three stages: Inputs, Analyzers, and Output. Inputs include documents, images, video, and audio. Analyzers perform preprocessing, enrichments, and reasoning. Output is structured Markdown or JSON consumed by search, agents, copilots, and apps.</figure>",
"locationMetadata": {
"pageNumberFrom": 1,
"pageNumberTo": 1,
"ordinalPosition": 0,
"source": "D(1,0.6348,0.3598,7.2258,0.3805,7.223,1.2662,0.632,1.2455)"
},
"imagePath": "aHR0cHM6Ly9henNyb2xsaW5nLmJsb2IuY29yZS53aW5kb3dzLm5ldC9tdWx0aW1vZGFsaXR5L0NVLnBkZg2/normalized_images_0.jpg"
}
相关内容
- Foundry Tools中的Azure内容理解是什么?
- 内置技能
- 打造一套技能组合
- 索引器 - 创建 (REST API)