Azure AI 搜尋服務 (preview) 中的多向量欄位支援

註

Azure AI 搜尋服務 可透過 Azure 入口網站、REST API 及 Azure SDK 取得。 它同時也是 Foundry IQ 的基礎,這是一個管理式知識層,能將企業內容轉化為可重複使用、權限感知的知識庫,供 Microsoft Foundry 入口網站中的代理使用。

Important

標記(預覽)的功能、能力或屬性不受服務等級協議涵蓋,也不建議用於生產工作負載,且在正式上架前可能會有所變動或受限。 Azure AI 搜尋服務 預覽條款適用於所有預覽功能,無論是獨立功能還是正式推出功能的一部分。

Azure AI 搜尋服務 中的多向量欄位支援功能(預覽)允許您在單一文件欄位中索引多個子向量。 此功能對於多模態資料或長篇文件等使用情境非常有價值,因為若只用單一向量表示內容,會損失重要細節。

限制

  • 語意排名器不支援複雜欄位內的巢狀區塊。 因此,語意排序器不支援多向量場中的巢狀向量。

了解多向量場支援

傳統上,向量類型 Collection(Edm.Single) 只能用於頂層欄位。 隨著多向量欄位支援的引入,你現在可以在複雜集合的巢狀欄位中使用向量類型,實際上允許多個向量與單一文件關聯。

一份文件可包含最多 100 個向量,涵蓋所有複雜的集合欄位。 向量欄位的巢狀結構僅支援一層深度。

多向量場的指標定義

此功能不需要新的索引屬性。 以下是指數定義範例:

{
  "name": "multivector-index",
  "fields": [
    {
      "name": "id",
      "type": "Edm.String",
      "key": true,
      "searchable": true
    },
    {
      "name": "title",
      "type": "Edm.String",
      "searchable": true
    },
    {
      "name": "description",
      "type": "Edm.String",
      "searchable": true
    },
    {
      "name": "descriptionEmbedding",
      "type": "Collection(Edm.Single)",
      "dimensions": 3,
      "searchable": true,
      "retrievable": true,
      "vectorSearchProfile": "hnsw"
    },
    {
      "name": "scenes",
      "type": "Collection(Edm.ComplexType)",
      "fields": [
        {
          "name": "embedding",
          "type": "Collection(Edm.Single)",
          "dimensions": 3,
          "searchable": true,
          "retrievable": true,
          "vectorSearchProfile": "hnsw"
        },
        {
          "name": "timestamp",
          "type": "Edm.Int32",
          "retrievable": true
        },
        {
          "name": "description",
          "type": "Edm.String",
          "searchable": true,
          "retrievable": true
        },
        {
          "name": "framePath",
          "type": "Edm.String",
          "retrievable": true
        }
      ]
    }
  ]
}

樣本內嵌文件

這裡有一份範例文件,說明你如何在實務中使用多向量場:

{
  "id": "123",
  "title": "Non-Existent Movie",
  "description": "A fictional movie for demonstration purposes.",
  "descriptionEmbedding": [1, 2, 3],
  "releaseDate": "2025-08-01",
  "scenes": [
    {
      "embedding": [4, 5, 6],
      "timestamp": 120,
      "description": "A character is introduced.",
      "framePath": "nonexistentmovie\\scenes\\scene120.png"
    },
    {
      "embedding": [7, 8, 9],
      "timestamp": 2400,
      "description": "The climax of the movie.",
      "framePath": "nonexistentmovie\\scenes\\scene2400.png"
    }
  ]
}

在這個例子中,場景欄位是一個包含多個向量(嵌入欄位)及其他相關資料的複雜集合。 每個向量代表電影中的一個場景,並可用來尋找其他電影中的類似場景,以及其他潛在的應用場景。

支援多向量欄位的查詢

多向量場支援功能對 Azure AI 搜尋服務 的查詢機制帶來了一些變更。 然而,主要的查詢流程大致相同。 過去,vectorQueries 只能針對定義為頂層索引欄位的向量場。 透過此功能,我們放寬這項限制,並允許 vectorQueries 以巢狀於複雜型別集合內的欄位為目標 (最多一層深)。 此外,新增了一個查詢時間參數: perDocumentVectorLimit。

  • 設定 perDocumentVectorLimit 為 1 確保每份文件最多匹配一個向量,確保結果來自不同文件。
  • 設定 perDocumentVectorLimit 為 0 (無限制)則可以匹配同一文件中多個相關向量。
{
  "vectorQueries": [
    {
      "kind": "text",
      "text": "whales swimming",
      "K": 50,
      "fields": "scenes/embedding",
      "perDocumentVectorLimit": 0
    }
  ],
  "select": "title, scenes/timestamp, scenes/framePath"
}

在同一欄位中針對多個向量進行排名

當多個向量與單一文件相關聯時,Azure AI 搜尋服務 會使用它們之間的最高分數來排名。 系統使用最相關的向量來評分每份文件,避免被較不相關的文件稀釋。

在集合中檢索相關元素

當參數包含 $select 一組複數型態時,僅回傳與向量查詢相符的元素。 這對於擷取相關的元資料(如時間戳記、文字描述或圖片路徑)非常有用。

註

為了減少有效載荷大小,避免在參數中包含向量值本身 $select 。 如果不必要,考慮完全省略向量儲存。

除錯多向量查詢(預覽)

當文件包含多個嵌入向量,例如不同子欄位的文字與影像嵌入時,系統會使用所有元素中最高的向量分數來排名文件。

要除錯每個向量的貢獻,請使用 innerHits 除錯模式(最新預覽版 REST API 中提供)。

POST /indexes/my-index/docs/search?api-version=2026-08-01-preview
{
  "vectorQueries": [
    {
      "kind": "vector",
      "field": "keyframes.imageEmbedding",
      "kNearestNeighborsCount": 5,
      "vector": [ /* query vector */ ]
    }
  ],
  "debug": "innerHits"
}

響應形狀範例

"@search.documentDebugInfo": {
  "innerHits": {
    "keyframes": [
      {
        "ordinal": 0,
        "vectors": [
          {
            "imageEmbedding": {
              "searchScore": 0.958,
              "vectorSimilarity": 0.956
            },
            "textEmbedding": {
              "searchScore": 0.958,
              "vectorSimilarity": 0.956
            }
          }
        ]
      },
      {
        "ordinal": 1,
        "vectors": [
          {
            "imageEmbedding": null,
            "textEmbedding": {
              "searchScore": 0.872,
              "vectorSimilarity": 0.869
            }
          }
        ]
      }
    ]
  }
}

田野描述

欄位 描述
ordinal 集合內元素的零基索引。
vectors 每個元素中可搜尋的向量欄位各有一個項目。
searchScore 該欄位的最終分數 (在經過任何重新評分和加權之後)。
vectorSimilarity 原始相似度由距離函數回傳。

註

innerHits 目前僅報告向量場。

與 debug=vector 的關聯

以下是關於此物業的一些事實:

  • 現有的 debug=vector 開關保持不變。

  • 當與多向量欄位一起使用時,會 @search.documentDebugInfo.vector.subscore 顯示用於排序父文件的最高分數,但不顯示每個元素的細節。

  • 用 innerHits 來深入了解各個元素如何對配樂產生影響。