為 Azure AI 搜尋服務中的查詢,從 Azure Cosmos DB for Apache Gremlin 編製索引資料 (預覽版)

Note

Azure AI 搜尋服務 可透過 Azure 入口網站、REST API 及 Azure SDK 取得。 它同時也是 Foundry IQ 的基礎,這是一個管理式知識層,能將企業內容轉化為可重複使用、權限感知的知識庫,供 Microsoft Foundry 入口網站中的代理使用。

重要

標記(預覽)的功能、能力或屬性不受服務等級協議涵蓋,也不建議用於生產工作負載,且在正式上架前可能會有所變動或受限。 Azure AI 搜尋服務 預覽條款適用於所有預覽功能,無論是獨立功能還是正式推出功能的一部分。

重要

這些功能支援與其他 Microsoft 服務 及第三方服務的連結。 使用這些服務須遵守其各自的條款,可能導致資料處理或儲存超出 Azure 合規邊界,以及資料流入 Azure 合規邊界。

你有責任管理資料是否會超出組織的合規與地理邊界及相關影響,並確保適當的權限、邊界與核准被提供。

你有責任仔細審查並測試你在特定使用情境中所建置的應用程式,並做出所有適當的決策與客製化。 這包括實施你自己負責任的 AI 緩解措施,例如元提示、內容過濾器或其他安全系統,並確保你的應用程式符合適當的品質、可靠性、安全性與可信度標準。 欲了解更多資訊,請參閱Azure AI 搜尋服務透明度說明。

Azure Cosmos DB for Apache Gremlin indexer (preview)會從 Azure Cosmos DB for Apache Gremlin 匯入內容,並讓內容可於 Azure AI 搜尋服務 中搜尋。

本文補充建立索引子,並提供 Cosmos DB 專屬資訊。 它使用 REST API 展示所有索引器共有的三步驟工作流程:建立資料來源、建立索引、建立索引器。 資料擷取發生在你提交建立索引器請求時。

由於術語可能令人混淆,值得注意的是Azure Cosmos DB索引與Azure AI 搜尋服務索引是不同的操作。 在 Azure AI 搜尋服務 中建立並載入搜尋索引時,會在您的搜尋服務中生成該索引。

先決條件

定義資料來源

資料來源定義指定了索引資料、憑證及識別資料變更的政策。 資料來源被定義為獨立的資源,以便多個索引器都能使用。

在此通話中,請指定預覽版 REST API 版本,以建立一個可透過 Azure Cosmos DB for Apache Gremlin 連接的資料來源。 你可以用 2021-04-01-preview ,也可以以後再用。 我們推薦 使用最新的預覽版 REST API。

  1. 建立或更新資料來源 以設定其定義:

     POST https://[service name].search.windows.net/datasources?api-version=2026-08-01-preview
     Content-Type: application/json
     api-key: [Search service admin key]
     {
       "name": "[my-cosmosdb-gremlin-ds]",
       "type": "cosmosdb",
       "credentials": {
         "connectionString": "AccountEndpoint=https://[cosmos-account-name].documents.azure.com;AccountKey=[cosmos-account-key];Database=[cosmos-database-name];ApiKind=Gremlin;"
       },
       "container": {
         "name": "[cosmos-db-collection]",
         "query": "g.V()"
       },
       "dataChangeDetectionPolicy": {
         "@odata.type": "#Microsoft.Azure.Search.HighWaterMarkChangeDetectionPolicy",
         "highWaterMarkColumnName": "_ts"
       },
       "dataDeletionDetectionPolicy": null,
       "encryptionKey": null,
       "identity": null
     }
    
  2. 將「type」設為 "cosmosdb" (必需)。

  3. 將「credentials」設為連線字串。 下一節將介紹所支援的格式。

  4. 將「container」添加到集合中。 「name」屬性是必填的,它指定了圖的 ID。

    「查詢」屬性是可選的。 預設情況下,Apache Gremlin Azure Cosmos DB 的 Azure AI 搜尋服務 索引器會將圖中的每個頂點設為索引中的文件。 邊緣則被忽略。 查詢預設為 g.V()。 或者,您也可以將查詢設定為只為邊緣編製索引。 要為邊建立索引,請將查詢設定為 g.E()。

  5. 如果你的資料是易失性,且你希望索引器在後續執行中只偵測新項目和更新項目,請設定「dataChangeDetectionPolicy」。 依預設會啟用累加式進度,並使用 _ts 作為最高水位標記資料行。

  6. 如果你想在刪除來源項目時從搜尋索引中移除搜尋文件,請設定「dataDeletionDetectionPolicy」。

支援的憑證與連線字串

索引器可以透過以下連線連接到集合。 對於目標為 Azure Cosmos DB for Apache Gremlin 的連線,請務必在連線字串中加入「ApiKind」。

避免在端點網址中顯示埠號。 如果您包含連接埠號碼,連線將會失敗。

完整存取連接字串
{ "connectionString" : "AccountEndpoint=https://<Cosmos DB account name>.documents.azure.com;AccountKey=<Cosmos DB auth key>;Database=<Cosmos DB database id>;ApiKind=Gremlin" }
你可以在Azure入口網站的Azure Cosmos DB帳戶頁面左側面板中選擇Keys來獲取連接字串。 務必選擇完整的連線字串,而非僅選擇一個金鑰。
管理身份連接字串
{ "connectionString" : "ResourceId=/subscriptions/<your subscription ID>/resourceGroups/<your resource group name>/providers/Microsoft.DocumentDB/databaseAccounts/<your cosmos db account name>/;(ApiKind=[api-kind];)" }
此連接字串不需要帳號金鑰,但您必須先設定搜尋服務以透過受控識別連接,並建立一個角色指派來授予 Cosmos DB 帳號閱讀者角色的許可權。 更多資訊請參見 使用受控身分識別設定索引器連接至 Azure Cosmos DB 資料庫。

在索引中新增搜尋欄位

在 搜尋索引中,新增欄位以接受原始 JSON 文件或自訂查詢投影的輸出。 確保搜尋索引結構與你的圖表相容。 對於Azure Cosmos DB內容,你的搜尋索引結構應該對應於資料來源中的Azure Cosmos DB項目。

  1. 建立或更新索引,以定義儲存資料的搜尋欄位:

     POST https://[service name].search.windows.net/indexes?api-version=2026-08-01-preview
     Content-Type: application/json
     api-key: [Search service admin key]
     {
        "name": "mysearchindex",
        "fields": [
         {
             "name": "rid",
             "type": "Edm.String",
             "facetable": false,
             "filterable": false,
             "key": true,
             "retrievable": true,
             "searchable": true,
             "sortable": false,
             "analyzer": "standard.lucene",
             "indexAnalyzer": null,
             "searchAnalyzer": null,
             "synonymMaps": [],
             "fields": []
         }, {
             "name": "label",
             "type": "Edm.String",
             "searchable": true,
             "filterable": false,
             "retrievable": true,
             "sortable": false,
             "facetable": false,
             "key": false,
             "indexAnalyzer": null,
             "searchAnalyzer": null,
             "analyzer": "standard.lucene",
             "synonymMaps": []
        }]
      }
    
  2. 建立文件鍵欄位(「鍵」:真)。 對於分割式集合,預設文件鍵為 Azure Cosmos DB _rid,Azure AI 搜尋服務會自動重新命名為 rid,因為欄位名稱不能以底線字元開頭。 此外,Azure Cosmos DB _rid 值包含在 Azure AI 搜尋服務 鍵中無效的字元。 因此,這些 _rid 數值是用 Base64 編碼的。

  3. 新增欄位以提供更多可搜尋的內容。 詳情請參見 建立索引 。

映射資料類型

JSON 資料型別 Azure AI 搜尋服務 欄位類型
Bool Edm.Boolean、Edm.String
看起來像整數的數字 Edm.Int32、Edm.Int64、Edm.String
看起來像浮點數的數字 Edm.Double、Edm.String
弦 埃德姆·斯特林
原始型態陣列如 [“a”、“b”、“c”] Collection(Edm.String)
看起來像日期的字串 Edm.DateTimeOffset, Edm.String
GeoJSON 物件如 { “type”: “Point”, “coordinates”: [long, lat] } Edm.GeographyPoint
其他 JSON 物件 無

配置並執行 Azure Cosmos DB 索引器

一旦索引和資料來源建立完成,你就可以開始建立索引器了。 索引器配置指定控制執行時行為的輸入、參數與屬性。

  1. 透過命名索引器並引用資料來源與目標索引,建立或更新它:

    POST https://[service name].search.windows.net/indexers?api-version=2026-08-01-preview
    Content-Type: application/json
    api-key: [search service admin key]
    {
        "name" : "[my-cosmosdb-indexer]",
        "dataSourceName" : "[my-cosmosdb-gremlin-ds]",
        "targetIndexName" : "[my-search-index]",
        "disabled": null,
        "schedule": null,
        "parameters": {
            "batchSize": null,
            "maxFailedItems": 0,
            "maxFailedItemsPerBatch": 0,
            "base64EncodeKeys": false,
            "configuration": {}
            },
        "fieldMappings": [],
        "encryptionKey": null
    }
    
  2. 如果欄位名稱或類型有差異,或搜尋索引中需要多個來源欄位版本,請指定欄位對應。

  3. 請參閱 建立索引器 以了解更多其他屬性的資訊。

索引器在建立時會自動執行。 你可以把「停用」設為 true,來避免這種情況。 要控制索引器的執行,請按 需求執行索引器 或 將其列入排程。

檢查索引器狀態

要監控索引器狀態與執行歷史,請發送 「取得索引器狀態 」請求:

GET https://myservice.search.windows.net/indexers/myindexer/status?api-version=2026-08-01-preview
  Content-Type: application/json  
  api-key: [admin key]

回應內容包括狀態及已處理項目數量。 它應該看起來像以下範例:

    {
        "status":"running",
        "lastResult": {
            "status":"success",
            "errorMessage":null,
            "startTime":"2022-02-21T00:23:24.957Z",
            "endTime":"2022-02-21T00:36:47.752Z",
            "errors":[],
            "itemsProcessed":1599501,
            "itemsFailed":0,
            "initialTrackingState":null,
            "finalTrackingState":null
        },
        "executionHistory":
        [
            {
                "status":"success",
                "errorMessage":null,
                "startTime":"2022-02-21T00:23:24.957Z",
                "endTime":"2022-02-21T00:36:47.752Z",
                "errors":[],
                "itemsProcessed":1599501,
                "itemsFailed":0,
                "initialTrackingState":null,
                "finalTrackingState":null
            },
            ... earlier history items
        ]
    }

執行歷史包含最多 50 次最近完成的執行,並依逆時間順序排序,使最新的執行先行。

索引新文件和已更改的文件

索引器完成搜尋索引後,你可能會希望後續的索引器執行時,只針對資料庫中新增或變更的文件進行增量索引。

要啟用增量索引,請在資料來源定義中設定「dataChangeDetectionPolicy」屬性。 此特性告訴索引器您的資料使用了哪種變更追蹤機制。

對於 Azure Cosmos DB 索引子,唯一支援的原則是 HighWaterMarkChangeDetectionPolicy,並使用 Azure Cosmos DB 提供的 _ts (時間戳記) 屬性。

以下範例展示了帶有變更偵測政策的資料 來源定義 :

"dataChangeDetectionPolicy": {
    "@odata.type": "#Microsoft.Azure.Search.HighWaterMarkChangeDetectionPolicy",
    "highWaterMarkColumnName": "_ts"
},

已刪除文件的索引

當圖表資料被刪除時,你可能也想從搜尋索引中刪除其對應的文件。 資料刪除偵測策略的目的是有效識別已刪除的資料項目,並從索引中刪除整份文件。 資料刪除偵測政策並非用來刪除部分文件資訊。 目前唯一支援的政策是 Soft Delete 政策(刪除會以某種旗標標記),其在資料來源定義中具體說明如下:

"dataDeletionDetectionPolicy": {
    "@odata.type" : "#Microsoft.Azure.Search.SoftDeleteColumnDeletionDetectionPolicy",
    "softDeleteColumnName" : "the property that specifies whether a document was deleted",
    "softDeleteMarkerValue" : "the value that identifies a document as deleted"
}

以下範例建立一個具有軟刪除政策的資料來源:

POST https://[service name].search.windows.net/datasources?api-version=2026-08-01-preview
Content-Type: application/json
api-key: [Search service admin key]

{
    "name": "[my-cosmosdb-gremlin-ds]",
    "type": "cosmosdb",
    "credentials": {
        "connectionString": "AccountEndpoint=https://[cosmos-account-name].documents.azure.com;AccountKey=[cosmos-account-key];Database=[cosmos-database-name];ApiKind=Gremlin"
    },
    "container": { "name": "[my-cosmos-collection]" },
    "dataChangeDetectionPolicy": {
        "@odata.type": "#Microsoft.Azure.Search.HighWaterMarkChangeDetectionPolicy",
        "highWaterMarkColumnName": "`_ts`"
    },
    "dataDeletionDetectionPolicy": {
        "@odata.type": "#Microsoft.Azure.Search.SoftDeleteColumnDeletionDetectionPolicy",
        "softDeleteColumnName": "isDeleted",
        "softDeleteMarkerValue": "true"
    }
}

即使啟用刪除偵測政策,也不支援刪除索引中複雜的(Edm.ComplexType)欄位。 此政策要求 Gremlin 資料庫中的「active」欄位必須為整數、字串或布林型。

將圖表資料映射到搜尋索引中的欄位

Azure Cosmos DB for Apache Gremlin 索引器會自動映射幾個圖形資料:

  1. 索引器會映射 _rid 到索引中的欄位 rid (如果存在),而 Base64 會將其編碼。

  2. 索引器會映射 _id 到索引中的欄位 id (如果存在)。

  3. 當您使用 Azure Cosmos DB for Apache Gremlin 查詢 Azure Cosmos DB 資料庫時,可能會注意到每個屬性的 JSON 輸出中都有一個 id 和一個 value。 索引子會自動將該屬性的 value 對應到搜尋索引中一個欄位;如果存在的話,該欄位的名稱會與該屬性相同。 以下範例中,450 被映射到搜尋索引中的一個 pages 欄位。

    {
        "id": "Cookbook",
        "label": "book",
        "type": "vertex",
        "properties": {
          "pages": [
            {
              "id": "48cf6285-a145-42c8-a0aa-d39079277b71",
              "value": "450"
            }
          ]
        }
    }

你可能會發現需要使用 Output Field Mappings 來將查詢輸出映射到索引中的欄位。 你可能會想用輸出欄位映射(Output Field Mappings)代替 欄位映射 ,因為自訂查詢可能包含複雜的資料。

舉例來說,假設你的查詢產生了這樣的輸出:

    [
      {
        "vertex": {
          "id": "Cookbook",
          "label": "book",
          "type": "vertex",
          "properties": {
            "pages": [
              {
                "id": "48cf6085-a211-42d8-a8ea-d38642987a71",
                "value": "450"
              }
            ],
          }
        },
        "written_by": [
          {
            "yearStarted": "2017"
          }
        ]
      }
    ]

如果你想將上述 JSON 中的 值 pages 映射到索引中的欄位 totalpages ,可以在索引器定義中加入以下 輸出欄位 映射:

    ... // rest of indexer definition 
    "outputFieldMappings": [
        {
          "sourceFieldName": "/document/vertex/pages",
          "targetFieldName": "totalpages"
        }
    ]

注意輸出欄位映射是以 /document 開始,並且不包含 JSON 中屬性鍵的參考。 這是因為索引器在匯入圖資料時會將每個文件放在節點下方/document,且索引器也自動允許你透過簡單引用pages來參考 的pages值,而不必參考陣列pages中的第一個物件。

下一步