Formerly known as Azure AI Services or Azure Cognitive Services is a unified collection of prebuilt AI capabilities within the Microsoft Foundry platform
Hi @Riley McCann ,
Thank you for the detailed explanation.
What you are seeing is a current limitation in Foundry's agent evaluator support for Azure AI Search, rather than a configuration issue with your evaluation.
Microsoft's current Agent Evaluators documentation states that Azure AI Search currently has limited support with the agent evaluators and recommends avoiding the Groundedness, Tool Output Utilization, Tool Call Accuracy, Tool Input Accuracy, and Tool Call Success evaluators when the agent conversation includes calls to Azure AI Search.
This explains why the Groundedness evaluation can complete without usable Azure AI Search tool results in the evaluation data, even though the agent itself successfully performed the search. There is currently no supported portal setting to make the native Azure AI Search retrieval results available to these evaluators in the same way as supported user-defined tools.
For evaluating groundedness, the recommended approach is to provide the retrieved content explicitly as context in the evaluation dataset. The Groundedness evaluator can then evaluate the response against that supplied context. For best results, Microsoft recommends providing the query, response, and context fields.
For example:
{
"query": "your test question",
"context": "the content retrieved from your Azure AI Search index",
"response": "the agent response"
}
You can then map the evaluator inputs to:
query -> {{item.query}}
context -> {{item.context}}
response -> {{item.response}}
If your goal is specifically to evaluate the quality of the Azure AI Search retrieval itself, Microsoft also provides the Retrieval and Document Retrieval evaluators for this purpose.
Therefore, for the current Foundry implementation, I recommend evaluating the retrieved content separately by supplying it as context, rather than relying on the native Azure AI Search tool results to provide the context to the Groundedness evaluator.
I hope this clarifies the current limitation and the supported evaluation approach.
References:
Microsoft Learn: Agent evaluators Agent evaluators for generative AI
Microsoft Learn: RAG evaluators Retrieval-Augmented Generation (RAG) evaluators
Microsoft Learn: Evaluation dataset schema Evaluation dataset schema in Microsoft Foundry