Information retrieval not working in my agents.

Mats van Borre 0 Reputation points
2026-10-07T08:49:16.0366667+00:00

Hi,

I am building several knowledge-based agents in Microsoft Copilot Studio using connected document sources.

Across multiple agents, I am seeing the same recurring problem: the correct information exists in the connected sources, but Copilot Studio does not consistently retrieve all relevant passages needed to answer a question completely.

The issue appears to be retrieval completeness rather than answer generation itself.

Typical pattern:

  • The user asks a question that requires information from one or more connected documents.
  • Copilot retrieves some relevant passages.
  • Other relevant passages that are also present in the connected sources are not retrieved.
  • The model then produces an incomplete, overly cautious, or incorrect answer based on the partial context it received.
  • If the missing concept is searched more directly or with different terminology, the relevant passage can often be found.

This behaviour occurs across different agents and different document sets, not just one specific source.

I have already tried:

  • detailed descriptions for knowledge sources;
  • instructions to search synonyms and alternative terminology;
  • instructions to perform additional searches when the first search is insufficient;
  • instructions to consider all relevant roles, conditions, exceptions and applicability;
  • instructions not to stop after finding only one part of the answer;
  • stricter source-use and comparison instructions.

These changes sometimes improve results, but they do not reliably solve the retrieval issue.

In the activity trace, I can often see that Copilot generates a refined knowledge query, but that query retrieves only part of the information needed. Sometimes the correct source is searched, but the relevant passage from that source is still not returned.

My questions are:

  1. Is there a supported way to improve retrieval completeness in Copilot Studio?
  2. Can the automatically generated/refined knowledge query be controlled or influenced more directly?
  3. Can an agent be forced to perform multiple independent knowledge searches for one question and then combine the retrieved results?
  4. Is there a way to increase the number or diversity of retrieved chunks?
  5. Can Copilot Studio be configured to make sure multiple relevant knowledge sources contribute to the answer instead of stopping after the first strong match?
  6. Would Azure AI Search provide more control over retrieval behaviour than the standard Copilot Studio knowledge system?
  7. Are there recommended Microsoft architectures or best practices for knowledge-based agents where completeness across multiple documents is important?

The main issue is that the information is present in the knowledge sources, but the retrieval layer does not consistently provide all relevant context to the model.

I would like to understand whether this is an expected limitation of the current Copilot Studio knowledge retrieval system, or whether there are configuration options or recommended patterns that can improve this.

Thanks.

Microsoft Copilot | Other
0 comments No comments

1 answer

Sort by: Most helpful
  1. AI answer

    2026-10-07T08:49:48.3+00:00

    Several parts of this behavior are expected limitations of the current SharePoint-based knowledge retrieval path in Copilot Studio.

    1. Retrieval completeness is limited by the built-in retrieval behavior.
    • For SharePoint searches, Copilot Studio uses only the top three search results to summarize and generate a response.
    • If relevant content is spread across more documents or pages than those top results, answers can be incomplete even when the information exists in the source set.
    • Copilot Studio also depends on Microsoft Search indexing, so missing or delayed indexing can prevent relevant passages from being retrieved.
    1. Direct control over the automatically refined knowledge query is not described as a supported configuration.
    • The available guidance focuses on improving source quality, indexing, filters, authentication, and supported content rather than tuning the generated query itself.
    • No supported setting is described for forcing a custom refined query, increasing retrieved chunk count, or requiring broader retrieval diversity.
    1. Forcing multiple independent searches and combining them is not described as a supported built-in capability.
    • The documented behavior is that the system searches SharePoint and uses the returned top results.
    • There is also no documented setting to force multiple knowledge sources to contribute before answer generation.
    1. Practical ways to improve retrieval quality in Copilot Studio are:
    • Verify the SharePoint location actually contains the relevant content.
    • Confirm the content is in supported formats.
    • Use only modern SharePoint pages; modern pages with SPFx components are not supported, and classic .aspx page content is not used for answers.
    • Check whether documents were uploaded recently and are not yet indexed.
    • In SharePoint, test retrieval by searching for a unique keyword from the target document. If it does not appear, indexing is incomplete.
    • If needed, have a SharePoint site administrator trigger reindexing at Site Information > View all site settings > Search and offline availability > Re-index site.
    • Review agent-level SharePoint filters in Knowledge > source > Edit > Advanced settings.
    • Review topic-level Create generative answers node settings in Properties > Data source, especially Search only selected sources. If this is on, only selected sources are searched.
    • If using SharePoint lists, ensure Dataverse search is enabled, keep list count to 10 or fewer, and avoid very large lists because lists over 35,000 rows can reduce quality and increase latency.
    1. Source and file constraints can also look like retrieval incompleteness.
    • Without a Microsoft 365 Copilot license in the same tenant, SharePoint generative answers process files only up to 7 MB, and Tenant graph grounding with semantic search must be off.
    • With a Microsoft 365 Copilot license in the same tenant, the maximum is 200 MB, and Tenant graph grounding with semantic search must be on.
    • Larger SharePoint files can still be returned by Microsoft Graph search but are not processed by generative answers.
    • Files can also fail to answer if they are not ready, use unsupported characters in the file name, have confidential sensitivity or password protection, are unsupported file types, or are not shared appropriately.
    1. Azure AI Search provides more control only in the sense that Copilot Studio’s native knowledge experience has documented limitations.
    • A Microsoft Q&A answer states that Copilot Studio’s built-in Azure AI Search knowledge experience primarily uses document or snippet content and limited citation information, and does not support mapping arbitrary metadata fields directly into generated answers.
    • That same answer points to custom knowledge sources for more advanced retrieval and orchestration scenarios.
    • Based on the available material, standard Copilot Studio knowledge retrieval has limited exposed controls, while custom knowledge source patterns are the documented direction for advanced retrieval behavior.
    1. Best-fit architecture when completeness across multiple documents matters:
    • First, tighten the standard SharePoint path: indexing, supported formats, page type, filters, authentication scopes, and source selection.
    • If completeness still remains a hard requirement, use a more custom retrieval architecture rather than relying only on the default SharePoint knowledge path.
    • The available guidance explicitly mentions alternatives such as custom knowledge sources, Microsoft 365 Semantic Indexing, or connecting data to Azure OpenAI for generative answers.

    In short: this is largely an expected limitation of the current built-in retrieval system. The context supports improving retrieval quality through indexing, source design, filters, and supported content checks, but it does not describe supported controls for directly steering the refined query, forcing multiple independent searches, increasing retrieved chunk diversity, or guaranteeing multi-source aggregation in the default knowledge system.


    AI-generated content may be incorrect. Read our transparency notes for more information.

    Was this answer helpful?

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.