An Azure service that is used to collect, analyze, and act on telemetry data from Azure and on-premises environments.
Based on the timings you've captured, I don't think the evidence currently points to the KQL query itself being slow.
Your successful comparisons are particularly useful here. For example, one request took approximately 8.826 seconds before response headers arrived, while the reported query-engine execution time was only about 26.8 ms. Your later comparison similarly took about 2.44 seconds to receive response headers, while query execution was approximately 11 ms.
That suggests most of the observed latency is occurring outside the measured query-engine execution interval. It could be somewhere in request receipt, authentication/authorization, routing, queueing, front-end processing, or response handling, but the client-side evidence alone cannot determine which component is responsible.
The original failure also differs slightly from a normal Azure Monitor Logs server-side query timeout. Microsoft documents a 3-minute default server timeout, configurable up to 10 minutes using Prefer: wait=<seconds>. When that server-side timeout is exceeded, the API returns HTTP 504. In your case, the client stopped waiting at its unchanged 20-second deadline and received no response headers, so I wouldn't classify that event as a confirmed Azure Logs query timeout/504.
Also, don't increase the server-side Prefer: wait value as the first fix here. Your query-engine timings are measured in milliseconds, while the unexplained delay occurs before the client receives the response headers.
For the next diagnostic run, add:
Prefer: include-statistics=true
Prefer: include-dataSources=true
include-statistics returns query execution/resource statistics, while include-dataSources identifies the regions, workspaces, clusters, and tables involved in processing the request.
Most importantly, generate and retain a unique client request ID for every attempt, including the attempt that times out. The successful provider/client request IDs you supplied are useful comparison points, but as you've correctly noted, they don't identify the original failed request.
Microsoft Q&A participants don't have access to Azure's internal service traces needed to break a request down into receive → authentication/routing → queue → execution → response timings. The IDs and UTC timestamps you've captured are therefore exactly the sort of information I would preserve for an Azure support investigation.
For the next occurrence, capture the UTC start/end time, workspace ID, client request ID, source environment, authentication method, client-side elapsed time, whether any HTTP status/headers were received, and the corresponding successful request immediately before/after it. Since the original failure produced no response headers, a client request ID is especially important for attempting server-side correlation.
One additional comparison would be useful without changing the workspace configuration: run the identical query against the identical workspace from the same client several times, then repeat it from the Azure-hosted managed-identity path you've already tested. If the Azure-hosted requests remain consistently fast while only the original client path intermittently stalls before headers, that provides useful evidence for separating a service/query issue from the client/network path. It still wouldn't by itself prove where the delay occurs.
Azure Monitor also imposes query concurrency and rate limits, but those normally produce defined HTTP responses such as 429 when the applicable limits are exceeded. The documented Analytics-table limit is five concurrent queries per user, with queued queries terminated after three minutes, which doesn't match a client-side 20-second timeout with no HTTP response.
So at this stage, preserve the current configuration and focus on obtaining a correlatable ID from an actual timed-out request rather than changing permissions, retention, or the query.
References:
Azure Monitor Logs - Query timeouts and errors
Azure Monitor Logs - Prefer options and query statistics
Help make this community better for everyone: If this answer helped or resolved your issue, please accept it or upvote it. If not, share more details in a comment so we can continue the discussion and find the right solution. Thank you.