Azure Managed Grafana: Scheduled PDF report generation fails intermittently and silently (no error, no email)

Regina Michel 0 Zuverlässigkeitspunkte
2026-08-17T10:25:16.5833333+00:00

We use Azure Managed Grafana (major version 12) for periodic, scheduled PDF reports (Grafana Reporting feature) that are emailed automatically to our end users.

Report generation fails intermittently and non-deterministically. The exact same report definition (same dashboard, same panel count, same schedule) sometimes succeeds and sometimes fails, with no error surfaced anywhere — no error email, nothing in the Grafana UI, and no way to retrieve report generation status/history via the API. The only symptom is that the expected email simply doesn't arrive. We currently work around this by manually checking every week whether all expected reports were sent, and re-triggering failed ones by hand — sometimes needing up to 5 attempts before a report goes out successfully.

We are aware of the documented global 200-second hard limit for report generation and the ~10s average per-panel rendering time, and that this isn't configurable on Azure Managed Grafana. That explains why reports with more panels are more likely to fail, but not why identical, repeated runs of the same configuration produce different outcomes, nor why failures are completely silent with no diagnostic information.

I've filed the detailed analysis (including empirical success-rate measurements across different scheduling strategies) as a GitHub issue in the Grafana repository: https://github.com/grafana/grafana/issues/130837

Is there any known guidance on making scheduled reporting more reliable on Azure Managed Grafana?

Thanks a lot for your help.

Von Azure verwaltetes Grafana
Von Azure verwaltetes Grafana

Ein Azure-Dienst, der zum Bereitstellen von Grafana-Dashboards für Analyse- und Überwachungslösungen verwendet wird

0 Kommentare Keine Kommentare

1 Antwort

Sortieren nach: Am hilfreichsten
  1. Suchitra Suregaunkar 16,780 Zuverlässigkeitspunkte Externe Microsoft-Mitarbeiter Moderator
    2026-08-27T07:17:07.63+00:00

    Hello Regina Michel

    Thank you for posting your query on Microsoft Q&A platform.

    You've already found the 200-second cap and the ~10s average per-panel render time. The key detail is that the ~10s is an average, not a fixed cost — the documented figure assumes the data query for that panel completes in under one second (Use reporting and image rendering in Azure Managed Grafana). In practice, per-panel time = browser render + query time, and query latency against Azure Monitor, Log Analytics or ADX varies from run to run depending on backend load, caching and throttling.

    That's why the boundary isn't a clean panel-count threshold. A report sitting at 12–13 panels lands close to 200 seconds, so small variations in query latency decide whether it finishes or gets cut off. Your data (9 panels always fine, 14–15 always failing, 10–13 mixed) matches that behaviour exactly.

    Your top-of-the-hour result is the other half of the picture. 52.9% when everything fires at :00 versus 75% when staggered five minutes apart is a strong indication of contention on the image renderer, which has fixed capacity per workspace. The same doc notes that overloading the renderer can make it unstable, and that alert screenshots draw on the same renderer.

    One more failure mode worth ruling out: Grafana does not send attachments larger than 10 MB by default, to stop mail servers rejecting them (Create and manage reports). This limit isn't customer-configurable on Azure Managed Grafana. If the PDFs that do arrive are anywhere near 10 MB, that could be causing some of your silent drops independently of the timeout — worth a quick check.

    I want to be straight with you here rather than point you at a document that doesn't exist: there is no Microsoft article that resolves the missing error notification or the missing report run history. That behaviour sits in the Grafana reporting feature itself, which is why your GitHub issue is the correct place for it.

    It's also not something you can work around by digging into logs on Azure Managed Grafana. Diagnostic settings for the Microsoft.Dashboard/grafana resource type currently support Grafana login events only — audit logs and metrics are documented as not supported.

    Reference: Monitor Azure Managed Grafana using diagnostic settings.

    And renderer internals aren't reachable because the Grafana Server Admin role isn't available to customers and the Admin API is disabled (Service limits, quotas, and constraints).

    So the realistic approach is to keep reports comfortably inside the budget rather than near the edge, and to detect misses outside Grafana.

    Aim for around 8–10 panels per report, not 20. The 20-panel figure in the docs is an upper bound, and it drops further if you attach CSV. Targeting roughly half the time budget gives you headroom for query variance. Splitting one large dashboard into two or three smaller report dashboards is more effective than trimming a borderline one.

    Cut query time, not just panel count. Shorten the report time range, pre-aggregate where you can, and use table format instead of time series for heavy panels — the limits doc specifically recommends this for Azure Data Explorer, along with avoiding many panels querying the same cluster at once. Your sweep showed the time series panels were driving the failures, which fits.

    Widen the stagger. Five minutes already got you to 75%. Since a single report can legitimately hold the renderer for up to 200 seconds, 10–15 minutes between reports is safer, and moving them off the top of the hour avoids competing with dashboard refreshes and alert screenshots. If you have alert rules with screenshots enabled, limiting those to the ones that truly need them frees up renderer capacity too.

    Drop CSV attachments from reports that don't strictly need them.

    Replace the weekly manual check with automated detection. Since run history isn't available through the API, verify at the delivery end instead, add a monitored mailbox as an extra recipient, then use a scheduled Logic App to check whether each expected report arrived in its expected window, with an Action Group alert when one doesn't. It won't retry for you, but it turns a weekly manual sweep into a notification within minutes.

    Hope this helps clarify things.

    Best regards,

    Suchitra

    War diese Antwort hilfreich?

    0 Kommentare Keine Kommentare

Ihre Antwort

Antworten können von Fragestellenden als „Angenommen“ und von Moderierenden als „Empfohlen“ gekennzeichnet werden, wodurch Benutzende wissen, dass diese Antwort das Problem des Fragestellenden gelöst hat.