Container Apps internal service-to-service call times out (TCP connect timeout) despite healthy target container and correct DNS resolution

JochenLeppert-3710 0 Zuverlässigkeitspunkte
2026-09-22T14:46:00.24+00:00

Problem description

Internal HTTP calls between Container Apps in the same Managed Environment, using the short internal service name (http://appname:port), fail with a raw TCP connect() timeout - not an HTTP-level timeout. This happens even though:DNS resolves the internal name to a valid internal IP.The target container is confirmed healthy (HealthState: Healthy), stable, and has been running for hours without restarts.The exact same call pattern to a different, freshly created "hello world" test app in the same environment succeeds within ~1 second.Ingress configuration (targetPort aside) is identical between the working and failing app: external=false, transport=Auto, no IP restrictions, no client cert requirement, no custom domains.The target app's own log stream shows zero new log entries corresponding to the failed call's exact timestamp window - no incoming request is logged at all.

Environment

Azure Container Apps, Consumption workload profile, Single revision mode, Dapr disabledVNet-integrated environment (infrastructureSubnetId set to a delegated subnet), no custom domains bound, no peering involved in the current testRegion: France Central

What we've already tried

Deleted and fully recreated the entire Managed Environment and all application Container Apps from scratch, onto the same subnet as before.Verified DNS resolution succeeds to a correct internal IP for the failing hostname.Confirmed calling localhost inside the same container works fine (rules out app-level binding issues).Deployed a brand-new isolated environment on a brand-new VNet/subnet in a separate resource group - internal calls worked correctly there.After moving production onto a new VNet/subnet (matching the working isolated test): the exact same timeout reappeared immediately.Removed VNet peering entirely: no change.Restarted all application revisions: no change.Ran a clean paired test with exact UTC timestamps: a call to a fresh test app succeeded in ~10s with a real HTTP response; immediately afterward, the failing call to an existing app timed out with a Python socket.connect() TimeoutError (the TCP three-way handshake itself never completes - this is not an application-layer issue).

Question

Is there a known issue or diagnostic path for internal service-to-service TCP connections silently failing at the handshake level within a Container Apps Managed Environment, specifically for some apps but not others in the same environment, that isn't explained by DNS, ingress configuration, VNet/subnet, or app health state? We have an open Azure support case (Developer plan) but have been asked to also post here for continued help. Happy to share more diagnostic output if useful.

Azure
Azure

Eine Cloud Computing-Plattform und -Infrastruktur für die Erstellung, Bereitstellung und Verwaltung von Anwendungen und Diensten über ein weltweites Netzwerk aus Rechenzentren, die von Microsoft verwaltetet werden.


Ihre Antwort

Antworten können von Fragestellenden als „Angenommen“ und von Moderierenden als „Empfohlen“ gekennzeichnet werden, wodurch Benutzende wissen, dass diese Antwort das Problem des Fragestellenden gelöst hat.