Azure SQL private endpoint: Linux App Service + managed identity passes DNS/TCP/token checks, but SqlClient OpenAsync fails with ConnectionReset 10054

Pravin Goorun 21 Reputation points
2026-10-11T09:46:14.5133333+00:00

QUESTION

I am investigating a Development-only Azure SQL connection problem from a .NET application hosted on Linux Azure App Service. I would appreciate help identifying the next targeted diagnostic, rather than repeating deployment attempts or weakening the security configuration.

The latest verified result is that private DNS resolution, TCP connectivity to port 1433, and managed-identity token acquisition all succeed, but opening the SQL session fails with SqlException Number 0, State 0 and SocketError ConnectionReset (10054).

The latest live evidence below was collected on 8 October 2026. No real customer or financial data has been used. Resource names and identifiers have been deliberately omitted from this public question.

  1. ENVIRONMENT AND INTENDED CONNECTION PATH
  • Linux Azure App Service, Basic B1 plan, .NET 10.
  • The diagnostic worker is a temporary Web App sharing the existing B1 plan, with its own system-assigned managed identity.
  • The worker is integrated with the existing application VNet integration subnet. It runs as an App Service startup process and attempts SQL access from the application runtime, not from my laptop.
  • Azure SQL Database, General Purpose serverless, using the free offer with paid-overage protection configured to AutoPause.
  • SQL public-network access is Disabled; Microsoft Entra-only authentication is enabled; minimum TLS version is 1.2.
  • An approved private endpoint targets the intended SQL logical server.
  • The private endpoint NIC's private IP matches the private DNS A record, and the private DNS zone is linked to the integrated VNet.
  • No NSG, user-defined route table or NAT gateway was found on either the integration subnet or the private-endpoint subnet.
  • The SQL connection uses the normal logical-server hostname, <server>.database.windows.net, not the private-link alias or an IP address.
  • The database name is explicitly supplied.
  • Microsoft.Data.SqlClient version 6.1.6.
  • Encryption and certificate validation remain enabled. No SQL username/password is used.
  • The original SQL connection policy is Default. Explicit Proxy was tested temporarily and then restored to Default.
  1. WHAT SUCCEEDS BEFORE SQL OPEN

From the diagnostic application process:

  • The logical-server hostname resolves to a private IPv4 address, consistent with the checked private endpoint/DNS configuration.
  • A bounded TCP connection to port 1433 succeeds.
  • ManagedIdentityCredential.GetTokenAsync succeeds for scope https://database.windows.net/.default.

The token is kept in memory only. These checks establish DNS, basic port reachability and token acquisition; they do NOT prove a successful SQL login, an authorised database principal, or a completed TLS/TDS negotiation.

  1. OBSERVED ERRORS AND THEIR ORDER

The following are recorded diagnostic fields, not verbatim full exception messages. Our diagnostic host deliberately suppresses raw exception text, connection values and tokens.

Latest repeated result, 8 October 2026:

Operation: opening SQL connection before schema work

SqlException.Number: 0

SqlException.State: 0

SocketError: ConnectionReset

Numeric socket error code: 10054

An earlier attempt on 8 October, after successful token acquisition, reported:

SqlException.Number: -2

SqlException.State: 0

Classification: connection-stage timeout

Historical results from 3 October, using earlier diagnostic packages:

SqlException.Number: 40613

Context: database availability/serverless wake-up attempt

SqlException.Number: 18456

Context: one earlier run reached a login-denial result before temporary

         administrator assignment, and still reported 18456 after assignment

         and restart. The verification of fresh post-restart process results

         was subsequently strengthened, so this is historical evidence,

         not a reliable explanation of the latest 10054 result.

Only 40613 receives a bounded connection-open retry (up to eight attempts, separated by 15 seconds). Schema operations are not retried by that helper. The latest reset results are not reported as successful migrations.

  1. ISOLATION TESTS ALREADY PERFORMED

A. Credential selection

Restricted the credential selection to ManagedIdentityCredential on a temporary worker. This did not resolve the failure. Token acquisition was subsequently verified independently as described above.

B. Explicit SQL Proxy and longer timeout

Tested explicit Proxy with a 120-second connection timeout. DNS/TCP and token checks passed, but SQL open still failed. Default policy was restored afterwards.

C. Private-only versus route-all App Service routing

On a disposable worker only, disabled route-all while retaining VNet integration for private SQL traffic. The modern outbound routing flags were also verified as false. With Proxy and timeout held constant, token acquisition and private DNS/TCP checks succeeded but SQL open still failed. The permanent application routing was not changed.

D. Direct token handoff to SqlClient

Passed the separately acquired managed-identity token directly to SqlConnection.AccessToken. Removed the Authentication keyword from that constructed connection because it cannot be combined with AccessToken. Disabled pooling for this one-shot diagnostic process. Encryption, hostname and certificate validation were retained. With original route-all and Default SQL policy, the result was still Number 0 / State 0 / ConnectionReset 10054.

E. Strict encryption

On a disposable worker, tested Encrypt=Strict while retaining certificate validation. The result was again 10054. I am not treating this as proof that TLS is correct or that all TLS-related causes are ruled out.

F. Independent client comparison -- INCONCLUSIVE

Attempted to run official Microsoft go-sqlcmd v1.10.0 via the temporary host's Kudu command endpoint, using managed identity, validated encryption and only SELECT 1. Kudu command execution returned exit 127 in an initial attempt and did not produce trustworthy readiness/SQL-stage evidence in subsequent bounded attempts. Therefore, an independent client has NOT been proven to reach SQL. This is not a cross-client reproduction of the SQL reset and does not rule out a SqlClient/runtime-specific problem.

  1. DIRECT-TOKEN TEST SHAPE

This is a simplified illustration of the tested connection pattern, not a standalone reproduction or a claim that this identity has database permissions:

var credential = new Azure.Identity.ManagedIdentityCredential();

var token = await credential.GetTokenAsync(

    new Azure.Core.TokenRequestContext(

        new[] { "https://database.windows.net/.default" }),

    cancellationToken);

var settings = new Microsoft.Data.SqlClient.SqlConnectionStringBuilder

{

    DataSource = "tcp:<server>.database.windows.net,1433",

    InitialCatalog = "<database>",

    Encrypt = Microsoft.Data.SqlClient.SqlConnectionEncryptOption.Mandatory,

    TrustServerCertificate = false,

    ConnectTimeout = 120,

    Pooling = false

};

// No Authentication keyword, User ID or Password on this connection.

await using var connection =

    new Microsoft.Data.SqlClient.SqlConnection(settings.ConnectionString)

    {

        AccessToken = token.Token

    };

await connection.OpenAsync(cancellationToken);
  1. IMPORTANT AUTHORISATION AND DIAGNOSTIC LIMITS
  • In the latest 10054 attempts, the disposable worker had not yet been granted SQL administrator authority. Our safety workflow stops before that elevation if the initial connection probe produces an unexpected failure. Database authorisation for that identity is therefore not proven.
  • I do not assume that successful token acquisition proves database access, or that a raw TCP success proves a working SQL session.
  • Earlier 18456 results show that login denial was observed in some runs, but the latest blocker is a reset/timeout rather than a specific SQL login error.
  • Recent connection counters did not supply a more specific error. A Resource Health query was unavailable; that is not evidence of an Azure outage.
  • All temporary diagnostic Web Apps were removed, the original administrator was restored/retained, SQL public access remains disabled, and the original connection policy was restored.
  • Application/packaging tests pass, but that does not establish successful Azure SQL connectivity. No successful schema migration or application publication is being claimed.
  1. TARGETED QUESTIONS
  1. Given successful private DNS, TCP 1433 and managed-identity token acquisition, what supported diagnostic best distinguishes TLS/TDS pre-login failure, SQL gateway/session routing failure and an Entra login/authorisation failure when SqlClient reports Number 0 with ConnectionReset 10054?
  2. In this private-endpoint configuration, could an identity without a mapped SQL principal produce this reset rather than 18456? What is the safest controlled way to distinguish transport failure from missing database authorisation without assuming either is already correct?
  3. Are there additional App Service VNet integration/private endpoint/SQL connection-policy checks that remain meaningful after both explicit Proxy and private-only routing have been tested? Please identify the specific effective setting or diagnostic, rather than recommending public SQL access.
  4. What is the supported way to run a second SQL client in the SAME App Service application-runtime network/identity context, given that the Kudu comparison was inconclusive? Please do not assume Kudu execution is equivalent to a completed runtime-side test.
  5. Is there a documented compatibility issue relevant to Linux App Service, .NET 10, Microsoft.Data.SqlClient 6.1.6 and Azure SQL private endpoints? If a version change is suggested, which documented issue would it test?
  6. Can existing Azure platform diagnostics establish whether the SQL gateway receives the attempt and where the reset occurs, without exporting tokens, complete connection strings or application/customer data? Which specific metric/log category or supported trace would help, and would it introduce charges?

I am looking for one evidence-led next diagnostic that preserves private SQL access, Entra-only authentication, validated encryption and least privilege. I can provide necessary resource identifiers only through an appropriately verified private Microsoft support channel, not in this public thread. I will review any proposed additional resource, role assignment or paid diagnostic before proceeding.

Thank you for helping narrow down the connection stage and the safest next step.

Azure SQL Database

1 answer

Sort by: Most helpful
  1. AI answer

    2026-10-11T09:49:00.0533333+00:00

    The most targeted next test is a single, time-correlated OpenAsync attempt from a minimally authorized managed identity while collecting SqlClient EventSource, an application-container packet capture, and Azure SQL connection metrics.

    1. Check the remaining private-endpoint prerequisite

    Before tracing, inspect the effective server-level firewall rules for the source address that the SQL gateway sees. Microsoft’s private-endpoint troubleshooting guidance states that Azure SQL can still enforce server-level firewall rules for private-endpoint traffic.^1^

    Do not assume that the private-endpoint subnet is the correct source range:

    1. Determine the effective source address from the packet capture and Azure configuration.
    2. If no matching rule exists, add the narrowest temporary rule for that address/range through SQL logical server → Networking → Firewall rules, or have the administrator use sp_set_firewall_rule.
    3. Repeat one bounded OpenAsync attempt.
    4. Remove the temporary rule afterward if it is not part of the intended design.

    This does not require enabling public access.

    2. Establish a controlled authorization baseline

    Have the Microsoft Entra administrator connect to the explicit target database and create a contained user for the disposable worker’s managed identity. Grant only CONNECT, not reader, writer, DDL, or administrator roles:

    USE [TargetDatabase];
    GO
    CREATE USER [DiagnosticWorkerIdentity] FROM EXTERNAL PROVIDER;
    GRANT CONNECT TO [DiagnosticWorkerIdentity];
    GO
    

    Use an unambiguous principal-creation method supported by your tenant if display names are duplicated. Azure SQL authorization is controlled through database users, roles, and explicit permissions, with least privilege recommended.^2^

    A missing database principal normally belongs in the authentication/authorization branch, but 10054 alone does not prove either authorization failure or transport failure. This controlled mapping removes that ambiguity without granting schema privileges.

    3. Capture both client layers during exactly one attempt

    SqlClient EventSource

    If the App Service image contains a compatible .NET SDK and permits diagnostic-tool installation and process attachment, attach at information level:

    dotnet tool install --global dotnet-trace
    export PATH="$PATH:$HOME/.dotnet/tools"
    dotnet-trace ps
    dotnet-trace collect --process-id <PID> \
      --providers Microsoft.Data.SqlClient.EventSource:1FFF:4
    

    If the image has only the runtime or blocks tool installation/attachment, package dotnet-trace with the diagnostic deployment or launch the diagnostic assembly under the documented collector form instead:

    dotnet-trace collect \
      --providers Microsoft.Data.SqlClient.EventSource:1FFF:5 \
      -- dotnet DiagnosticWorker.dll
    

    Verbose level 5 produces more sensitive and voluminous output; use it only for a short window and protect the trace. These are the documented SqlClient tracing forms.^3^

    Application-container packet capture

    For a built-in Linux App Service where SSH/root package installation is available:

    apt-get update && apt-get install -y tcpdump
    tcpdump -D
    tcpdump -i eth0 -nn -s 0 -w /home/sql-open.pcap \
      host <private-endpoint-ip> and port 1433
    

    Reproduce one connection, stop tcpdump, and download the capture from /home. This is the documented Linux App Service procedure.^4^

    The capture can show:

    • which peer transmitted the TCP RST;
    • whether the reset occurs before or after visible TDS pre-login/TLS exchange;
    • retransmissions or timeout behavior.

    It cannot generally identify which internal Azure gateway component generated the reset or why; that requires Microsoft Support correlation. The SQL Server 10054 TLS article also recommends examining Client Hello and Server Hello in a client/server trace, although it applies specifically to SQL Server rather than establishing an Azure SQL root cause here.^5^

    4. Correlate with Azure SQL metrics

    Record UTC start/end timestamps and inspect SQL database → Monitoring → Metrics using one-minute granularity:

    • Successful Connections
    • Failed Connections: System Errors
    • Failed Connections: User Errors

    Split failures by Error and ValidatedDriverNameAndVersion; successful connections can also be split by SslProtocol.^6^

    Reading metrics requires Microsoft.Insights/metrics/read, commonly supplied through Monitoring Reader, scoped as narrowly as practical. Publication delay, one-minute aggregation, and concurrent attempts prevent guaranteed one-to-one attribution; absence of a metric is supporting evidence, not proof that the gateway received nothing.

    Standard platform metrics have no charge. Sending audit or diagnostic logs to Log Analytics can incur ingestion and retention charges.^7^

    Do not rely on SQL auditing to exclude Entra authentication failure: failed Microsoft Entra logins do not appear in Azure SQL audit logs because credential verification occurs before the requested database is accessed.^8^

    Independent client

    Deploy sqlcmd (Go) inside the same diagnostic app/container image and invoke it from the application startup process—not Kudu—with the same system-assigned identity:

    sqlcmd \
      -S <server>.database.windows.net \
      -d <database> \
      --authentication-method ActiveDirectoryManagedIdentity \
      -Q "SELECT 1"
    

    That managed-identity mode is documented for system-assigned identities.^9^

    No published issue in the supplied material ties this exact combination to 10054. Version 6.1.6 is supported, but 6.1.7 is the latest 6.1 LTS patch; test it only as a controlled comparison, not as a presumed fix.^10^

    If the capture shows a gateway-originated RST and metrics/traces do not expose the reason, open a private support request with UTC timestamps, database/server resource IDs, .pcap, .nettrace, driver version, and any tracing ID—never the token or complete connection string.


    References

    1. Troubleshoot connectivity issues and other errors
    2. Authorize database access to Azure SQL Database
    3. Enable event tracing in SqlClient
    4. Troubleshoot virtual network integration with Azure App Service
    5. An existing connection was forcibly closed by the remote host (OS error 10054)
    6. Monitor Azure SQL Database with metrics and alerts
    7. Azure Monitor cost and usage
    8. Auditing for Azure SQL Database
    9. Authenticate with Microsoft Entra ID in sqlcmd
    10. SqlClient Driver Support Lifecycle - ADO.NET Provider for SQL Server
    AI-generated content may be incorrect. Read our transparency notes for more information.

    Was this answer helpful?

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.