Persistent packet loss between Japan ISP (BIGLOBE, 60.237.255.0/24) and Azure Japan East backbone (104.44.0.0/19)

takayoshi maeda 0 評価のポイント
2026-05-20T06:32:10.66+00:00

Environment

  • Source: Japan / BIGLOBE (KDDI group) FTTH, IP range 60.237.255.0/24 (AS2518)
  • Destination: Azure Database for PostgreSQL Flexible Server
    • Region: Japan East
    • Endpoint IP: 20.37.121.241 (Azure backbone 104.44.0.0/19, AS8075)
    • Ports: 5432 (direct) + 6432 (built-in PgBouncer, transaction mode)

Symptoms (since 2026-05-19)

Sustained congested-peering packet loss to the above endpoint.

mtr 100-packet result under load

  • Hop 4 (104.44.44.175): 94% packet loss, 7200 ms RTT
  • Hops 5+: minor

Connection failure pattern

  • HikariCP / Agroal pool init: PSQLException: 接続試行は失敗しました storms
  • :5432 (direct) and :6432 (built-in PgBouncer) show comparable loss rates
  • Failures during TLS handshake / DNS / SYN-ACK phase

Routing isolation

From the same BIGLOBE IP:

  • AWS Tokyo endpoints: 0% loss
  • Akamai CDN (Tokyo edge): 0% loss
  • Azure Japan East (skillbridgedb): high loss only here

Additional cross-region test (2026-05-20):

  • u-note1 → Azure JP West (20.78.169.17 via B1ms test VM): also routes through 104.44.44.175 at hop 4, 72% loss
  • Confirms Microsoft AS8075 edge 104.44.44.175 is the shared bottleneck for all Japan-region routing from BIGLOBE

→ Specific Microsoft Azure Japan-facing peering segment is the bottleneck.

Impact

  • ~15,000 pool init failures / minute observed on production batch servers
  • Production batch jobs in churn / retry storms

Already applied (client-side mitigations)

  • Migrated to Azure built-in PgBouncer (transaction mode, default_pool_size=20)
  • JDBC URL params: connectTimeout=10&socketTimeout=300&tcpKeepAlive=true
  • HikariCP retry: 5s → 10s → 15s linear backoff (3 attempts)
  • Tested JP West bastion: same Microsoft entry hop 104.44.44.175 → no improvement (cross-region bastion ruled out)

Correction note

The initial version of this post incorrectly identified the source ISP as So-net (AS2527). Confirmed via whois on

60.237.0.0/16 that the actual ISP is BIGLOBE (AS2518, KDDI group). All routing observations remain the same; only the

ISP attribution is corrected.

Requested action

  1. Verify congestion on the peering segment BIGLOBE (AS2518, KDDI group) ↔ Microsoft (AS8075), specifically the hop

via 104.44.44.175

  1. Capacity expansion or alternative transit routing in Japan East
  2. Confirm whether NTT Communications transit (133.205.138.0/24 → 104.44.44.175) and BIGLOBE direct peering

(60.237.255.0/24 → 104.44.44.175) are both saturated

Thank you in advance.

Azure Database for PostgreSQL
Azure Database for PostgreSQL

アプリの開発およびデプロイ用の Azure マネージド PostgreSQL データベース サービス。


1 件の回答

並べ替え方法: 最も役に立つ
  1. Pilladi Padma Sai Manisha 11,715 評価のポイント Microsoft 外部スタッフ モデレーター
    2026-05-21T10:09:07.2366667+00:00

    Hey takayoshi maeda,

    It definitely looks like the packet loss is happening right at the Microsoft public-peering edge in Japan East (AS8075). Since your AWS and Akamai tests show 0% loss, and all your So-net/NTT traffic to Azure hits that same hop with 70–94% loss, it’s almost certainly congestion on Microsoft’s public-peering link.

    Here’s what you can do:

    1. Gather detailed peering metrics
      • In Azure Portal go to Network Watcher > Metrics, select your peering connection, and chart “Next Hop Loss” over time.
      • Or run the PowerShell reachability report:
        
             Get-AzNetworkWatcherReachabilityReport -Location "Japan East"
        
        
      This will confirm packet drops and help pinpoint peak congestion windows.
    2. Share your data with the Microsoft Peering team
      • Email [email protected] with your ASN (2527) and the traceroute/MTR results.
      • Ask them to verify if that public-peering link is at capacity, and to schedule an expansion or shift your peering to an alternate POP in Japan East.
    3. Consider alternative Microsoft entry points
      • Azure Peering Service automatically picks the lowest-latency, lowest-loss Microsoft POP for you. Rolling this out can bypass the congested AS8075 link.
      • If you’re on ExpressRoute, add Microsoft peering for your prefixes—this uses dedicated circuits and avoids the public-peering bottleneck.
    4. Short-term mitigations
      • Switch your PostgreSQL Flexible Server to a Private Endpoint. Traffic will ride over the Azure backbone (ExpressRoute/Microsoft backbone), not the public-peering link.
      • If changing ISPs is an option, test another transit provider in Japan to see if your path diverges before that congested hop.

    Hope this helps you nail down the root cause and get that peering segment beefed up soon!

    References:

    この回答は役に立ちましたか?

    0 件のコメント コメントはありません

お客様の回答

質問作成者は回答に "承認済み"、モデレーターは "推奨" とマークできます。これにより、ユーザーは作成者の問題が回答によって解決したことを把握できます。