An Azure real-time data ingestion service.
I think the ASA-side fix is to repartition both streams by EntityId with the same partition count before the JOIN, rather than trying to align partition mapping across producers.
This browser is no longer supported.
Upgrade to Microsoft Edge to take advantage of the latest features, security updates, and technical support.
ASA job (compat 1.2) joins two Event Hubs on EntityId (both 4 partitions) into a third hub. With the input Partition key column declared (embarrassingly-parallel), the JOIN silently loses ~78% of matches — no errors. Measured: parallel ≈ 22% match, non-parallel ≈ 100%.
Cause: the two inputs are produced by clients that hash the key differently, so the same EntityId lands on different partitions: EventHubBufferedProducerClient uses client-side lookup3, EventHubProducerClient/gateway uses another hash, and ASA output uses a third.
Q: Can I force all publishers (buffered producer, regular producer, and ASA output) to use the same partition-key hashing function, so the same EntityId always lands on the same partition and a parallel JOIN matches?
An Azure real-time data ingestion service.
I think the ASA-side fix is to repartition both streams by EntityId with the same partition count before the JOIN, rather than trying to align partition mapping across producers.
Hi @Andrei Josephsen ,
Since the partition key hash algorithm is an internal implementation detail and cannot currently be configured, there is no supported method to enforce Kafka Murmur2 partition mapping within the service.
The following mitigation options may be considered:
Limitation
At present, the partition key hash algorithm is not exposed as a customer-configurable setting. Therefore, different services may map the same partition key value to different physical partitions, and this behavior is expected.
Recommended Resolution
If deterministic partition mapping across Kafka and Azure services is a business requirement, republishing or implementing a custom routing layer remains the most practical workaround. However, this approach introduces additional latency and operational cost. A product enhancement request for configurable hashing algorithms would be the preferred long-term solution.
Thanks & Regards,
Sivasankar.