Compare the similarity of 2 given addresses

Jonathan 380 Reputation points
2026-09-29T03:59:32.17+00:00

Hi,

The way below is good to compare the similarity of 2 given addresses? We expect to achieve that while 80% of the words (in the different position) are appearing well within both strings. Possible to achieve that? Or any other better to do that?

User's image

Developer technologies | C#
Developer technologies | C#

An object-oriented and type-safe programming language that has its roots in the C family of languages and includes support for component-oriented programming.

0 comments No comments

2 answers

Sort by: Most helpful
  1. Brian Pham (WICLOUD CORPORATION) 85 Reputation points Microsoft External Staff Moderator
    2026-09-29T05:38:44.76+00:00

    Hi @Jonathan ,

    The Levenshtein Distance implementation you shared is a valid approach and works well for finding small spelling differences, typos, or character-level changes between two strings.

    However, for address comparison it has some limitations. Since it compares the entire address character by character, it can produce a lower similarity score when the same address components appear in a different order

    To better align with the requirement that approximately 80% of the address words should match regardless of their position, I implemented a hybrid approach:

    public static double WordSimilarity(string first, string second)
            {
                MatchCollection firstWords = Regex.Matches(first ?? "", @"[\p{L}\p{N}]+");
                MatchCollection secondWords = Regex.Matches(second ?? "", @"[\p{L}\p{N}]+");
     
                if (firstWords.Count == 0 || secondWords.Count == 0)
                    return 0;
     
                var firstMatched = new bool[firstWords.Count];
                var secondMatched = new bool[secondWords.Count];
                int matches = 0;
     
                for (int secondIndex = 0; secondIndex < secondWords.Count; secondIndex++)
                {
                    for (int firstIndex = 0; firstIndex < firstWords.Count; firstIndex++)
                    {
                        if (!firstMatched[firstIndex] &&
                            string.Equals(firstWords[firstIndex].Value,
                                secondWords[secondIndex].Value,
                                StringComparison.OrdinalIgnoreCase))
                        {
                            firstMatched[firstIndex] = true;
                            secondMatched[secondIndex] = true;
                            matches++;
                            break;
                        }
                    }
                }
     
                for (int secondIndex = 0; secondIndex < secondWords.Count; secondIndex++)
                {
                    if (secondMatched[secondIndex])
                        continue;
     
                    string secondWord = secondWords[secondIndex].Value;
                    if (secondWord.Length < 4 || ContainsDigit(secondWord))
                        continue;
     
                    for (int firstIndex = 0; firstIndex < firstWords.Count; firstIndex++)
                    {
                        string firstWord = firstWords[firstIndex].Value;
                        if (!firstMatched[firstIndex] && firstWord.Length >= 4 &&
                            !ContainsDigit(firstWord) && WithinOneEdit(firstWord, secondWord))
                        {
                            firstMatched[firstIndex] = true;
                            matches++;
                            break;
                        }
                    }
                }
     
                return (double)matches / Math.Max(firstWords.Count, secondWords.Count);
            }
     
            public static bool IsMatch(string first, string second)
            {
                return WordSimilarity(first, second) >= 0.8;
            }
     
            private static bool ContainsDigit(string word)
            {
                foreach (char character in word)
                {
                    if (char.IsDigit(character))
                        return true;
                }
                return false;
            }
     
            private static bool WithinOneEdit(string first, string second)
            {
                if (Math.Abs(first.Length - second.Length) > 1)
                    return false;
     
                int firstIndex = 0;
                int secondIndex = 0;
                int edits = 0;
     
                while (firstIndex < first.Length && secondIndex < second.Length)
                {
                    if (char.ToUpperInvariant(first[firstIndex]) == char.ToUpperInvariant(second[secondIndex]))
                    {
                        firstIndex++;
                        secondIndex++;
                        continue;
                    }
     
                    if (++edits > 1)
                        return false;
     
                    if (first.Length >= second.Length)
                        firstIndex++;
                    if (second.Length >= first.Length)
                        secondIndex++;
                }
     
                return edits + (first.Length - firstIndex) + (second.Length - secondIndex) <= 1;
            }
        }
    

    Here's what the implementation does:

    1. Split both addresses into individual words and numbers -The regular expression extracts each address component separately instead of treating the address as one long string.
    2. Perform exact word matching first -The code compares each word from one address against the other. -Matching is case-insensitive. -The position of the words does not matter, so addresses containing the same components in different orders can still achieve a high score.
    3. Perform a second pass for unmatched words -If a word was not matched exactly, the code checks whether the words differ by only a single edit using the WithinOneEdit() method.
    4. Avoid fuzzy matching on short or numeric values

    Is it fine even if the words are in the different order?

    For your question about does the different order, my code is still work fine if the order of the address is mixed

    Compared to applying Levenshtein Distance to the entire address string, I believe this is more suitable for address matching scenarios. If you found my response helpful or informative, I would greatly appreciate it if you could follow this guide for your confirmation.

    Thank you.

    Was this answer helpful?


  2. Senthil kumar 2,500 Reputation points
    2026-09-29T05:20:14.99+00:00

    Hi @Jonathan

    Please try the below method will achieve the your result.

    public static double WordMatchPercentage(string a, string b)
    {
     var words1 = a.ToLower().Split(' ', ',', '.')
     .Where(x => !string.IsNullOrWhiteSpace(x))
     .ToHashSet();
     
     var words2 = b.ToLower().Split(' ', ',', '.')
     .Where(x => !string.IsNullOrWhiteSpace(x))
     .ToHashSet();
     
     int matched = words1.Intersect(words2).Count();
     
     int maxWords = Math.Max(words1.Count, words2.Count);
     
     return (double)matched / maxWords * 100;
    }
    

    Thanks.

    Was this answer helpful?


Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.