An Azure service that provides curated open data for machine learning workflows.
The Azure Open Datasets “Public holidays” dataset already includes worldwide public holiday data (1970–2099) for 38 countries or regions. The list of countries in the question (Australia, Austria, Belgium, Brazil, Canada, China, Denmark, France, Germany, Greece, India, Ireland, Italy, Japan, South Korea, Luxembourg, Malaysia, Malta, Mexico, Netherlands, New Zealand, Poland, Portugal, Russia, South Africa, Spain, Sweden, Switzerland, Thailand, United Kingdom, United States) is covered as part of those 38 countries/regions.
To retrieve holidays for those specific countries, filter the dataset by the countryRegionCode or countryOrRegion columns.
Example using azureml-opendatasets (Python) for a date range and then filtering:
from azureml.opendatasets import PublicHolidays
from datetime import datetime
from dateutil.relativedelta import relativedelta
end_date = datetime.today()
start_date = datetime.today() - relativedelta(years=1)
hol = PublicHolidays(start_date=start_date, end_date=end_date)
hol_df = hol.to_spark_dataframe()
# Filter for selected countries by ISO country codes
selected_codes = [
'AU','AT','BE','BR','CA','CN','DK','FR','DE','GR','IN','IE','IT','JP','KR',
'LU','MY','MT','MX','NL','NZ','PL','PT','RU','ZA','ES','SE','CH','TH','GB','US'
]
hol_selected = hol_df.filter(hol_df.countryRegionCode.isin(selected_codes))
Example using PySpark directly on the Parquet dataset in Azure Open Datasets:
# Azure storage access info
blob_account_name = "azureopendatastorage"
blob_container_name = "holidaydatacontainer"
blob_relative_path = "Processed"
blob_sas_token = r"" # empty for public access
# Configure Spark to read from the public blob
wasbs_path = 'wasbs://%s@%s.blob.core.windows.net/%s' % (
blob_container_name, blob_account_name, blob_relative_path
)
spark.conf.set(
'fs.azure.sas.%s.%s.blob.core.windows.net' % (blob_container_name, blob_account_name),
blob_sas_token
)
# Load the Parquet data
df = spark.read.parquet(wasbs_path)
# Filter for the same set of countries
selected_codes = [
'AU','AT','BE','BR','CA','CN','DK','FR','DE','GR','IN','IE','IT','JP','KR',
'LU','MY','MT','MX','NL','NZ','PL','PT','RU','ZA','ES','SE','CH','TH','GB','US'
]
holidays_selected = df.filter(df.countryRegionCode.isin(selected_codes))
Key columns to use:
-
countryOrRegion: full country/region name (for example, “Australia”, “Germany”). -
countryRegionCode: ISO-like country/region code (for example,AU,DE). -
date: holiday date (timestamp). -
holidayName: full holiday name. -
normalizeHolidayName: normalized holiday name. -
isPaidTimeOff: available for US, GB, and India; indicates if most people have paid time off.
The dataset covers holidays from 1970-01-01 to 2099-01-01 and is stored in Parquet format in the East US region.
References: