How can I implement incremental API ingestion from Odoo ERP to Azure Data Lake Storage Gen2 using Azure Data Factory?

Kavika Roy 0 Reputation points
2026-09-25T14:43:28.43+00:00

I'm working for a project on an Azure data platform where Odoo ERP is the source and ADLS Gen2 is used as the data lake. The Odoo API contains manufacturing, inventory, procurement, and financial data that needs to be refreshed regularly for Power BI reporting.

Since Odoo is API-based and doesn't provide traditional database CDC, I'm trying to determine the best way to implement incremental ingestion in Azure Data Factory.

Should I use a watermark based on the record's write_date/last-modified timestamp, and how should I handle pagination, updated records, deleted records, and failed pipeline runs without creating duplicates in ADLS Gen2?

Azure Data Factory
Azure Data Factory

An Azure service for ingesting, preparing, and transforming data at scale.

0 comments No comments

1 answer

Sort by: Most helpful
  1. Allan Solomon Mejia 10,225 Reputation points
    2026-09-25T15:01:47.1033333+00:00

    Hello @Kavika Roy

    Yes, a watermark-based incremental pipeline is the standard Azure Data Factory pattern when the source exposes a monotonically increasing last-modified value. Using a watermark column such as last_modify_time to retrieve only records created or updated since the previous successful load.

    For your architecture, use Odoo REST/API → ADF REST connector → ADLS Gen2 with a small control store containing the last successfully committed watermark.

    A typical run would:

    1. Read the previous watermark.
    2. Determine an upper watermark for the current run.
    3. Request records where the source modification timestamp is greater than the old watermark and less than or equal to the new watermark.
    4. Follow the API's pagination until you retrieve all pages.
    5. Write the incremental batch to a new ADLS path/file, preferably partitioned by ingestion date/run ID.
    6. Update the stored watermark only after the ingestion completes successfully.

    This follows Microsoft's documented watermark pattern, where ADF reads the old and new watermark values, copies the delta, and updates the watermark after the copy.

    For pagination, ADF's REST connector supports next URLs, query parameters, headers, JSON response values, end conditions, and maximum request counts. The exact rule depends on how the Odoo API you're using exposes pagination.

    For failed runs, don't advance the watermark until the complete incremental window succeeds. A retry can then reprocess the same window. But note that for scenarios other than supported binary file copies, a rerun can restart the copy rather than resume at an individual API-record boundary.

    To avoid duplicates, keep the raw ADLS ingestion immutable and include the pipeline run/window in the path. Perform deduplication/upsert downstream using the Odoo record ID plus modification timestamp rather than expecting ADLS itself to provide record-level upsert semantics. Derived from standard data-lake architecture.

    Deletes require separate treatment. A timestamp watermark detects records returned by the source as new or modified; it does not inherently identify records that have disappeared from the source. I cannot find Microsoft documentation defining deletion/change-tracking semantics for Odoo. You would need to determine whether the specific Odoo API exposes deleted/archived records, a deletion indicator, or another mechanism you can reconcile against your data lake. Do not assume write_date solves deletion detection.

    One additional safeguard is to use a small overlap in the extraction window and deduplicate downstream. That can protect against timestamp-boundary issues, but this is an architectural recommendation rather than an ADF requirement documented by Microsoft.

    References:

    Incrementally copy data using Azure Data Factory

    Incremental copy using a watermark

    ADF REST connector and pagination support

    Copy Activity overview


    Help make this community better for everyone: If this answer helped or resolved your issue, please accept it or upvote it. If not, share more details in a comment so we can continue the discussion and find the right solution. Thank you.

    Was this answer helpful?

    0 comments No comments

Your answer

Answers can be marked as 'Accepted' by the question author and 'Recommended' by moderators, which helps users know the answer solved the author's problem.