An Azure service for ingesting, preparing, and transforming data at scale.
Hello @Kavika Roy
Yes, a watermark-based incremental pipeline is the standard Azure Data Factory pattern when the source exposes a monotonically increasing last-modified value. Using a watermark column such as last_modify_time to retrieve only records created or updated since the previous successful load.
For your architecture, use Odoo REST/API → ADF REST connector → ADLS Gen2 with a small control store containing the last successfully committed watermark.
A typical run would:
- Read the previous watermark.
- Determine an upper watermark for the current run.
- Request records where the source modification timestamp is greater than the old watermark and less than or equal to the new watermark.
- Follow the API's pagination until you retrieve all pages.
- Write the incremental batch to a new ADLS path/file, preferably partitioned by ingestion date/run ID.
- Update the stored watermark only after the ingestion completes successfully.
This follows Microsoft's documented watermark pattern, where ADF reads the old and new watermark values, copies the delta, and updates the watermark after the copy.
For pagination, ADF's REST connector supports next URLs, query parameters, headers, JSON response values, end conditions, and maximum request counts. The exact rule depends on how the Odoo API you're using exposes pagination.
For failed runs, don't advance the watermark until the complete incremental window succeeds. A retry can then reprocess the same window. But note that for scenarios other than supported binary file copies, a rerun can restart the copy rather than resume at an individual API-record boundary.
To avoid duplicates, keep the raw ADLS ingestion immutable and include the pipeline run/window in the path. Perform deduplication/upsert downstream using the Odoo record ID plus modification timestamp rather than expecting ADLS itself to provide record-level upsert semantics. Derived from standard data-lake architecture.
Deletes require separate treatment. A timestamp watermark detects records returned by the source as new or modified; it does not inherently identify records that have disappeared from the source. I cannot find Microsoft documentation defining deletion/change-tracking semantics for Odoo. You would need to determine whether the specific Odoo API exposes deleted/archived records, a deletion indicator, or another mechanism you can reconcile against your data lake. Do not assume write_date solves deletion detection.
One additional safeguard is to use a small overlap in the extraction window and deduplicate downstream. That can protect against timestamp-boundary issues, but this is an architectural recommendation rather than an ADF requirement documented by Microsoft.
References:
Incrementally copy data using Azure Data Factory
Incremental copy using a watermark
ADF REST connector and pagination support
Help make this community better for everyone: If this answer helped or resolved your issue, please accept it or upvote it. If not, share more details in a comment so we can continue the discussion and find the right solution. Thank you.