Issue
It is possible to occasionally experience large spikes of duplicate events processed twice by Segment Reverse ETL, often arriving about a month apart. These duplicates appear in downstream tools and Unify with brand new messageIds, preventing standard deduplication from recognizing them as the same event.
Product
Segment Reverse ETL
Environment
Segment Console
Cause
To prevent infrastructure overload, Segment intentionally applies a randomized start-time variation (jitter) of up to 15 minutes to every scheduled Reverse ETL sync. For example, a sync scheduled for 8:00 AM may start anywhere between 8:00 AM and 8:15 AM.
If your data warehouse transformation jobs (such as dbt runs) are scheduled too close to your Reverse ETL syncs, this 15-minute jitter can cause the Reverse ETL extraction to occasionally overlap with the exact moment your data warehouse tables are dropping, rebuilding, or temporarily empty. When Reverse ETL reads an empty or partial view, it clears those records from its internal state tracking table. During the next successful sync, Reverse ETL treats the returning rows as brand-new "Added Records", generates new messageIds starting with retl-, and sends the duplicate data downstream.
Resolution
To prevent scheduling overlaps caused by the 15-minute jitter, adjust your scheduling strategy using one of the following methods:
Option 1: Adjust your Sync Schedule with a Cron Expression
- Navigate to Connections > Sources and select the Reverse ETL tab.
- Select your Source, open your Model, and click on your Mapping to edit its settings.
- Proceed to the Set sync schedule step.
- Select the Cron expression option.
- Input a Cron schedule that provides a buffer of at least 20 to 30 minutes between your warehouse transformation job and your Reverse ETL sync to safely absorb the maximum 15-minute jitter variation.
Option 2: Use the dbt Cloud Sync Strategy
- Navigate to the Set sync schedule step of your mapping.
- Select the dbt Cloud option.
- This configures Segment to sync only after a job has been completed in dbt Cloud, eliminating the risk of extracting data from an empty or rebuilding view.
Additional Information
Segment's batch engine and downstream destinations typically cache and deduplicate events based on the messageId. Because Reverse ETL generates a fresh messageId when it evaluates a record as "new" following a state clear, downstream systems will correctly ingest the re-extracted rows as completely distinct events rather than deduplicating them.