Changing Sales Engagement Platforms Mid-Experiment: Keep Your A/B Test Interpretable
TL;DR: A platform migration can alter assignment, sender infrastructure, sequence timing, exclusions, reply classification, and outcome capture at once. Freeze the experiment manifest before moving anything. Finish the old test, restart under the new system, or analyze pre- and post-migration cohorts separately. Do not pool results and call the larger number a winner.
An outbound A/B test answers a narrow question only when the variants differ in the intended way and outcomes are measured consistently. Changing the sales engagement platform during the test can quietly change more than the visible copy. The new system may use different enrollment rules, sender mailboxes, scheduling windows, reply labels, suppression logic, or CRM write-back. If those changes coincide with a variant, the result is not a clean estimate of the copy change.
Our Sequencing supports coordinated email, phone, and social outreach. Our Analytics covers activity, sequences, plays, attribution, and export methods. Neither capability guarantees that a third-party platform’s experiment metadata will migrate automatically. Treat transferability as a test requirement, not an assumption.
Freeze the experiment before changing the system
Before migration, export an immutable manifest for every enrolled contact and every planned step. It should record the randomization unit, variant assignment, enrollment timestamp, sender, sequence version, step schedule, exclusions, completed actions, reply events, and outcome definition. Preserve the source platform identifiers alongside the CRM person and account identifiers so a later analyst can reconcile the two systems. Document the extraction time and any records excluded from the export.
| Field | Why it matters | Migration check |
|---|---|---|
| Experiment and variant ID | Keeps original assignment stable | Same person retains same assignment |
| Randomization unit | Prevents person-level and account-level mixing | Unit is explicit for every record |
| Enrollment and step history | Distinguishes exposure before and after switch | Completed steps do not replay |
| Sender and mailbox | Separates copy effects from sender changes | Mailbox identity and health recorded |
| Eligibility and exclusions | Explains who entered the denominator | Old and new rules compared |
| Reply and meeting outcome | Keeps classification and measurement consistent | Raw events and labels retained |
| CRM account and person IDs | Supports reconciliation without false matches | Identity conflicts quarantined |
This is an editorial checklist, not a native-export schema for a specific platform. If a field cannot be exported, record the gap and its effect on interpretation. A screenshot of a dashboard total is not a substitute for contact-level assignment and exposure history.
Decide whether to finish, restart, or split the test
The cleanest option is often to finish the old cohort on the old system while preventing new enrollment, then start a new experiment after migration. That avoids making platform a mid-test variable for active contacts. When the old system cannot remain active, move operationally but analyze the phases separately. Restart if assignment, exposure, or outcomes cannot be reconstructed. There is no universal answer; choose according to the data you can preserve and the operational risk of parallel systems.
| Condition | Recommended analysis | Reason |
|---|---|---|
| Old cohort can finish safely | Freeze intake and finish in old system | Keeps delivery conditions stable for existing assignments |
| Move required, history and variant IDs preserved | Separate pre- and post-switch cohorts | Platform period may still change outcomes |
| Assignment or exposure history missing | Restart the test | Cannot attribute the observed difference to the intended variant |
| Sender, timing, and audience changed together | Treat as a new operating experiment | Multiple changes confound a copy A/B result |
Do not combine cohorts solely to recover sample size. A larger pooled dataset can be less interpretable than two smaller, well-described datasets. If there is no defensible comparison, report the descriptive outcomes and the limitation instead of declaring a winner. The UK Government Data Quality Framework gives a useful general lens for completeness, consistency, and provenance, although it is not an outbound-testing standard.
Keep assignment and exposure separate
Assignment says which variant a contact was meant to receive. Exposure says what they actually received. A migrated contact may retain variant B in the CRM but receive variant A copy in the new tool, miss a step, or receive both versions. Store both values. For each record, identify whether it was assigned, enrolled, sent the first step, sent later steps, replied, or was suppressed. Do not treat an enrollment record as proof of delivery.
The randomization unit matters. If several people at one account are assigned independently, a buyer may encounter both variants; if the account was the unit, preserve that assignment across all contacts. Reassignment after migration can contaminate the test. Hold identity conflicts in a review queue rather than hashing on a new platform ID and silently reallocating the contact.
Compare delivery and measurement rules side by side
Build a preflight diff between the old and new platforms. Include send windows, timezones, holiday calendars, reply stops, bounce handling, link tracking, call task creation, social-task scheduling, and CRM sync behavior. Then verify each with controlled records. The migration is not complete when records import; it is complete when the team can explain what an eligible person will receive next and where the outcome will be recorded.
- Audience: Compare inclusion and exclusion logic before allowing any new enrollment
- Message: Compare subject, body, personalization inputs, and fallback text for each variant
- Delivery: Compare sender identity, step order, business-hour rules, and throttling
- Stop rules: Confirm reply, opt-out, bounce, and open-opportunity behavior in the new scheduler
- Outcome: Preserve raw reply text and event timestamps before mapping categories
- Attribution: Define whether meetings and opportunities are counted by person, account, or enrollment
Our Find What Your Best Sequences Do Differently use case compares sequence length, timing, channels, audience, and classified replies. The example data on that page is illustrative. Those are precisely the dimensions to hold stable or record as changed during migration, not evidence that one particular configuration wins generally.
Protect active conversations during the move
An active reply is not experimental inventory. Before moving an enrollment, inspect the latest mailbox thread, owner, scheduled steps, and opportunity state. Pause or suppress any step that could conflict with a live conversation. The new platform should not replay already-completed touches, and the old one should not continue sending after the record moves. Assign one transition owner and record the cutover state for each person.
Our Spot Stalled Sequences and Quiet Deals use case calls out people receiving email after replying. During a migration, this failure can occur if the old scheduler remains active or the new platform imports a stale enrollment. Use that workflow as an audit idea, not the illustrative examples as a performance claim.
Validate CRM and export paths rather than assuming parity
If the CRM is the source of ownership or opportunity context, compare how each platform reads and writes those fields. Our Salesforce integration guide covers mapping, read permissions, separate write enablement, and supported objects. We describe the read schedule as approximate, not a real-time guarantee. Confirm the fields used for exclusions and outcomes on controlled records before live enrollment.
For analysis, retain raw events from both systems in a common dataset with source-platform and migration-period fields. Do not overwrite the original timestamp or variant label during normalization. Create a reproducible mapping from old reply classifications to new ones, and flag categories that cannot be equated. If the definition of a positive reply changed, the headline metric changed even if the dashboard label did not.
Interpret results with a clear limitation statement
Report denominators and exclusions for each phase, then show the planned analysis and any deviations. If the test finished entirely before migration, the platform switch did not affect its delivery but may affect later CRM attribution. If active contacts moved, describe the proportion with complete history and any changed send conditions. If those details are unavailable, call the outcome descriptive. Do not write “variant B won” when a sender or timing change traveled with variant B.
The decision rule should be written before viewing the results: primary outcome, observation window, minimum data quality, and how ties or ambiguous classifications will be treated. This is not a demand for an arbitrary significance threshold. It is a way to prevent a convenient post-migration metric from replacing the question the team originally intended to answer.
Check the integrity of the allocation itself
Before interpreting outcomes, compare assignment counts and eligibility by variant in each period. A sudden shift in the balance of assigned contacts may reflect an import filter, a changed randomization rule, or account-level duplication rather than chance. Then compare assigned, exposed, and outcome-observed populations. If one variant lost more records during migration, a favorable response rate among survivors can be misleading. Investigate the missing records rather than removing them silently from the denominator.
Document any manual reassignment, sender replacement, or suppression override. These events may be necessary for operations, but they are deviations from the original test. Report them separately. When the migration changes who can be enrolled, compare like-for-like eligibility criteria or explicitly narrow the question to the new population. A test about message wording cannot answer whether a new targeting engine improved results if targeting changed at the same time.
Five mistakes to avoid
- Pooling old and new platform cohorts without a migration-period field
- Treating assigned contacts as though every scheduled step was delivered
- Changing sender, copy, and audience while labeling the exercise a copy-only test
- Replaying completed touches or leaving both schedulers active for one person
- Mapping unlike reply categories into one positive-reply metric without review
Glossary
- Assignment: The variant designated for a person or account before exposure
- Exposure: The message or step actually delivered or performed
- Manifest: A frozen record of experiment design, membership, steps, and outcomes
- Confounding: A simultaneous change that prevents isolating the effect of the intended variant
- Migration period: A labeled phase in which platform behavior or measurement may differ
To put the relevant workflow into practice, sign up for Unify and review the proposed actions before enabling live outreach.
Frequently asked questions
Can an active A/B test continue after switching platforms?
Only if assignment, exposure, eligibility, and outcome history can be preserved and the changed platform conditions are explicitly accounted for. Otherwise finish the old cohort or restart.
Should pre- and post-migration results be pooled?
Not by default. Keep the periods separate unless the delivery and measurement conditions are demonstrably comparable and the analysis plan justifies pooling.
Which fields are essential to export?
Preserve experiment and variant IDs, randomization unit, contact and account IDs, send history, sender, timestamps, exclusions, raw replies, and outcome definitions.
What if the new platform classifies replies differently?
Keep raw reply text and original labels, map categories explicitly, and flag non-equivalent definitions. Do not compare unlike positive-reply metrics as one measure.
How do we prevent duplicate sends during migration?
Freeze enrollment, record each person’s last completed and next scheduled step, transfer one owner, then verify that only one scheduler remains active.
When should the team restart the test?
Restart when variant assignment, actual exposure, or primary outcomes cannot be reconstructed, or when several delivery conditions changed together.
Sources
- Unify Sequencing
- Unify Analytics
- Find What Your Best Sequences Do Differently
- Spot Stalled Sequences and Quiet Deals
- How to integrate Unify with Salesforce
- The Government Data Quality Framework, GOV.UK
Written by Austin Hughes, Co-founder and CEO of Unify.

