Back to Blog

The Data Migration Plan Engineering Leaders Should Demand Before Anyone Moves a Row

Big DataInformation TechnologyLeadership

Last updated: October 9, 2026

A data migration plan is a written document that names what moves, how each field maps, which numbers must reconcile before cutover, who approves each gate, and the exact conditions that trigger rollback. Most of what gets circulated as a migration plan is a Gantt chart with the word “migration” on it. That is a schedule. It is not a plan, and it will not tell you whether to stop.

I have staffed data migration work across payments platforms, lenders, insurance back offices, and a few healthcare billing shops. The pattern does not change much. Somebody has already picked the target system. Somebody has already told the board a date. And the document that is supposed to govern the riskiest two days of the year is four slides long.

So this is the plan I would ask for before anyone runs a single load job. Seven sections. Each one has an owner whose name is typed into the document, not implied by an org chart. If you are moving transaction data, the staffing side of this eventually lands on data engineer and data scientist staffing, and I will get to that. The plan comes first.

Data engineer monitoring a platform data migration cutover window overnight at a single workstation

What a Data Migration Plan Actually Is

A data migration plan is the control document for moving records from one system to another. It defines the record boundary, the field-level mapping and its transformation rules, the reconciliation checks that must pass, the rollback triggers, and the named owner of every approval gate. It is the artifact an auditor asks for.

That last sentence is the one people skip past. Do not.

In regulated environments the plan is not an internal nicety. The FFIEC’s Development, Acquisition, and Maintenance booklet, issued to examiners through OCC Bulletin 2024-26 in September 2024, exists so examiners can “evaluate a financial institution’s controls and risk management processes” over project management and the system development life cycle. Nobody from the OCC is going to ask you for your Gantt chart. They are going to ask who signed off, against what evidence, and what you would have done if the evidence had come back wrong.

If you cannot answer that from a document, you are going to answer it from memory, in a room, eighteen months later.

What Separates a Plan From a Template

Search “data migration plan template” and you will get thirty results that all contain the same nine headings. Scope. Mapping. Testing. Cutover. Sign-off. They are not wrong. Just unfalsifiable. “Data validation passed” is not a gate. It is a mood.

The difference is whether each line of the document can fail. A real plan has numbers in it, and the numbers have owners, and there is a written answer for what happens when a number comes back outside its range at 2:40 in the morning on a Saturday. Everything below is built around that test.

Section One: The Record Boundary

Before mapping, before tooling, before anybody opens a terminal, write down what is not moving.

Not what moves. What does not.

This sounds procedural. It is actually the single biggest cost lever in the project, and it is almost always decided by default instead of by decision. The default is “bring everything,” because nobody wants to be the person who threw away a record somebody needed. Seven years of transaction history comes across, along with the dead accounts, the test records somebody created in 2019, the duplicate customer rows nobody reconciled, and four tables that feed a report one person runs each December.

So tier the data instead. Three buckets is enough:

  1. Moves and must reconcile to the cent. Open balances, active accounts, anything with a dollar figure a controller has already signed and reported. This tier gets the strictest gate and the most rehearsal time.
  2. Moves, reconciles on counts only. Closed records inside the retention window, historical transactions, resolved cases. You need them present and queryable. You do not need them balancing against a general ledger.
  3. Does not move. Archive it, and write down where the archive lives and who can read it.

The third bucket is where the arguments happen. Let them happen. A lender I worked with found that 61% of their customer rows had no activity in four years. Moving them would have added roughly nine days of mapping and rehearsal work to the project. They archived the lot and kept a read-only extract for the two people who actually pull that history.

Retention is a legal input here, not an engineering one, and the number of migration projects I have watched stall because an engineering lead guessed at a retention window rather than asking the person who would carry the consequence is genuinely higher than it should be. Ask once, in writing. Then put the answer in the plan with that person’s name beside it. This is also the point where you decide what happens to the source system afterward, which is its own project that nobody budgets for, and which I have watched run for years past the migration date. We wrote up what decommissioning a legacy system actually takes separately because it kept coming up.

Section Two: Field Mapping Is a Contract, Not a Spreadsheet

Every field that moves gets a row. Source name, target name, transformation logic, version, owner, and the edge cases that were tested. All six. No blanks.

A straight copy still gets a row. That is the part teams resist, and it is the part that saves you. “Copy as-is” written down and versioned is a commitment somebody made on a date. An unwritten straight copy is an assumption, and when the target truncates a 60-character field to 50 you will spend a day finding out which assumption broke.

Version the mapping document. Actually version it, in the repository, next to the code, because the mapping is code and it behaves like code. Somewhere in the project a mapping rule will change after the first rehearsal, and if the change lives in a comment thread on a shared drive you have lost the chain of custody on your own transformation logic.

Two edge cases earn a test every time, in my experience of watching these go sideways:

  • Dates. Timezone handling, null dates, the 1900-01-01 placeholder somebody used as a default, and anything stored as text.
  • Money. Precision, rounding direction, and negative values that were encoded as a string with a trailing minus sign. That one is real. It cost a client a full rehearsal cycle.

Encoding, too. If the source is an older SQL Server instance and the target is PostgreSQL or Snowflake, assume character set problems until a test proves otherwise. Somebody’s name has an accent in it. It always does.

Two reviewers comparing printed source and target ledgers during data migration reconciliation

Section Three: The Reconciliation Gate

Matching row counts prove almost nothing. They are the first check, not the gate.

This is where most plans stop. It is also where most of the damage gets through.

Google says so itself. Its migration guidance states that directly comparing the source and target databases “is not a solid approach for verifying completeness and consistency,” because items get extracted and then legitimately filtered out before insert. So track what you excluded. Deliberately, as output from the pipeline, and reconcile source against target plus exclusions. If your pipeline cannot produce that exclusion list, your pipeline cannot be reconciled.

Here is the gate I would hold a cutover against.

CheckWhat it actually provesPass condition
Count parity with exclusionsNothing vanished silently between extract and insertSource count minus approved exclusions equals target count, exactly
Control totalsThe money is the same moneyZero variance on every tier-1 balance, to the cent
Referential integrityChild records still point at parents that existZero orphaned keys on tier-1 tables
Distinct-key cardinalityYou did not quietly deduplicate or fan out rowsDistinct business keys match source plus exclusions
Business owner sampleThe records are right, not merely present25 records the owner selected, reviewed in the target UI, zero defects
Report parityDownstream consumers see what they saw yesterdayTop five production reports match old system output

The last row catches more problems than the first five combined. Every time. A migration can be technically perfect and still break the business, because the report that the operations team runs every morning was built on a view that assumed a column ordering nobody documented.

One note on the business owner sample. Let them pick. If you pick them, you will pick the clean ones, and you will not mean to. Reconciling to a number finance already published is the same discipline we wrote about in the ERP data migration guide, where the general ledger gives you an externally signed figure to hold yourself against. On a platform migration you may have to go find that number first.

Section Four: The Dual-Write Window, and What It Costs You

Dual-write is where plans get optimistic.

The appeal is obvious.

The theory is clean. Seductively so. Point the application at both systems, let writes land in each, and you can cut back to the source at any point because the source never went stale. Rollback becomes a configuration change instead of a restore. Every engineering leader likes this idea the first time they hear it.

Google’s guidance is worth reading before you commit to it. Two writes are two separate transactions, so a failure can land after the first commits and before the second does, leaving you inconsistent in a way neither system can detect on its own. Worse, the clients “do not know the source database transaction commit order,” so they can reorder writes and produce divergence that looks like corruption. Google’s own Database Migration Service does not support a dual-write setup at all.

None of that makes dual-write wrong. It makes it a build, not a toggle. That distinction has budget consequences. If you want it, the plan needs to say who is writing the parity checker, because a dual-write window without automated divergence detection is just two databases you now get to be wrong about simultaneously.

Three things belong in this section if you are doing it:

  • The divergence check. What compares the two writes, how often, and what it does when they differ. Automated. Not a daily query somebody runs.
  • Who repairs a divergence, and whether repair is forward-only or requires a replay. Write the decision down before you need it.
  • The window length, with an end date. Dual-write windows that stay open become permanent architecture, and then you are paying to run two systems and reconcile them forever.

For event-driven stacks there is often a cheaper path. If both the old and new consumers can read from the same Kafka topics, you may not need application-level dual writes at all, which removes the ordering problem rather than managing it. That is worth half a day of design time before you accept the complexity.

Section Five: Rollback Criteria, Written Before the Window Opens

Write the rollback triggers down in advance, as numbers, and have the person with authority agree to them while everyone is calm and rested.

Calm and rested matters.

The reason is not process hygiene. It is that at hour five of a cutover, with the window closing and a reconciliation check failing by a small amount, the room will negotiate. Somebody will say it is probably a rounding artifact. Somebody will point out how much it cost to get here. A threshold agreed to in advance removes the negotiation, and that is its entire purpose.

TriggerThresholdAction
Tier-1 balance varianceAny non-zero amountNo-go. Do not proceed to the switch.
Orphaned keys, tier-1Greater than zeroNo-go
Load or reconcile runtimeExceeds rehearsed time by 25%Hold and reassess against the window, do not push through
Critical workflow smoke testAny failure in the named setRoll back
Error rate on new writesAbove pre-agreed baseline for 15 minutesRoll back inside the window
Business owner sampleAny tier-1 defectNo-go

Those thresholds are a starting point, not a standard. Google’s cross-region migration guidance makes the same point about writing the contingency down in advance, saying a plan should specify when to retry and resume the copy or fill in the gap, and when to do a complete rollback and recopy. Borrow the shape, not the figures. Set yours from your own service levels and your own tolerance. The numbers matter less than the fact that they exist and that somebody with authority agreed to them in writing.

Name the Point of No Return

Every cutover has a moment after which rollback stops being a switch and becomes a restore. Usually it is the first production write to the new system that has no equivalent in the old one. Put that moment in the plan, by name, with a timestamp estimate.

Before it, you can go back cheaply. After it, going back means replaying or discarding real customer activity, and that is a business decision, not an engineering one. Teams that have not identified this moment cross it without noticing. Then they discover their rollback plan was written for a world that ended ninety minutes ago.

Test the reverse path too. A working forward migration does not prove the reverse works, and a rollback procedure that has never been executed is a hypothesis. Rehearse it once. Badly is fine.

Section Six: Decision Rights

This is the section that is almost always missing, and it is the reason migrations stall instead of fail.

Each gate needs one named person who can say no and one named person who can override. Not a team. Not “IT leadership.” A person. With a phone number in the runbook.

GateEvidence requiredWho approves
Record boundary and retentionWritten retention determination, tiering sheetLegal or compliance, plus the data owner
Mapping sign-offVersioned mapping doc, edge-case test resultsMigration lead and target system owner
Rehearsal acceptanceDefect log, timed runbook, reconciliation outputMigration lead
Reconciliation passAll six checks green, exclusion list attachedController or equivalent business owner
Security and accessRole mapping reviewed, privileged access removedSecurity reviewer
Final go or no-goEvery gate above, signedOne named executive sponsor

Note who owns the reconciliation gate. Not engineering. The person who has to stand behind the numbers afterward is the person who should be allowed to reject them, and if your controller has never seen the reconciliation output before cutover night, you have a sign-off problem that no amount of testing fixes.

One more thing about the reconciliation lead. Make it somebody whose job that week is to say no. If the same person is accountable for hitting the date and for validating the data, the date wins. It wins quietly, through a hundred small judgment calls, and nobody ever decides to let it happen.

Section Seven: Rehearsals, and How Many

Two minimum. Three if the data is regulated or the window is short.

Nobody has ever regretted the third one.

The first rehearsal is for finding out that your assumptions about the source data were wrong. It will be ugly, it will overrun, and that is the rehearsal doing its job. The second is for timing the runbook and proving the defect fixes held. A third exists to prove the second was not luck.

Keep the failure logs from every rehearsal, including the ones you fixed. When somebody asks in six months why a field is formatted the way it is, that log is the answer, and it is the kind of evidence that shows the process had rigor rather than momentum.

There is a reason to be aggressive about rehearsal count and ruthless about project duration at the same time. McKinsey’s research with Oxford’s BT Centre for Major Programme Management, covering more than 5,400 IT projects, found that large IT projects ran on average 45% over budget while delivering 56% less value than predicted, and that every additional year of scheduled duration increased cost overruns by about 15%. Long migrations do not get safer. They get more expensive and they accumulate scope, which is why a tight boundary in section one pays for itself twice.

Migration lead halting a go or no-go review before a data migration cutover

Who Actually Executes This, and What It Costs

Now the part I get called about.

Usually late. Sometimes very late.

A plan like this implies a squad, and the squad is almost never the team you have. The wider bench sits under our IT staffing services. Your platform engineers know the target system, and we keep a separate bench for data engineers who have run a migration before. Your application owners know the business rules. Neither of them has run a reconciliation gate against a controller before, and neither of them is going to stop building product for eleven weeks.

Four roles genuinely need somebody who has finished one of these before, and the first two are usually filled through senior data engineer staffing:

RoleTypically engagedWhy it is usually contract
Migration leadFull project, 10 to 16 weeksYou need pattern recognition from three or four prior cutovers, and you need it once
Data engineers, 2 to 3Weeks 2 through 12Pipeline and transformation build is a burst of work that ends
Reconciliation leadWeeks 5 through cutoverNeeds independence from the delivery date, which is hard internally
QA or data quality analystWeeks 4 through cutoverRehearsal defect triage is full-time while it lasts, then it is not

On rates, our published 2026 benchmarks put a contract data engineer at $75 to $120 an hour as a W-2 pay rate, with a typical agency bill rate of $108 to $175, and the full breakdown across roles is in our tech contractor hourly rate guide. Senior data engineers on a permanent basis run $147,000 to $179,000 and up, per our data engineer salary guide. If you want to sanity-check a band for your own market before you post anything, the salary benchmark assistant will do it in a couple of minutes.

Worth knowing what the broader market looks like underneath those numbers. The U.S. Bureau of Labor Statistics put the May 2025 median annual wage for database architects at $139,500 and database administrators at $104,620, with combined employment projected to grow 4% from 2025 to 2035 and about 7,300 openings a year. Modest growth against a small pool, which is the part that catches hiring managers out, because a 4% growth rate sounds like an easy market right up until you start screening for the narrow experience this work actually requires. The practical consequence is that the person who has actually run a reconciliation gate is not sitting in an inbox waiting. Across our own searches, KORE1 averages 17 days to fill an IT role, and migration specialists sit at the slower end of that because the qualifying question is narrow.

Most teams end up using contract staffing for the build phase specifically, which keeps the core engineering team on the roadmap instead of on a one-time data move. We put the squad-shaped version of that into data engineering staff augmentation because the request comes in that shape so often. If the move is a warehouse rather than an application platform, the phase-by-phase team composition is different enough that we covered it in the data warehouse migration guide.

Before You Approve the Window

How long should a plan like this take to write?

Two to three weeks of real effort. Mostly section one. The boundary decisions require legal, finance, and business owners to actually answer questions, and that calendar time is the constraint. The mapping document is slower than people expect but it is predictable work. If somebody hands you a complete migration plan in two days, they filled in a template.

We already have a project plan from our implementation partner. Is that the same thing?

Usually not. A partner’s project plan covers their scope of work, which typically ends at loading data you supplied and validated. Read the reconciliation and data-quality language in the statement of work closely. In most of the engagements I see, data cleanup and reconciliation sit on the client side, and that is where the surprise headcount comes from.

Is dual-write worth the complexity for a mid-market platform move?

Only if you cannot afford to lose writes made after the switch. Dual-write buys fast rollback, and it costs you a parity checker, divergence repair procedures, and ordering problems that are genuinely hard to reason about. Can your window absorb a restore-based rollback? Then take the simpler path. Payments and anything balance-bearing is where the complexity earns its keep, and you budget engineering time for it accordingly.

Our own team wants to run this. Should I let them?

Partly. They should own the business rules, the target configuration, and the final sign-off, because that knowledge has to stay in-house afterward. What usually does not work is asking the people responsible for the product roadmap to also build migration pipelines and run three rehearsals. Something gives, and it is normally the rehearsals.

What is the first sign a migration is quietly going wrong?

Rehearsal dates moving while the cutover date stays fixed. That is the tell. Visible months out, usually. The compression always lands on validation, because validation is the only part of the plan with no external dependency forcing it to happen. Watch for the exclusion list growing late, too. Late exclusions are often defects being reclassified as scope.

Who should own the go/no-go call if we do not have a CIO?

The executive who owns the business process, not the one who owns the technology budget. In a mid-market company that is frequently the CFO or a COO. The requirement is that they are senior enough to absorb the cost of a no-go and far enough from the delivery date to make that call honestly.

The One Page I Would Ask For First

If you only get one artifact out of all of this, ask for the gate sheet. Six rows. The evidence each gate requires, and the name of the person who approves it.

It fits on a page. An afternoon of work, maybe less. And the conversation it forces, about who is allowed to stop a cutover that is already in motion, is the conversation that determines how the weekend goes. Teams that can produce that page tend to have thought about the rest. Teams that cannot are usually further from ready than their schedule suggests.

We staff the squads that execute these, and our recruiters average more than 15 years in the market, which mostly means we have seen which of these roles clients regret filling last. If you are scoping a move and want to talk through what the team should look like before you commit to a date, talk to a recruiter on our team.