Last updated: September 28, 2026
By Tom Kenaley, President and Senior Partner, KORE1
A data warehouse migration moves the tables, SQL code, pipelines, and reports off a legacy platform like Teradata, Netezza, or Hadoop and onto a cloud warehouse such as Snowflake, Databricks, or BigQuery. It runs in four phases. Each phase needs different people, and most companies staff it with one migration architect, three to six contract data engineers, and a reconciliation lead.
Almost nobody starts one of these because they woke up wanting a new warehouse. Nobody.
They start because a renewal quote arrived. Or a hardware refresh got priced. Or the vendor put an end date on the thing. IBM ended support for the PureData System for Analytics N3001, the last generation of the classic Netezza appliance, on April 30, 2023. Cloudera’s own support lifecycle table shows Hortonworks HDP 3.1 hitting end of support in December 2021. Teradata customers tend to get their nudge from the contract instead, when a multi-year renewal lands on the CFO’s desk with a number nobody budgeted for.
So the clock is set by finance. The team is set by whoever happens to be free.
That second part is the problem. Plenty of good guides already cover migration strategy, lift and shift versus re-architecting, and which converter to buy. Very few say who should be doing the work in each phase, how many of them, and when they should leave. That’s the question we get on the phone, usually about six weeks after the decision was made and about six weeks too late. Our data warehouse engineer staffing team fills these seats, so read the staffing advice knowing we get paid when you take it. Some of it, though, is advice to hire fewer contractors than you planned.

What a Data Warehouse Migration Actually Moves
A data warehouse migration is the planned transfer of an analytics platform’s data, schema, transformation code, scheduling, security rules, and downstream reports from one warehouse to another, ending with the old system switched off. The data is the easy part. The code and the business logic buried inside it are what take the months.
People scope the tables. Then they find everything else.
On a typical Teradata estate, the inventory looks something like this, and the bottom half of the list is where schedules slip:
- The tables and views, usually thousands of them, with a surprising share nobody has queried in a year.
- BTEQ scripts. Hundreds, sometimes, glued together with shell and called from Control-M or Autosys.
- Stored procedures and macros that encode business rules the finance team forgot were rules.
- ETL mappings in Informatica PowerCenter or DataStage, each one a small program in its own right.
- FastLoad, MultiLoad, and TPT jobs feeding the whole thing overnight.
- Every report in Tableau, Power BI, MicroStrategy, or Cognos pointed at the old connection string.
- Extracts to partners and regulators. The ones with a file spec from 2011 that somebody downstream still validates character by character.
Netezza estates swap BTEQ for NZPLSQL and nzload. Hadoop estates swap it for HiveQL, Spark jobs, and Oozie workflows, plus a pile of Python that ran on an edge node nobody wants to admit exists. Different nouns. Same shape.
The Team, Phase by Phase
Here’s the part I wish more project plans started with. McKinsey’s research on cloud migrations found companies spending 14 percent more on migration than planned each year, with 38 percent seeing delays of more than a quarter. That study covered cloud moves broadly, not only warehouses, but the warehouse version of the story is familiar. The overruns rarely come from the technology. They come from staffing each phase with the people who were right for the previous one.
| Phase | Typical length | Who you add | Who you keep inside | Done means |
|---|---|---|---|---|
| Assessment | 3 to 8 weeks | Migration architect, one senior data engineer | Your warehouse lead, a finance or BI owner | Inventory with a keep, rewrite, or retire call on every object |
| Code conversion | 3 to 9 months | 3 to 6 contract data engineers, the architect | Two internal engineers pairing on the new platform | Every in-scope job runs on the target and passes unit checks |
| Parallel run | 2 to 4 months | QA and reconciliation lead, 1 to 2 engineers for fixes | Report owners who sign off | Two clean month-end closes with numbers that match |
| Decommission | 1 to 3 months | Usually nobody new | One engineer, plus legal and records | Old platform read-only, archived, then off before the renewal date |
The ranges are what we see on mid-sized estates, a few thousand tables and a few hundred jobs. A bank with twenty years of Teradata can double every line. A 400-table Netezza box can halve them.
Assessment Needs One Person Who Has Finished a Migration
Not started one. Finished one.
The migration architect’s job in the first month is to decide what not to move, and you can only make that call confidently if you’ve been through a parallel run and watched which shortcuts came back to bite. On Teradata, the evidence lives in the query log, DBQL, which records who ran what and when. On Hadoop it’s scattered across YARN history and whatever the BI tools logged. An architect who has done this pulls twelve months of usage in week one and walks into the steering meeting with a list.
In one assessment we staffed for a specialty retailer, a little over a third of the tables had no reads in the prior year. Some were abandoned. Dead, really. Some were quarterly, which is exactly why you check twelve months and not three. Retiring those, instead of converting them, took roughly two months off the conversion phase, and nobody on the project would have found it without someone whose first instinct was to ask for the logs.
Keep it small. One architect, one strong senior engineer doing the digging, and a person from finance or BI who can tell you which reports the CFO actually opens. Resist the systems integrator pitch that puts eight people on discovery. Eight is a lot. Most of them will be learning your estate on your dime.
The architect should also own the target-platform decision if it isn’t made yet. Snowflake versus Databricks is a real fork. It changes the conversion tools, the skills you screen for, and who you hire afterward, and our Snowflake engineer staffing and Databricks engineer staffing teams screen for noticeably different things.

Code Conversion Is Where the Contractors Earn Their Rate
This is the headcount phase. It’s also the phase where people most overestimate the tools.
Every major target ships a converter now. Snowflake has SnowConvert, Google has the BigQuery batch SQL translator, and AWS has the Schema Conversion Tool, which lists both Teradata and Netezza as sources. They’re good. Genuinely good. Use them. They’ll turn most of your DDL and a large share of ordinary SQL into something that compiles on the target in an afternoon.
Compiles. That word is doing a lot of work.
| The converter usually handles | A person usually handles |
|---|---|
| Table DDL and data type mapping | BTEQ control flow, error handling, and the shell around it |
| Standard SELECT, JOIN, and window functions | Stored procedures, macros, and recursive queries |
| Most dialect functions and QUALIFY clauses | Semantic differences that compile fine and return different answers |
| A report of what it couldn’t convert | Everything on that report, plus the things that should have been on it |
The third row is the expensive one. In Teradata session mode, which is how most older estates run, string comparisons ignore case unless a column says otherwise. Snowflake’s string comparisons are case-sensitive by default. So a join on a customer segment code that worked for fifteen years, because half the source systems sent “Gold” and the other half sent “GOLD,” quietly drops rows on the new platform. No error. No warning. The query runs faster than it ever did. Teradata SET tables are the other classic. They refuse exact duplicate rows, and on an INSERT-SELECT they drop them without a word, so a target that allows duplicates will happily keep rows the old system never had. Converters flag some of this. The engineer has to know to look for the rest.
That’s the skill you’re paying for. Not “knows Snowflake.” Plenty of people know Snowflake. You want contract data engineers who have read both dialects closely enough to predict where the answers will differ, and who test for it without being told. When we screen for these seats, we ask candidates to explain a conversion that compiled and was still wrong. The good ones have a story ready and it’s usually specific enough to be slightly embarrassing.
How many? A rough rule from the searches we’ve run: one senior contractor per 80 to 150 moderately complex jobs, fewer if the code is mostly straight SQL, more if it’s procedural. Three to six covers most mid-sized estates. Past eight, the architect spends the day reviewing pull requests instead of making decisions, and your throughput goes flat. One client went to eleven on an Informatica-to-dbt job. Against my advice. By month three their architect had forty-odd pull requests a day stacking up in his queue, finished work kept losing to unreviewed work, and four people got cut before the pace came back. Eleven was worse than seven.
Two of your own engineers should pair with them the whole way. I wouldn’t bend on that. The contractors leave, and somebody has to own the new platform at 2 a.m. in month thirteen.
The Parallel Run Needs Someone Whose Job Is to Say No
Every other person on a migration is rewarded for finishing. The reconciliation lead is the one person rewarded for finding reasons you aren’t finished, and that job should belong to one named person, not to “the team.”
The method isn’t mysterious. Row counts first, by table and by partition. Then column aggregates, meaning sums, minimums, and maximums on the columns that end up in financial reports. Then row-level hashes on the tables that matter most, compared key by key. Google’s open-source Data Validation Tool does all three and connects to Teradata, Snowflake, Hive, and BigQuery among others, which saves writing the harness yourself. The tooling part is easy. Cheap, too.
The hard part is temperament. A good reconciliation lead is a little stubborn and very organized, and they don’t accept “that’s a rounding difference” until they’ve seen the rounding.
We placed one on a claims warehouse move for a health plan in Orange County. Week three of the parallel run, every count matched and every sum matched to the cent, and the program director wanted to call it. Our reconciliation lead wouldn’t sign. She’d noticed that one member-level extract matched in total but not in distribution. Inside it, 1,140 members had shifted between two plan categories, and the net effect happened to cancel out. It was the case-sensitivity problem from the section above, sitting in a lookup table. Totals were right. The file going to the state would have been wrong.
That’s the hire. Somebody who reads a matching total as a question. Rare trait.
Run in parallel through at least two month-end closes, and a quarter-end if the calendar allows. The first close shakes out the obvious problems. The second one shows you whether the fixes held. Keep one or two of the conversion contractors on through this phase for fixes, and release the rest, because paying six people to wait for mismatches is how the budget line from the McKinsey number above gets written.

Decommission Is Short, and It Has a Hard Deadline
Brief section. On purpose. The last ten percent of a migration deserves more room than I can give it here. The staffing point is simple, though. Decommission rarely needs anyone new, and it always needs someone named.
Set the old platform to read-only the day the second clean close signs off. Archive what legal and records retention require, in a format someone can still open in seven years. That step trips people up more than it should, since the retention schedule usually lives with a records manager nobody on the data team has ever met, and learning in the last month that seven years of claims history has to stay queryable is not a fun meeting. Then switch it off. Actually off. Do all of that before the renewal date that started the project, because paying for two warehouses for another term is the most common way a migration’s savings disappear. One internal engineer usually owns this, with the architect on call for a few hours a week.
Contract, Project, or Your Own People
Short version. Contract for the surge, keep the platform ownership in-house, and convert one or two of the best contractors if you’re light on permanent staff at the end.
The longer version depends on how clearly the work is scoped. Most migrations fit contract staffing well. That means individual engineers who join your team, work in your repos, and roll off on a date. When the scope is fixed enough to hand over whole, say converting 600 BTEQ scripts to dbt models against an agreed test suite, project staffing on a statement of work can make more sense, because you’re buying an outcome instead of hours.
Cost is the question behind the question. Senior contract data engineers with migration experience typically bill $100 to $150 an hour, the same band our warehouse staffing page quotes, and a proven migration architect runs above that. Compare that with a permanent database architect, whose median pay the Bureau of Labor Statistics put at $139,500 in May 2025, before benefits and before the six months it can take to find one. For a twelve-month project, contract usually wins. Easily. For the platform after the project, it doesn’t. Our salary benchmark tool will give you a local number for the permanent seats, and if you’re weighing that hire now, our guide on how to hire a data warehouse engineer covers the screening side.
One thing I’d skip. Hiring the permanent team first and asking them to migrate. It sounds efficient. It isn’t. In practice you hire for the migration skills, which you need for a year, instead of the platform skills, which you need for a decade.
Questions Data Leaders Ask Before the Renewal Date
Realistically, how long does a Teradata or Netezza migration take?
Nine to eighteen months is the range we see for a mid-sized estate, from the start of assessment to the old platform being switched off. The spread comes almost entirely from how much procedural code there is and how many reports have to be re-pointed.
Large banks and insurers with decades of Teradata history run longer. A single Netezza appliance with mostly straight SQL can finish in well under a year if the parallel run goes cleanly.
Could our own data team do this without contractors?
Some teams can, and the ones that manage it usually freeze most new analytics work for a year to free the hours.
That freeze is the real cost. If the business will tolerate twelve months with no new dashboards or models, go ahead. Most won’t, which is why the common shape is internal engineers owning the target platform while contractors carry the conversion volume.
Lift and shift or re-architect, does it change who we hire?
It changes the ratio more than the roles. Lift and shift leans on conversion engineers and a reconciliation lead, while re-architecting adds data modelers and analytics engineers who rebuild the logic in dbt or Spark rather than translating it.
Re-architecting also stretches the parallel run, because the new numbers are supposed to differ in places and somebody has to prove each difference is intended.
Does the migration architect need a vendor certification?
Short answer: it helps a little, and a finished migration helps a lot more. A SnowPro Advanced or Databricks certification shows platform depth, but it says nothing about whether the person has ever reconciled two warehouses and switched one off.
Ask for the last migration they completed, how big it was, what they chose not to move, and what went wrong in the parallel run. Candidates who have done it answer the last question fastest.
What does a reconciliation lead check first?
Counts, then sums, then hashes. Row counts by table and partition catch missing loads, column aggregates catch type and rounding problems, and row-level hashes on the critical tables catch the rest.
After that, they compare distributions on anything that feeds a regulatory file or a financial statement, because totals can match while the rows underneath have moved.
Roughly what will the contract team cost?
$100 to $150 an hour is the usual band for senior contract data engineers with real migration history, and a migration architect bills above that. Four engineers and an architect for nine months is a meaningful line item, and it’s still typically smaller than one more multi-year renewal on the old platform.
The fastest way to overspend is keeping the full conversion team through the parallel run. Release them as their jobs pass reconciliation. Job by job.
Before You Sign the Renewal
If the quote is already on someone’s desk, the useful move this month is small. Get one experienced migration architect in for the assessment, pull a year of query logs, and find out what you’re actually moving. The headcount plan for the next three phases comes out of that, not the other way round.
KORE1 has been staffing data and infrastructure teams since 2005. Our average time to fill an IT role is 17 days, and 92% of the people we place are still with that client a year later, which matters more than usual when the person is holding the only map of your old warehouse. If a renewal date is setting your timeline, talk to our data engineering recruiters about the architect seat first. Everyone else can start a few weeks behind. The architect usually can’t.

