Last updated: September 7, 2026
By Robert Ardell, Co-Founder and Strategic Advisor, KORE1
For data teams, contract versus full-time should be decided one layer at a time: contract the platform build and the migration spike, keep the semantic layer and pipeline on-call full-time. Treating the whole team as one staffing decision is what produces a warehouse nobody can explain and a bill nobody can stop.
I co-founded KORE1 in 2005. Two decades of watching this question get asked, and the framing has barely moved even though the work underneath it has changed completely.
The call goes roughly the same way every time. A hiring manager says they are deciding between a contractor and a full-time data engineer, singular, as though the data team were one seat with one answer attached to it. It almost never is. A modern data org is four or five distinct jobs wearing similar titles, and the honest answer is different for each one. Contract is right for some of them nearly always. Wrong for others nearly always. The interesting part is the middle.
Most of what gets published on this topic is written one level up from the actual decision. Bounded project or permanent function. Budget certainty. Cost of a bad hire. That framing is fine and we have written our own version of it for contract versus full-time IT hiring generally. It just does not survive contact with a data team, for two reasons that are specific to this function and almost nowhere else.
I should say where I sit before you weigh any of this. KORE1 runs a data engineer staffing desk, we bill for both models, contract and direct, and we have since the beginning, so the choice you make is close to revenue-neutral on our side. Read that as a conflict or as the reason I can be straight about it. Your call. Most of what follows is something you can run without calling anyone, and a fair number of the teams I talk to should.

Why the Usual Math Breaks on a Data Team
Two things make data different.
The first is that data work is lumpy by design. Not by accident, not because of bad planning. A warehouse migration, a Redshift to Snowflake move, a lakehouse consolidation after an acquisition, these are genuine spikes that consume three or four engineers for two or three quarters and then release almost all of that capacity. Software teams have crunch. Data teams have a shape, and the shape repeats. Every company I have watched go through a platform move has staffed up for the build and then spent the following year quietly figuring out what to do with the people.
The second thing is stranger, and it is the one that costs money. When a software contractor leaves, you have a gap. When a data contractor leaves, you have a gap and a running meter.
Their pipelines still execute. Their models still refresh at 4 a.m. Their queries still scan whatever they were told to scan, at whatever cost the warehouse charges for it, on a schedule they set in a config file nobody has opened since. The invoice stopped in March. The Snowflake credits did not.
That asymmetry does not exist in most engineering functions and it is the thing that ought to reshape the decision.
Decide It One Layer at a Time
Stop asking whether your data team should be contract or full-time. Ask it per layer. The answer changes, and it changes for reasons you can actually defend in a budget meeting.
| Layer | What it owns | Default model | What breaks if you get it wrong |
|---|---|---|---|
| Ingestion and platform | Connectors, orchestration, warehouse and lakehouse infrastructure, Terraform | Contract works well | Very little. The artifacts are code, and code survives the person who wrote it. |
| Transformation and semantic layer | dbt models, metric definitions, the meaning of every number leadership sees | Full-time | Everything downstream. Nobody can say why revenue is defined the way it is, so nobody trusts it. |
| Analytics and BI | Dashboards, ad hoc analysis, the relationship with the business unit | Full-time core, contract the backlog spike | Output that answers the question asked instead of the question meant. |
| ML and data science | Features, model development, evaluation, retraining | Split by stage | A model in production with no owner and no evaluation harness. |
| Governance and access | PII classification, row-level security, lineage, audit response | Full-time | Accountability with no name attached to it when the auditor asks. |
Five rows. Two of them settle themselves, one contracts cleanly, and the other two are where judgment actually lives.
Run your own org against that table before you run any rate comparison. Most teams discover they have it exactly inverted, because contract budget is easier to get approved than headcount, and platform work is the thing executives most want to see a full-time name attached to. So the contractor ends up owning the metric definitions while the FTE writes Terraform. That is backwards and it is extremely common.
The Layer Where Contract Almost Always Wins
Migration work. Platform build-outs. Anything with a cutover date on it.
The federal labor data makes this case better than I can. The Bureau of Labor Statistics projects employment of database administrators and architects to grow 4 percent from 2025 to 2035, from 144,500 jobs to 150,900. Unremarkable on its own. Then look at how BLS splits it. Database administrators, the maintenance seat, come in at zero percent growth for the decade. 75,000 jobs down to 74,900. Database architects, the seat that builds things, add 9 percent and go from 69,500 to 76,000. Same occupation code, two halves, walking away from each other.
Flat maintenance. Growing build. That is the shape of a function where the interesting work arrives in waves and the steady state gets automated down.
Compare it to data scientists, projected to grow 35 percent over the same decade, 275,600 jobs to 371,000, with roughly 24,800 openings a year. Database administrators and architects combined generate about 7,300 annual openings. Both of those numbers are real and they point in opposite directions, which is exactly why one staffing model across the whole team is a mistake.
A platform contractor is a good trade when the work has an end. We have staffed a lot of these. In our own placements a mid-size warehouse migration tends to run somewhere between five and nine months from kickoff to cutover, and the team is roughly twice its steady-state size for the middle third of that. Those are our numbers from our own delivery, not a published benchmark, so weigh them accordingly. But the pattern is consistent enough that I would plan against it.
Hire four full-time engineers for that spike and you have made a ten-year commitment to solve a seven-month problem. Then the migration lands and you are managing two people out of a company they joined in good faith, which is a bad outcome for everyone and an expensive one for you.

The Layer Where Contract Almost Never Wins
Now the other end.
Somewhere in your stack there is a definition of “active customer.” It excludes internal test accounts. It excludes the two enterprise logos that churned but stayed on a legacy contract through the end of the fiscal year. It counts a seat as active at 30 days rather than 28 because of an argument the VP of Sales won in 2023 and everyone else conceded.
None of that is in the repo. Some of it is in a dbt model as a WHERE clause with no comment above it. The rest lives in the head of whoever was in the room.
That is the semantic layer, and it is the single worst thing to staff on a statement of work. Not because contractors do it badly. Plenty do it well. The problem is that the artifact they produce is only half code, and the other half walks out with them on the last Friday of the engagement.
I watched a Costa Mesa logistics company learn this the hard way about three years ago. Good contractor, sharp, built out their entire Looker semantic model over eight months. Left on schedule, everything documented to the standard anyone would call reasonable. Six months later finance and operations were reporting on-time delivery rates that differed by eleven points, and it took an internal analyst most of a quarter to work out that the two teams were inheriting different definitions of when a delivery clock stopped. Both were defensible. One was in the dbt project, one was in a Looker view, and the person who could have said which was intentional had been gone for half a year.
No villain in that story, which is the part worth sitting with. The contractor delivered what the statement of work described, the engagement closed on schedule with documentation everyone signed off on, and the knowledge evaporated anyway, because the knowledge was never actually the deliverable.
Keep that layer in-house. If budget forces your hand, at minimum pair the contractor with a full-time analytics engineer who owns the definitions and reviews every merge. We staff a lot of analytics engineer seats specifically for this, and it is the seat most under-hired relative to how much damage it prevents.
The Cost Line Nobody Puts in the Comparison
Here is the arithmetic most teams run. It is not wrong. It is just incomplete.
| Line | Full-time data engineer | Contract data engineer |
|---|---|---|
| Direct cost | $125,000 to $135,000 base, national mid-band | $108 to $175 per hour billed |
| Annualized at 2,080 hours | Roughly $160,000 to $175,000 loaded | $225,000 to $364,000 if run a full year |
| Time to first commit | Search plus notice period, often 10 to 14 weeks | Days to weeks, subject to data access |
| Cost if you end it early | Severance, morale, and a re-run search | Notice period in the SOW |
| Compute they leave running | Owned, tuned, and someone’s performance review | Unowned and recurring |
The published bill rates there come from our own tech contractor rate benchmarks, and the salary band from our salary benchmark tool. Look at the annualized row and the conclusion writes itself. Contract is roughly 1.4 to 2 times the cost of an equivalent FTE if you run it for a year.
Which is exactly why nobody should run a data contractor for a year without a reason they can say out loud.
Now the last row, the one that is not on anyone’s comparison sheet. A data engineer’s output is not just code. It is code that consumes metered compute, forever, until somebody turns it off. An hourly full-refresh on a table that only needed an incremental. A dbt model materialized as a table because it was faster to develop that way. A dashboard scheduled to refresh every fifteen minutes for an executive who looks at it on Mondays.
The FinOps Foundation’s State of FinOps 2026 report, which surveyed 1,192 practitioners who collectively manage more than $83 billion in annual cloud spend, found that 98 percent of them now manage AI spend as part of the job, up from a small minority two years ago. AI cost management is the number one skill those teams say they are trying to add. It got that way because this exact thing kept happening.
None of that argues against contractors. It argues for one clause in the SOW. Cost review at exit, warehouse credit consumption attributable to the engagement reviewed before the final invoice clears. We started recommending it in 2024 after a client in Irvine found a single scheduled job, written by a contractor who had been gone eleven months, that was the third-largest line in their Snowflake bill. Nobody had touched it. It had simply been running.
Access Is the Hidden Clock on Every Data Contractor
Speed is the standard argument for contract. In data specifically, it is weaker than people think.
A software contractor can be productive in a sandbox on day two. A data contractor cannot do anything at all until they are inside the production data, and if that data contains protected health information, cardholder data, or anything a SOC 2 auditor will ask about, you are looking at a real security review with a real queue in front of it.
The risk is not hypothetical. The 2026 Verizon Data Breach Investigations Report found that 48 percent of breaches now involve a third party, a 60 percent jump over the prior year’s dataset. Your security team has read that report. It is why the access request that “should take a day” takes five weeks in a regulated shop.
So the real comparison is not ten weeks of search versus two weeks of contractor onboarding. In a healthcare or fintech environment it is closer to ten weeks versus six, and at that point the speed advantage has mostly evaporated while the cost premium has not.
Unregulated environment, straightforward stack? Contract is genuinely fast, and this section does not apply to you. Know which one you are before you build the timeline.
On-Call Decides More of This Than Budget Does
This is the question I would ask first if I only got one.
When the ingestion job fails at 3 a.m. on a Sunday and the executive dashboard is empty at 7 a.m. Monday, whose phone rings?
Most contract agreements do not include production on-call, and the ones that do price it accordingly. That is reasonable. It is also the fact that quietly settles the staffing model for any pipeline the business depends on, regardless of what the budget says. If the answer to that question is a person whose engagement ends in six weeks, you do not have an owner. You have a temporary arrangement that the org has started depending on permanently.
I have seen teams discover this at the worst possible moment more times than I can count. Board meeting Tuesday. Numbers wrong Monday night. The one person who understood the reconciliation logic rolled off the project in April.
Ask the on-call question before the rate question. It reorders everything.

The Shape That Actually Works
Almost every data org I would call healthy has landed on roughly the same structure, whether or not they arrived at it on purpose.
A small full-time core that owns meaning and uptime. Usually that is an analytics engineer or two who own the semantic layer, one platform engineer who owns the pipelines that page, and whoever runs governance. Three to five people at a mid-size company. Then contract capacity that flexes around the spikes, which is where contract staffing earns its premium honestly, because you are buying speed and elasticity rather than trying to buy a permanent function at an hourly rate. That core-plus-flex shape is the one our guide to budgeting a blended team prices out across functions.
The core is small on purpose. Not cheap, small. Those seats should be your best-paid and most stable people, and if you are choosing where to spend a direct hire budget, spend it there rather than on the fourth pipeline engineer.
Contract-to-hire deserves a mention because it fits this function unusually well. Data hiring has an evaluation problem, which is that a great interview performance and a great first ninety days correlate weakly. Working with someone inside your actual stack for three months tells you more than any loop will. We have written separately about what good contract-to-hire conversion rates look like, and what it comes down to is that teams who plan for conversion from day one tend to get it, while teams treating C2H as a hedge mostly do not.
One caution, since I have watched this go wrong. Contract-to-hire only works on layers where you would have hired full-time anyway. Running it on the platform-migration layer just means you have offered a permanent job to someone hired for temporary work, and either they accept a role that will not exist in a year or they decline and you have soured the last three months of the engagement.
Questions Data Leaders Bring Us
Our rate comparison came out in favor of contract, so what did it leave out?
The compute line. Contract rate comparisons almost always stop at the invoice, but a data engineer’s work keeps consuming warehouse credits after the engagement ends, and nobody owns tuning it. Add a cost review at exit and the math changes. On a short engagement the difference is small. Past about six months it stops being small, and I have seen a single unowned scheduled job outrank most of a team’s actual workload on the monthly bill.
Which data roles should never be contract?
Three of them, and they are not the ones most people guess. Anyone who owns metric definitions, anyone carrying production on-call, and anyone accountable for data governance or audit response. Those three share a trait, which is that their real output is judgment and continuity rather than artifacts, and neither of those transfers in a handoff document no matter how carefully it is written.
How bad is it that our data team is almost entirely contractors right now?
Recoverable, usually, and more common than you would think. The urgent question is not the ratio, it is whether any full-time person can currently explain why your top five metrics are calculated the way they are. If someone can, you have time to fix the structure deliberately. If nobody can, that is the first hire, and it is an analytics engineer rather than another pipeline engineer. Make that hire before you renew a single SOW.
How long before a data contractor is actually productive?
Four to eight weeks in a regulated environment, one to two in an unregulated one, and access provisioning is nearly all of that gap. Build the timeline around the security review rather than the start date, because the start date is the number that gets quoted and the review is the number that governs.
Can a contractor own the dbt project?
Yes, with one condition. A full-time person reviews and approves every model that touches a metric leadership sees. Contractors can write most of your transformation code perfectly well, and many write it better than the team they are supporting. What they should not hold is unilateral authority over definitions, because that is the one thing you cannot re-derive from the repo after they are gone.
We have budget for exactly one data hire. Contract or full-time?
Full-time, almost every time. A single hire is by definition your core, which means they own meaning and uptime, and both of those require someone who is still there next year. Contract makes sense as a first move only when the work is unambiguously a bounded migration with a cutover date and a business that can wait afterward. That is rarer than the pitch deck suggests.
How I Would Actually Decide It
Write your five layers down. Ingestion and platform, transformation and semantic, analytics, ML, governance. For each one, ask what remains when the person leaves and who gets paged when it breaks at 3 a.m.
Layers where the answer is “working code” and “somebody else” can go contract. Layers where the answer is “an understanding nobody wrote down” and “this person” cannot, at any bill rate, in any market.
That test takes ten minutes and it will contradict your current org chart. Most do.
If you want an outside read on which of your layers are staffed the wrong way, or you are sizing a migration and want to know what the contract market actually costs right now, talk to our data staffing team. If the answer for a layer is contract capacity for a window, the capacity-by-the-month model for data teams has the block shapes and the crossover math. We have been placing data engineers and analytics talent across 30-plus U.S. metros since 2005, in both models, and we will tell you when the answer is a full-time hire we do not get paid to place. Getting the layer right before anyone interviews is a decent part of why our placements are still there at 92 percent a year in.

