Counted, not rated 7 links No form

AI Readiness Assessment: The Mid-Market Operations Scorecard

Seven links an AI decision travels before production, and the one where yours stops.

Operations and finance leaders reviewing a printed AI readiness assessment before moving a pilot into production

Last updated: August 7, 2026 · KORE1 AI & Data practice

An AI readiness assessment scores whether your operation can carry an AI decision, not whether your model works. This one counts seven links a decision travels, from input to owner. Your readiness is where the chain stops.

The pilot is not the hard part anymore. A capable analyst can get a useful model running against your own data inside a fortnight, and mid-market ops teams have been proving that to themselves for two years now. Then the demo ends and the thing sits there, working, touching nothing.

MIT’s Project NANDA looked at 300 public deployments alongside 150 leader interviews and found roughly 95% of generative AI pilots producing no measurable P&L impact. Not because the models were bad. Because nothing downstream of the model changed.

So this scorecard skips your model entirely. It walks one decision, a real one you would actually automate, through seven links of your operation and asks a countable question at each. No scale from one to five. No maturity tier. You get a number between 1 and 7, which is the first link that fails, and the role that closes it. When that role turns out to be a hire, our AI recruiters fill it off the same IT staffing bench, across 30-plus U.S. metros.

Operations manager tracing a single order-to-cash decision on paper during an AI readiness assessment
The rule

A Score Is Not a Start Date

Most readiness assessments hand you a number out of 100 and a quadrant. Useful for a board slide. Useless on Monday, because 68 out of 100 does not tell you what to do next or when you can start.

Chains don’t work that way and neither do operations. A decision that reads clean data, gets scored properly and has a written confidence limit still does nothing at all if the output lands in a spreadsheet that somebody re-keys into the ERP by hand. You didn’t get 3 out of 7. You got stopped at 4, and links 5 through 7 are hypothetical until you fix it.

Gartner puts the consequence in dollars. Their February 2025 research predicts organizations will abandon 60% of AI projects through 2026 for want of AI-ready data, at an average of $7.2 million per abandoned initiative. Mid-market budgets don’t absorb that twice.

Find the first break. Fix that one. Count again.

The scorecard

Walk One Decision Through Seven Links

Pick one decision. Invoice coding, lead routing, returns triage, shift scheduling, whatever your team argues about most. Then work the links in order, because each one is built on the one before it. Stop at the first failure and stop honestly.

L1

Input, and how many places it hides in

Ask
Sit with the person who makes this decision by hand today and watch them make it once, start to finish, without helping.
Count
Separate places they look before they decide. Email threads and the notebook in the drawer both count.
Stops at
3 or more
What your count means

Three sources is the line, and almost nobody guesses their own number correctly.

Watch, don’t ask. People describe the tidy version of their job and leave out the two lookups they stopped noticing years ago, which is exactly how a model ships without the context that made the human right. Under three, an integration handles it. At three and up you have a data problem wearing an AI problem’s clothes, and no amount of prompt work or model selection touches it, so the role you need first is a data engineer rather than anybody with the word AI in their title.

L2

Proof that somebody scored the output

Ask
Take 100 decisions the model has already made and have the person who owns that job today mark each one right or wrong.
Count
How many it got right, graded by a human who does the work.
Stops at
No graded set
What your count means

The number matters less than whether the number exists.

Plenty of teams get to link 2 and discover their model is right 91% of the time, which sounds fine until you ask what the nine cost, and the answer is a credit memo and a phone call from a customer who has been with you since 2011. Grade before you trust. The role here is an AI and machine learning engineer who builds evaluation as a habit rather than a milestone, or a decision scientist when the real tradeoff is commercial and somebody has to price what a wrong answer costs you.

L3

Limit, written down, agreed by three people

Ask
Ask three stakeholders, separately, how confident the model has to be before it acts without a human looking.
Count
Distinct answers you get back.
Stops at
Anything but one
What your count means

Two answers is worse than none. None is at least honest.

An unwritten limit gets set by whoever is most nervous in the room, usually downward, and a model throttled to review-everything is a very expensive way to keep doing the job by hand. Write one number in one sentence and put a person’s name beside it. The role is an AI product manager, and this is the link where mid-market teams most often decide they can absorb the work into an existing job, then find out four months later that nobody has actually made the call.

L4

Write, into the actual system of record

Ask
Trace one output from the moment the model produces it to the moment your ERP, CRM or WMS actually changes.
Count
Manual steps in between. Copy, paste, export and approve-by-email all count as one each.
Stops at
1 or more
What your count means

One manual step. That’s the whole threshold, and it is deliberately brutal.

A human sitting in the middle of the write is not a safety control, it is a queue with a person’s calendar attached to it, and every hour that queue waits is an hour the automation is not actually automating anything. This is where the largest share of mid-market pilots stall out, because the model was a six-week project and the integration is a nine-month one nobody scoped. McKinsey’s State of AI survey reaches the same place from the other direction. Its November 2025 wave found just 39% of organizations reporting EBIT impact at the enterprise level, and across waves of that survey the attribute most associated with impact has been workflow redesign rather than model choice. The role is an API and integration architect, or a platform-side specialist like a NetSuite integration specialist where the write lands in the ERP.

L5

Exception, with a queue and a name on it

Ask
Ask who works the wrong answers on the Tuesday after a holiday weekend, when 340 of them are waiting.
Count
Named people with that queue written into their actual job description.
Stops at
Zero
What your count means

Everybody nods at this one in planning. Then nobody’s name goes in the box.

An exception queue with no owner fills for about three weeks and then somebody quietly turns the automation off, which shows up in your metrics as a successful pilot that ended, rather than as the operating failure it actually was. Staff the queue before you turn the model on, not after it backs up. On mid-market teams this usually sits with an ops lead who already knows the process plus an MLOps engineer who can see the failures building in the monitoring a week before the queue makes them somebody’s Tuesday problem.

L6

Trail you can reconstruct nine months later

Ask
Pick one decision the system made last quarter. Ask for the inputs it read, the model version, and why it chose what it chose.
Count
Minutes to produce all three.
Stops at
Over a day, or never
What your count means

Your auditor will ask this. So will a customer’s lawyer, eventually.

Reconstruction gets impossible fast once models are retrained, and a team that cannot answer the question at nine months could usually have answered it at nine days if anyone had written the logging. Cheap early, expensive later, invisible in between, which is why it gets skipped. The role is an ML platform engineer who treats logging as part of the build instead of a follow-up ticket, working alongside a data governance analyst anywhere the decision touches money, health records or hiring.

L7

Owner, after the pilot team disbands

Ask
Name the person accountable for this decision twelve months from now, when the consultants are gone and the champion has been promoted.
Count
Names on your own payroll.
Stops at
Zero, or a vendor
What your count means

A vendor’s name in this box is a zero. Write it down anyway, because seeing it written is the point.

Ownership is the link that decays rather than breaks, so it passes at go-live and fails eighteen months later when the one person who understood the confidence limit takes a job somewhere warmer and nobody can explain why the threshold is 0.86. At mid-market scale this is rarely a full CIO. It’s a fractional CIO, or a chief AI officer once the portfolio is past two or three live decisions.

Stopping at 4 on a first pass is normal rather than embarrassing, it is roughly where most mid-market operations land, and the teams that eventually get a decision into production are usually the ones that wrote the number down instead of debating it. The counting is what keeps this from becoming a meeting about whose department is behind, because a link number is harder to argue with than an opinion about maturity.

The verdict

What Your Stop Link Means

Stops at 1–2 Data problem

Your model isn’t the blocker and buying a better one won’t move you. Fix inputs and evaluation first.

Stops at 3–4 Authority problem

Nobody has written down what it may do alone, or the write is still a human copying a value across.

Stops at 5–6 Operating problem

You can launch it. You can’t keep it. Exceptions and audit trail are what make it survive year two.

Reaches 7 Ready for one decision

Put exactly one in production, run it for a quarter, then count the chain again for the second.

The seven links and the bands are KORE1’s own framework, built from the AI and data searches we run alongside live implementations. Our placement numbers behind it are a 17-day average time-to-hire on IT roles and 92% retention at twelve months.

Operations director and integration architect agreeing who owns an AI decision in production
The hiring lever

When the Break Is a Seat, Not a Sprint

Five of the seven links fail for one reason. Nobody owns the thing.

That’s uncomfortable to hear from a staffing firm, so treat it as a claim to test rather than advice to take. Go back through your seven counts and mark which ones would be fixed by better software. Across the AI and data searches we run alongside live builds, the honest answer is usually one link, occasionally two, and the remaining five turn out to be a person who does not currently exist anywhere on your org chart.

Mid-market AI hiring splits into two clocks. Contract and project work covers the build window, where you need an integration architect or an ML engineer intensely for four months and then not at all, and contract staffing exists for exactly that shape. Permanent hires cover everything after, because links 5, 6 and 7 are not project work. They’re a job.

Nobody needs a fourteen-person AI function at 400 employees. What tends to work is one strong generative AI engineer or machine learning engineer on staff, an ops owner who already knows the process, and specialists brought in per link through IT staffing as the chain extends. We’ve been placing that shape for 20-plus years, most of it before anyone called it AI.

This week

Three Things to Do Before the Next Vendor Demo

None of these need budget, a platform decision or a steering committee. Each one is an afternoon.

Watch

Shadow one decision

Sit with the person who makes it today and count every place they look before they commit to an answer.

Write

Set one confidence limit

Put a single number in a single sentence, then get three stakeholders to say it back the same way.

Name

Staff the exception queue

Write a real person against the wrong answers before launch, and put the hours in their week.

Mid-market operations team running an AI readiness scorecard on paper before a vendor demonstration
Timing

Run It Before the Pilot, Not After the Invoice

When you count changes what the count is worth. Run the seven links before you scope anything and they shape the requirements, the vendor questions and the budget line for the integration nobody has priced yet. Run them after the pilot lands and they’re a list of things you’re now paying somebody to discover on your behalf.

There’s a second reason to go early, and it has to do with how demos work. A demo is built to be impressive, and an impressive demo makes links 4 through 7 feel like implementation detail, which is how a company with no named exception owner talks itself into an annual contract that quietly assumes one exists.

Count again at three points. Before you scope, at the end of the pilot, and 30 days before you widen it past one decision. Same seven questions each time. Numbers move earlier than confidence does, and having last quarter’s counts written down somewhere everybody can see them is what lets an ops leader push a date out by six weeks without the room quietly deciding they’re the pessimist.

Two companion diagnostics use the same counting method if your bottleneck sits elsewhere. The ERP readiness assessment covers implementation readiness, and the engineering velocity assessment covers whether your engineering org can ship what you decide.

Questions

Common Questions

What is an AI readiness assessment?

It’s a structured check of whether your operation, rather than your model, can carry an AI decision in production. This version counts seven links a single decision travels through, input, proof, limit, write, exception, trail and owner. Your readiness is the first link that fails, never an average across the seven.

How long should the scorecard take us to run?

About a week of calendar time and roughly five hours of real work. Shadowing one decision at link 1 is the longest piece at two hours or so. Grading 100 outputs at link 2 takes an afternoon if the model already exists. The rest are conversations you can hold between other meetings, and nobody has to block a sprint, buy a tool or tell a vendor you’re doing it.

Is anything here gated behind an email address?

No. Every link, threshold and band is on this page, with no form, no gate and no PDF to download. We publish it ungated because readiness conversations are reliably where staffing gaps surface, and filling those gaps is the business we’re actually in, whereas the scorecard itself is not something we would ever charge for.

We stopped at link 4. Is that bad?

Stopping at 4 is the most common first-pass result in mid-market operations, so read it as a starting position. Link 4 is the write into your system of record, and it’s the point where a six-week model meets a nine-month integration nobody scoped. Fix that one, count again in four to six weeks, then move on to link 5, and expect the second pass to go faster because the arguments about scope have already happened.

Can we run this before we’ve picked an AI vendor or platform?

Yes, and it’s the better time. None of the seven links depend on which model or vendor you choose. Platform affects who you hire and what the integration costs, not whether three people agree on a confidence limit or whether anyone works the exception queue. Counting first also sharpens the questions you take into demos, because a vendor answers “how does this write to our ERP” very differently once they know you’ve already traced the manual steps.

How is this different from a vendor’s AI readiness questionnaire?

Two differences. Every link here ends in something you count rather than a maturity level you rate yourself on, and every failed link names the role that closes it. Vendor questionnaires are scoped to the vendor’s own product, so they tend to skip straight past links 4 through 7, which happens to be where most of the real work sits and where the missing seat almost always turns out to be.

Does this work for agentic AI, not just a single model?

Yes, and the chain gets stricter rather than looser. An agent that takes multiple actions makes links 3, 5 and 6 harder, because the confidence limit now applies per step, exceptions arrive mid-sequence, and the audit trail has to capture a path instead of one answer. Count the seven links against the agent’s most consequential single action first, get that one clean, and only then widen the count to the full sequence the agent is allowed to run unattended.

Next step

Tell Us Where Your Chain Stopped.

Send us the link number and we’ll tell you whether it reads as a hire, a contract engagement or a process fix, in writing, before you commit to anything. No deck.