Last updated: August 6, 2026
By Kris Drouet, Engineering Executive, in partnership with KORE1
Your best engineers cannot reliably evaluate an AI vendor demo, because the demo is a controlled performance and their own judgment about AI is measurably unreliable. In a 2025 randomized trial, experienced developers ran 19% slower with AI tools while believing they had run 20% faster. Sending your strongest engineer into the room does not solve that problem. It hides it.
I used to think the fix here was simple. Send the best engineer.
Put the person with the deepest knowledge of your systems in the room, let the vendor talk for fifty minutes, then ask her what she thought on the walk back. Her read was the verdict. I would have defended that process in front of any board, and for a while I did, because the alternative was letting a procurement team grade a machine learning system on a rubric originally written for a CRM migration.
Then I started writing down the verdicts.
Not formally. A note in my own file after each vendor evaluation, one line, what my technical reviewer said in the hallway and what actually happened twelve to eighteen months later. I have maybe two dozen of those lines now. The hallway verdict and the eventual outcome agree less often than a coin flip would, and the thing that finally bothered me was not the miss rate. It was that the confidence never moved. My reviewers were exactly as certain in the cases they got wrong. Every single time.
I have argued the underlying call elsewhere, in the case for building versus buying AI capability, and the four-gate checklist covers the diligence you run between a yes and a signature. This one is about something narrower and more uncomfortable. It is about the fifty minutes themselves. And why the smartest person on your payroll is not a reliable instrument inside them.

A Demo Is a Performance. An Evaluation Is an Experiment.
AI vendor evaluation is the practice of testing a vendor’s system against your own data, your own volume, and your own failure conditions, rather than judging it by a prepared demonstration. A demo shows you a model on curated inputs under conditions the seller chose. An evaluation puts the same model somewhere it can lose.
Those are not two grades of the same activity. They are different activities. Nobody watching a demo is running an experiment, because an experiment requires a result the presenter did not pick, and in a vendor demo every input on the screen was selected by someone whose commission depends on how the next forty minutes go.
Which is fine. That is what a demo is for. Vendors sell. The failure is not that they present well. It is that we keep accepting a sales artifact as evidence and then acting surprised at renewal.
Your Engineers Already Failed This Test on Their Own Codebase
Here is the finding that changed how I staff these meetings.
METR, an independent research nonprofit, ran a randomized controlled trial in early 2025 with 16 experienced open-source developers working on repositories they already maintained, averaging more than 22,000 GitHub stars and over a million lines of code. Real issues from their own projects, 246 of them, about two hours each. Half the tasks allowed modern AI tooling, mostly Cursor Pro with Claude 3.5 and 3.7 Sonnet. Half did not.
Before starting, the developers predicted AI would speed them up by 24%. After finishing, they reported it had sped them up by 20%. The measured result was a 19% slowdown.
Read that gap again. These were not junior engineers being dazzled. They were maintainers, on code they knew better than anyone alive, using tools they had spent tens or hundreds of hours with, and they were wrong about the direction of the effect by roughly forty points. Not the magnitude. The direction.
METR is careful about what this does and does not prove, and I will be too. Sixteen is a small sample. Mature open-source repos with high standards are a hard setting for AI assistance, and the authors say plainly that their result does not establish that AI fails to speed up most developers, in most settings, or that it will still hold as models improve. Fair. All true.
But none of those caveats touch the part that matters for procurement. The developers were not just wrong about whether the tool helped. They could not perceive that they were wrong, with the work in front of them and the clock running. Now take that same instrument, give it fifty minutes, no access to the model weights, no error logs, no p95 latency, and a salesperson steering, and ask it for a yes or no on a seven-figure commitment.
That is the trap. It is not that your engineers are gullible. It is that the specific perceptual skill this asks for is one nobody has, and expertise does not supply it.
The Numbers on the Vendor’s Slide Are Partly Memorized
Say your engineer is disciplined and does the right thing. She ignores the polish and asks for benchmarks.
Reasonable. Also not as safe as it sounds.
A 2025 study called LessLeak-Bench compared roughly 1.7 trillion data pairs between model pre-training corpora and 83 software engineering benchmarks, checking how much of the test set had already been swallowed during training. The averages look mild. Python benchmarks leaked 4.8% on average, Java 2.8%, C and C++ 0.7%. The individual numbers are where it gets interesting. QuixBugs came in at 100%. All of it. BigCloneBench at 55.7%. SWE-bench Verified, the benchmark half the industry now quotes in pitch decks, at 10.6%.
Then the part that should worry a buyer. On the APPS benchmark, StarCoder-7B scored 4.9 times higher on the leaked samples than on the clean ones. Same model. Same benchmark. Nearly a fivefold difference depending on whether the question had been seen before.
Leaderboards have their own version of this. A team of researchers looked at Chatbot Arena and documented that a handful of large providers test private variants and disclose only the ones that score well, including 27 private variants tested by Meta in the run-up to the Llama-4 release. The published number is not the model’s performance. It is the best of an undisclosed number of attempts, and you are not told what the denominator was.
So when the deck says a benchmark figure, the honest translation is: this number was produced under conditions I cannot see, on questions the model may have already read, selected from a set of runs I do not know the size of. That is not fraud. Every vendor in the category does it, and most of them are not even being cynical about it. It just is not evidence. Not about your data.

What the Demo Prices, and What Production Actually Bills You For
The gap between the first two columns is where the money goes.
| What the demo shows you | What production charges you for | Ask this in the room |
|---|---|---|
| Curated inputs the vendor chose | Your messy records, your legacy encodings, your nulls | “Run it on this file. Now, on screen.” |
| Median response time on a warm path | p95 and p99 during your month-end spike | “What is p95 at 40 requests per second?” |
| A benchmark score with no methodology | Accuracy on the 8% of cases that generate 80% of your escalations | “Show me the confusion matrix, not the headline.” |
| A clean happy path, start to finish | Silent wrong answers nobody catches for a week | “How do we know when it is wrong?” |
| Per-seat or per-month pricing | Token cost multiplied by real volume, plus retries | “Price this at our actual annual call volume.” |
| “Integrates with your stack” | Two engineers for a quarter, plus an auth model nobody scoped | “Name a customer who integrated in under 60 days.” |
Agent Washing and the 130 Number
Gartner put a name to the broader version of this in June 2025. They call it agent washing. The definition is rebranding an existing product as agentic AI without the underlying capability changing much, and the products getting rebranded are mostly assistants, robotic process automation, and plain chatbots. Their estimate at the time was that of the thousands of vendors selling agentic AI, only about 130 were genuinely doing it. They also predicted more than 40% of agentic AI projects will be canceled by the end of 2027, on escalating cost, unclear value, and thin risk controls.
MIT’s NANDA initiative found something adjacent from the buyer’s side, reviewing more than 300 publicly disclosed AI initiatives alongside interviews at 52 organizations and survey responses from 153 senior leaders. Roughly 95% of enterprise generative AI pilots produced no measurable return. Their conclusion was that the divide between the 5% and everyone else was not model quality. It was approach. Not the model.
I want to be careful not to turn that into cynicism, because cynicism is just as lazy as enthusiasm and it costs you the real tools. The 130 exist. Some of them are very good, and the teams that bought them early are enjoying an advantage that is going to be hard to catch. The problem is strictly one of telling them apart. A category where most of the sellers are exaggerating and a few are not is precisely the category where a fifty-minute performance tells you nothing, because the exaggerators demo just as well. Often better. They have more time to practice. Nothing to build.
Four Questions That Break the Spell
These go in order, and the order is the point. Each one is only useful because the previous one narrowed the room.
- “Run it on this, right now.” Bring a file. Not a clean one. Bring the export with the 2014 records where somebody put the address in the name field, because that file is your actual business and the vendor’s sample data is not. The answer you are grading is not the output. It is the ten seconds before the output, when you find out whether they can even accept your input without a scoping call.
- What is p95 latency at our peak volume, and what does a single call cost at that volume? Two numbers. Most reps do not carry them, which is not disqualifying, but the speed with which they get you a real answer afterward tells you roughly everything about what support will feel like in month eight.
- Show me it failing. Not a hypothetical edge case you have already patched. A live wrong answer, produced in front of me, and then walk me through how your system knew it was wrong. If nothing in the architecture can flag a bad output, you have not sold me a system. You have sold me a very confident intern with API access.
- The last one is a name, not a question. Give me a customer at roughly our size and complexity who has been in production more than a year, and let me talk to their engineering lead without you on the call. Vendors resist this last clause harder than any pricing conversation I have ever had. That resistance is itself the data.
None of that is clever. It is the difference between watching and testing. Twenty extra minutes in a meeting, and I have never once regretted them.

Why Your Strongest Engineer Is the Wrong Person to Send Alone
This is where people push back. Fair enough.
A great engineer watching a demo does something involuntary and mostly admirable. She builds the system in her head as she watches. The vendor gestures at retrieval, and she fills in a vector store, a reranker, a cache. He waves at error handling, and she supplies a retry policy and a dead letter queue, because that is what she would have built, and her version is coherent. Usually better than theirs.
Then she grades the demo against the thing she just imagined.
That is the whole failure, and it is invisible from the inside. The stronger the engineer, the more competently she patches the vendor’s gaps with her own expertise, and the better the product appears. A weaker reviewer would have simply noticed that nobody explained what happens on a malformed payload. Expertise is not neutral here. It actively fills holes, and you end up buying the architecture your own engineer built silently in her head during a slide about roadmap.
The fix is not a worse engineer. It is a second chair with one assigned job, which is to write down every question the demo did not answer, and to say those out loud at the end while the vendor is still in the room. Split the role. One person evaluates what was shown. One person inventories what was skipped. I have run it this way for about three years and the second chair has caught more real problems than I have.
And put the artifacts in writing before the meeting, not after. Under the hood, this is the same clarity problem I write about constantly, just relocated to a conference room with a vendor in it. Decisions made in a room and never written down drift, and by the third meeting nobody remembers whether the vendor actually committed to on-prem inference or merely nodded when someone said it. Write it down. In the room.
What Good Actually Looks Like
The good version of this fits on an index card.
You give the vendor a bounded, paid, two-week evaluation against a slice of your real data, with a written pass mark set before it starts. One metric. One threshold. One date. You do not move the mark afterward, which sounds obvious and is the single hardest discipline in this entire process, because by week two somebody senior will have fallen in love with the tool and will start explaining why 71% was really more like a pass if you think about it correctly.
Somebody has to be immune to that. Usually that means the person who can kill the deal is not the person whose budget or reputation is riding on it going forward, and if you do not have someone like that on staff, that is a hiring problem wearing a procurement disguise. Bringing that judgment in as a direct hire tends to beat renting it, because the value compounds across every vendor cycle after this one. It is also increasingly the thing I get asked to help with, and it is why KORE1’s AI and machine learning staffing practice has been fielding more requests for evaluation-side talent than build-side lately. The delivery data underneath all of this is collected in KORE1’s 2026 engineering velocity and AI-in-production benchmark report, which is worth a read before your next vendor cycle.
The other honest answer is that sometimes you should just buy it. Not every decision deserves this apparatus. If the tool is cheap, the exit is clean, and the blast radius is one team, buy the thing and find out. Save the full process for the systems that get wired into decisions you cannot easily unwire.
What People Ask Me After a Demo Goes Well
How is this different from evaluating any other software vendor?
Regular software behaves the same on Tuesday as it did in the demo. An AI system’s output depends on inputs it has never seen, so a demo on curated data tells you almost nothing about your data. That is the whole difference.
There is a second difference that shows up later. Traditional software fails loudly, with a stack trace and a status page. AI systems fail quietly, returning a confident, well-formatted, completely wrong answer that nobody catches until a customer does. Your evaluation has to test for the quiet failure, and almost nobody’s does.
Our engineers loved it. Is that a signal at all?
It is a signal about the interface, not the system. Engineer enthusiasm reliably tracks how good the developer experience feels in the first hour, which is genuinely worth something and is also the thing vendors optimize hardest for.
Treat it as one input among several and notice what it is actually measuring. If your team is excited, ask them to write down specifically what impressed them. Half the time the honest answer is that the setup was fast and the docs were clean. Both real. Neither predicts whether the thing survives your volume.
The vendor will not run it on our data during the demo. Fair or not?
Usually fair on the day, and unacceptable by the second meeting. Security review, sandbox provisioning, and a data processing agreement are legitimate reasons to say not right now.
What you are grading is what happens next. A vendor with a real product schedules that session before you have to ask twice. A vendor without one starts producing reasons, and the reasons will be individually plausible and will keep arriving. If you are three meetings in and still watching sample data, you have your answer, and it did not come from the demo.
Does a proof of concept solve this?
Only if you set the pass mark before it starts and refuse to move it. A POC without a written threshold is just a longer demo, and it is worse than a short one because it manufactures sunk cost and internal advocates.
I would rather run a tightly scoped two-week evaluation with one number attached than a three-month pilot with a vague sense that things are going well. Vagueness always resolves in the vendor’s favor. Always. That is not a knock on vendors. It is just what happens when the only person tracking the criteria is the person selling.
We already bought something on the strength of a demo. How bad is it?
Probably survivable, and renewal is your second window. Go run the evaluation you skipped, against real data, with a real threshold, and do it about ninety days before the renewal date so the results land while you still have leverage.
You will get one of two outcomes. Either the tool clears the bar and you renew with actual evidence instead of a feeling, which is a genuinely better place to be than you were. Or it does not clear the bar, and you now have a documented case for renegotiating or leaving, built on your own numbers rather than an argument about whether the demo was misleading. Both beat drifting into year three.
The Meeting Is Not the Evidence
I still send my best engineer to the demo. That never changed. What changed is that I stopped treating what she says afterward as a finding, and started treating it as a hypothesis with a test attached.
Because the uncomfortable thing in the METR data is not that AI slowed those developers down. Models improve. That number will move, and it may already have. The uncomfortable thing is the confidence. Sixteen expert engineers, on their own code, could not feel a 19% slowdown while it was happening to them. Nothing about a vendor conference room makes that instrument sharper. Show me the data instead.
Sitting in one of these cycles right now? Or quietly unwinding something you bought last year on the strength of a very good Tuesday afternoon? Message me. Kris Drouet on LinkedIn. I read those.
Different problem if what you are missing is a person on the payroll who can sit through a great demo and stay unmoved. That is a search, not a process fix, and I say that as someone who spent years trying to solve it with process. KORE1 fills AI and platform roles in more than 30 U.S. metros, and the recruiters there average fifteen years in engineering staffing. Start the conversation with a recruiter.

