Notes · Buying AI

How to score an AI vendor before the demo convinces you

Tyler Kent · September 3, 2026 · 4 minute read

An AI vendor demo is built to be believed. The sales team ran it many times before you saw it, on cases chosen because they work. None of that is dishonest. It is just not evidence about your Tuesday morning. The fix is to change the order of operations: write down what you need before you watch anything.

The matrix comes first

A requirements matrix is a one-page list of what the tool must do, drawn from the workflow itself and two or three interviews with the people who run it. Each line is ranked must, should, or nice. A prior-authorization tool for a medical group might have "reads faxed referrals as PDF" as a must, "pushes status into the EHR work queue" as a should, and "Spanish-language patient letters" as a nice.

Write it before any demo. Then score every vendor against the same lines. A vendor that misses a must is out, regardless of how the demo felt. A shortlist of eight usually drops to four in this step, which is the cheapest cut you will make all quarter.

Test the survivors on your own de-identified cases, not theirs. Five cases from last month, run through each tool, scored by the staff who would use it. If a vendor will not run your cases, that is a score too.

Five questions about your data

Every healthcare buyer should have written answers to these five before pricing is discussed.

  1. BAA. Will they sign a business associate agreement, and does it cover the AI features specifically? Some vendors' BAAs predate their AI features and exclude them.
  2. Subprocessors. Who else touches the data? Most AI vendors sit on top of a model provider such as OpenAI, Anthropic, Google, or Microsoft. Ask for the full list, the BAA chain down to that provider, and where the data is hosted.
  3. Retention. How long are your prompts and outputs stored, by the vendor and by their model provider? Zero-retention and 30-day options exist. If the answer is "it depends," treat it as a no until it isn't.
  4. Training on your data. Is your data used to train or improve their models, or the model provider's? The answer needs to be no, in the contract, not in a settings toggle an administrator can miss.
  5. Exit and export. When you leave, what do you get back, in what format, at what cost, and how long until your data is deleted? Ask for a deletion certificate in the contract.

A vendor that answers all five in writing within a week is a vendor you can work with. One that sends a slide is telling you something.

Build versus buy, in dollars

Sometimes the honest answer is to build. Here is the shape of the math. Every number below is invented for illustration.

Say the workflow is referral-packet assembly at a 200-bed hospital, and two case managers each spend about 10 hours a week on it.

Buy: $60,000 a year in license fees, $25,000 one-time integration, and 0.2 FTE of internal ownership at about $20,000 a year. Three-year cost: $180,000 plus $25,000 plus $60,000, or $265,000.

Build: a $45,000 fixed-price build, about $12,000 a year in hosting and model API costs, and 0.3 FTE of ownership at about $30,000 a year. The ownership number is higher than for buying because you own the fixes. Three-year cost: $45,000 plus $36,000 plus $90,000, or $171,000.

On these numbers, building saves about $94,000 over three years. That is not the end of the analysis. Building wins only when the workflow is narrow, stable, and has an internal owner by name. If the vendor's product covers five workflows you will actually use, or nobody on staff can own a build, buying wins even at the higher number.

One more check before either column. Value is what the tool frees at the constraint, not what it costs. Twenty case-manager hours a week are worth recovering only if those hours were the thing holding up discharges. An hour saved somewhere that was not the bottleneck is worth close to nothing, and no spreadsheet will tell you that on its own.

What "none of these" looks like

An independent evaluation sometimes ends with no purchase. Three signals.

There is a fourth outcome worth naming: the capability is about to be a feature of something you already own. Epic shipped native AI charting in February 2026, and any ambient-scribe evaluation that ignored that fact was scoring the wrong question.

The last cycle's warning was Olive AI, an automation vendor valued at $4 billion that shut down in 2023 after customers found the ROI figures were estimates. A memo that says "wait" or "build" is doing its job.

The order of operations

Matrix, then the five data questions, then your own cases, then the demo, then the math. Run it in that order and the demo becomes what it should have been all along: a check on whether the product matches its paper, not the thing that decides.

Sources: Olive AI shutdown (valued at $4 billion, wound down October 2023, ROI figures found to be estimates); Epic native AI charting launch, February 2026 (why "a feature of what you already own" is a real evaluation outcome); Initial claim denial rate 11.8% in 2024 (context for the denial-appeal example).