You're a couple of weeks from signature. Procurement is satisfied, the reference calls went well, and the demo was excellent. The demo is always excellent. It's the one artefact the vendor has rehearsed more than anything else they will ever show you.
A meaningful part of my work is pre-signature reads: an independent look at an AI system before someone commits serious money to it. The questions below are the ones that do the most work in that setting. None of them requires you to be technical. What they require is a willingness to keep asking until you get an answer a non-specialist can repeat to their board.
Here's what the demo cannot tell you, and how to find it out.
What the system actually is
1. Which parts of this are a language model, and which are conventional software?
"AI-powered" covers everything from a thin layer over a general-purpose model to years of trained models and careful engineering. The split determines your costs, your failure modes, and how defensible the product is against the next competitor with an API key. A good answer names the split without flinching: language models for the fuzzy judgement calls, a trained statistical model for the scoring, plain code for everything else, and a reason why each piece is where it is. An evasive answer is "it's all proprietary AI", which usually means either they don't want you to know how thin it is, or they don't know themselves. Both are findings.
Ask too where inference actually runs. If the model underneath belongs to a foundation-model provider, you have a second vendor in the room whose pricing, terms and roadmap you never negotiated.
2. What does the system do when it's wrong, and how would we know?
Every AI system is wrong some of the time; the question is whether wrongness is visible and managed or silent and compounding. A good answer includes error rates measured on data like yours, a human review step where the stakes justify one, and monitoring that would surface drift before your customers do. It also says who is responsible for detecting errors in production, you or them. An evasive answer is an accuracy figure with no denominator. "97% accurate" means nothing until you know on what data, measured by whom, and what the remaining 3% costs in your workflow.
3. Can we watch it fail?
Ask to run the system on your own messy, real-world data, and ask the vendor to show you a case it gets wrong. A confident vendor will do both, because they know where the edges are and can talk about them like adults. A vendor who insists on their own curated data, or claims there are no interesting failure cases, is telling you the edges are unmapped.
Do this properly. Bring a held-out set of your own examples that the vendor hasn't seen, and run it against a pinned model version rather than whatever their default endpoint happens to be serving that day. Otherwise you've measured a moving target and learned nothing you can hold them to. The demo shows you the system at its best; you're buying its average day.
4. What can't it do?
Sales processes are built to show capability, so ask for the boundaries directly: which use cases customers have tried that the product wasn't suited for, which languages, data types or domains it handles poorly, where the outputs need the most human correction. A vendor who answers with specifics has tested extensively and knows their own edges. A vendor who pivots back to capability claims hasn't, and you'll be the one who finds the edges in production.
The data underneath, and yours
5. What was the model trained on, and who checked that data?
This is the question I'd keep if I could keep only one. The training data sets the ceiling on everything the system will ever do for you, and it is the one component no demo can show. If the data doesn't resemble your cases, or nobody has checked it since it was assembled, the problems arrive months after go-live, when the system meets inputs it never saw. A good answer covers provenance, how the data was cleaned, and how often it's refreshed. "Large proprietary datasets" is not an answer; it's a label on a closed box you're about to build your process on.
6. Is our data used to train anything, and who else touches it?
You want three things in writing: whether your data trains their models and how to opt out, which subprocessors sit behind them, including whichever foundation-model provider is actually doing the inference, and what retention looks like when the contract ends. Add where the data is stored, processed and logged, and whether residency is configurable by region if that matters to you. A good answer names the subprocessors and puts the training opt-out in the contract, not the sales call. "Enterprise-grade security" offered as a complete answer is a phrase to write down and push on, because it's doing the work of a paragraph they'd rather not write.
7. What did the audit behind your certificate actually cover?
A compliance certificate says the vendor passed an audit at a point in time, against a scope somebody defined. It doesn't say the AI components were inside that scope. A controls audit of the hosting environment can leave the model pipeline, the prompt logs and the training data entirely unexamined. Ask for the scope, the date and what specifically was assessed, then ask what has changed since. Treat the certificate as the start of that conversation.
The team behind the demo
8. Who built the core system, and are they still there?
With smaller vendors especially, the product often lives substantially in one or two heads. That's normal at their stage, but it means your real dependency is on people, not software. A good answer has more than one person who can explain the system's guts, and documentation that would survive a departure. If every technical thread routes back to a single founder, price that risk into the deal, because their next funding round or acqui-hire is your outage.
9. What breaks at ten times the load?
"Does it run?" is the demo's question. "Can it carry version 1.0 for a customer like us?" is yours, and they're different questions. Systems that demo well are frequently held together by manual steps and one-off scripts that nobody mentions: analysis living in engineers' notebooks, results assembled by hand. A good answer is honest about what would need rebuilding as volume grows, and roughly when. "It scales" as a complete sentence is the evasive version; genuine engineering answers have trade-offs in them.
Make the question concrete with numbers. Ask for p99 latency under realistic concurrency rather than median latency from a single-request demo, and ask what happens when you exceed the rate limit for your tier. Then ask for pricing at ten and a hundred times your pilot volume, and for everything outside the headline price: fine-tuning compute, dedicated capacity, premium support, storage, retrieval infrastructure. A per-call price that looks reasonable in a pilot can be ruinous in production, and the vendor's proposal is written at pilot volume.
10. Who are your vendors?
Your vendor has vendors. If they run inference on a cloud provider and that provider has an outage, your system is down. If they build on a foundation-model provider and that provider changes its terms, deprecates a model or raises prices, your vendor's product changes with it, whether or not they planned for that. Ask them to name the dependency chain, say where the single points of failure are, and describe the last time an upstream problem reached their customers and what they did about it. A vendor who has never thought about this hasn't run at the scale they're selling.
11. Can we speak to a customer who left?
Reference calls are curated; you'll be introduced to the happiest customers on the list. The more informative conversation is with someone who chose not to renew. Ask what the most common reasons for leaving are, and whether they'll connect you with a former customer who'll speak freely. The reaction to the request tells you as much as the answer. Refusal is a data point. A vendor confident enough to make the introduction has earned a different level of trust.
The exit
12. If we leave in two years, what do we walk out with?
Assume the relationship ends, because most do eventually, and negotiate the divorce while everyone still likes each other. You want your data back in a usable format, plus your configurations, prompts, evaluation sets, logged outputs and anything trained or tuned on your data, with the format, the timeline and any portability fee written into the contract. A good vendor treats this as routine hygiene. A vendor who responds with some version of "why would you leave?" has told you what the lock-in strategy is.
The technical half of this is on your side of the table. Put an abstraction layer between your application code and the vendor's API from day one, so that switching providers is a configuration change rather than a rewrite. If the vendor's SDK makes that awkward, or you lose functionality by going through a standard interface, that's a cost you'll pay at every future migration, and it belongs in the price comparison now.
13. Can you change the underlying model without telling us?
Vendors swap foundation models routinely: for cost, for capability, because a provider deprecated something. Each swap can change the system's behaviour overnight, which matters enormously if the output feeds anything regulated, audited or customer-facing. A good answer includes advance notice of model changes, versioning, the ability to pin a version while you test its successor, and a stated deprecation policy with a notice period and migration support attached. Ask what happens to customers who can't migrate inside the window. If the answer is "they lose access", your contract needs a different clause. "We're always improving the models" is the evasive form: improvement you can't test isn't improvement, it's variance.
How to use this
Don't run it as an interrogation. Pick the six or seven that matter most for your context and put them in writing before the final commercial conversation, so the answers land in the contract rather than the meeting notes.
Decide how you'll score the answers before the questions go out, because after a good demo everything sounds reasonable. Four criteria do most of the work. Specificity: numbers and names, or generalities. Evidence: documentation, or promises. Completeness: did they answer what was asked, or something adjacent. Directness: did they address the question, or reframe it into one they preferred. Weight the criteria towards your hard constraints. If data residency is non-negotiable, a vague answer on where your data lives should disqualify, however strong the demo was. Writing the rubric down first is what stops a polished performance from overriding a serious gap in governance or terms.
And watch the texture of the responses as much as the content. The strongest signal in any of this is whether hard questions make the vendor more specific or more fluent. Specific is engineering. Fluent is sales. Good vendors enjoy these questions; they've done the work the questions are probing for, and it shows.
If the contract is large enough that a wrong answer would hurt, an independent technical read before signature costs a fraction of the contract it protects. That's the work I do at Agathon: independent, no implementation upsell, no stake in which vendor wins.
FAQ
What questions should I ask an AI vendor before signing? Start with what the system actually is (which parts are a language model, what happens when it's wrong, what it can't do), what it was trained on and what happens to your data, who built it and what breaks at scale, and what you walk out with if you leave. The full list above is thirteen questions; six or seven put in writing before the commercial close will do most of the work.
How do I evaluate AI vendors properly? Test on your own data rather than the vendor's. Bring a held-out set the vendor hasn't seen, run it against a pinned model version, and ask to see a case the system gets wrong. Score every written answer against a rubric you agreed before the questions went out: specificity, evidence, completeness and directness. The demo is the least informative part of the process.
What are the red flags when choosing an AI vendor? Accuracy figures with no denominator. "It's all proprietary AI" as the whole answer to how the system works. "Large proprietary datasets" as the whole answer to what it was trained on. Refusal to run on your data, to show a failure case or to introduce a customer who left. And the general pattern: hard questions that produce more fluency rather than more specifics.
How do I avoid AI vendor lock-in? Negotiate the exit before you sign: your data, prompts, configurations, evaluation sets and anything tuned on your data, returned in a usable format on a stated timeline. On your side, put an abstraction layer between your code and the vendor's API from the start, so switching providers is configuration rather than a rewrite. And know the vendor's own dependencies, because their lock-in becomes yours.
What should an AI vendor contract include? Beyond standard commercial terms: a training opt-out and named subprocessors; data residency and retention; advance notice of model changes, version pinning and a deprecation policy with migration support; pricing at production volume with the extras itemised; support SLAs with severity definitions and an escalation path; and exit provisions covering what is returned, in what format, by when and at what cost.


