Guide

AI Due Diligence Checklist: Is the Target's AI Real?

The full checklist, free and ungated: 24 questions across six assessment areas that show whether a target's AI is real, defensible and worth the price.

Most of what ranks for "AI due diligence" is software that promises to read your data room faster. This page does the other job. You are pricing an investment or an acquisition where part of the valuation rests on the target's AI, and you need to know whether that AI is real, whether it is defensible, and whether the margins survive scale. The checklist below is the full working document. No form, no gate.

I assess AI companies for investors, and this checklist is the shape of that assessment: six areas, twenty-four questions, each with what a good answer sounds like and what a bad one is telling you. The areas mirror how I run a full engagement, from technical reality through to commercial alignment. Put the questions to the target's management and technical team directly.

These questions reveal depth. A strong team answers them without flinching; a rehearsed one reaches for the deck.

The questions are designed to reveal depth, not to catch founders lying. Strong technical teams welcome this kind of scrutiny because it separates them from competitors trading on marketing. A team that treats the questions as hostile has told you something too.

How to score the answers

Rate each of the twenty-four answers on the same scale:

  • 3 for a strong answer with specific evidence
  • 2 for an adequate answer with some gaps or vagueness
  • 1 for a weak answer that raises significant concerns
  • 0 when they are unable or unwilling to answer

Out of a possible 72:

  • 60 to 72. Strong technical foundation. Proceed to deeper diligence with confidence.
  • 48 to 59. Moderate concerns. Specific areas need investigation before you proceed.
  • 36 to 47. Significant gaps. The technical claims may not support the current valuation.
  • Below 36. Fundamental concerns. Reconsider the thesis, or require substantial remediation as a condition.

The score is directional, not definitive. A single critical weakness, such as complete key-person dependency, can outweigh strong scores everywhere else. The pattern of the failures matters more than the total, and the final section of this page translates patterns into deal decisions.

1. Technical reality

Start with what was built. In almost every AI product the model is bought: an API call to a frontier lab, available to the target's competitors on identical terms. So "we use advanced AI" tells you nothing. These four questions establish what sits around the bought component.

1. Describe the architecture without analogies

Put it exactly like that: "Describe your model architecture without analogies. Which specific techniques are you using, and why those over the alternatives?" A strong team explains its choices precisely, names the techniques, acknowledges the trade-offs, and can say what it rejected and why. The bad answer retreats to buzzwords, invented terminology, or the pitch deck's metaphors. Technical depth and marketing fluency sound different within two minutes of this question.

2. Which layers are yours, and which are rented?

Ask them to walk the system layer by layer: which parts are calls to someone else's model, and which parts did they build? A good answer is unembarrassed about the bought components and precise about where they end. The bad answer is a "proprietary AI" claim that dissolves, under the walk, into a prompt and an interface. That finding is not automatically fatal (plenty of useful products are thin), but a wrapper priced as a platform is a reprice at best.

3. Show me the evaluation framework

"What is your current performance on the metrics you care about, and what are the known failure modes?" Mature AI teams obsess over measurement, so a good answer cites specific metrics without needing to check, acknowledges where the system fails, and describes the improvement plan. The bad answer deflects to customer anecdotes, claims "high accuracy" with nothing behind it, or reveals there is no systematic evaluation at all. A team with no eval cannot assess its own system, which means neither can you.

4. Run the demo on my inputs

Bring your own examples and ask for the demo live on them. A demo that only runs on curated inputs is a rehearsed performance, and training data that suspiciously overlaps the demo scenarios is one of the oldest tells in AI diligence. A good team runs your inputs, shows you a failure, and explains it. A bad one has reasons the demo cannot deviate today.

2. Defensibility and moat

The model is never the moat. When defensible value exists it sits in a specific layer of the surrounding system: proprietary data, domain-specific retrieval, an optimisation layer that improves with use, or encoded expert judgement no vendor sells. These questions locate that layer, or establish that it is not there.

The model is bought. The moat, if it exists, is whatever the target built that a vendor does not already sell.

5. What couldn't a same-model competitor replicate?

"What can your technology do that couldn't be replicated with foundation models and good prompt engineering?" A good answer names specific capabilities that require proprietary training, proprietary data, or unusual architecture. The bad answer describes value that is replicable off the shelf, with differentiation living only in the interface. As foundation models improve, that kind of advantage erodes on someone else's release schedule.

6. What would £20m and eighteen months buy a rival?

"If a competitor raised £20m specifically to replicate your data advantage, how long would it take them, and what would stop them?" Good answers identify specific barriers, admit a realistic timeline, and describe how the data advantage keeps growing. Bad answers are overconfident, cannot articulate a barrier, or describe "proprietary" data that turns out to be licensable by anyone.

7. Why is the retrieval and data layer built this way?

Ask the target's domain expert, not the CEO, and watch how the answer arrives. An instant, crisp answer means the knowledge is a codifiable fact, and a vendor has probably already encoded it. The pause is the opposite signal: the expert reaches, draws on years of doing the thing, and gives you a judgement rather than a rule. That hesitation marks tacit, accumulated knowledge, which is the closest thing to a real moat you will see in a management session.

An instant answer to why is a commodity. The pause is the moat.

8. How has the differentiation moved in eighteen months?

Static advantages erode and growing ones compound, so you want a trajectory, not a snapshot. A good answer shows clear evolution and a widening gap against alternatives. The bad answer is the same pitch the target gave eighteen months ago, with no measurable improvement in the core capability.

3. Scalability

The demo works at demo volume. This section asks what breaks at ten times the load, and who pays when it does, because AI costs scale in ways traditional software does not.

9. What are the unit economics at ten times current volume?

"What breaks first?" is the useful half of the question. Good answers show cost modelling beyond the current state, name the component that fails first, and point to architectural decisions made for scale. Bad answers assume costs scale linearly or promise to optimise later. Impressive current margins can collapse at scale, and post-close that collapse is yours.

10. What does one transaction cost, and who sets that price?

Ask for cost per inference at current volume, then at ten times. The target's main variable cost is set by a supplier it does not control, so a good answer includes the number and a hedge: portability across providers, caching, or cheaper models on the low-stakes paths. If they cannot answer, the unit economics are an open question and you are underwriting it blind.

11. Walk me through production monitoring

"What has your error rate done over the past 90 days?" separates production-quality systems from demo-quality ones, because real production systems have real metrics. Good: dashboards exist, error rates are tracked and improving, anomalies get caught by machinery rather than by customers. Bad: metrics only from test datasets, or "we're still building that".

12. What is in the technical debt backlog?

Every real system carries debt; mature teams know where theirs is. A good answer is an honest inventory with a prioritised plan that appears in the roadmap. The bad answer claims there is no technical debt, or the debt only surfaces as you probe. Hidden debt becomes your problem after close, usually on an eighteen-month fuse.

4. Team capability

If the moat is real, it usually lives in people rather than code, which makes the team section less about hiring quality and more about concentration risk. What you are buying might be people, structured as if it were software, and people can leave.

13. If the lead engineer left tomorrow, how long to recover?

"How long before someone else could retrain your models, and who else has done it?" Good: several people have run the process, it is documented, and the recovery timeline is credible. Bad: one person has ever done it, the documentation is that person, and the mitigation offered is "they'd never leave".

14. Show me the model training documentation

Documentation quality correlates with system robustness, and the request doubles as a transparency test. A good answer produces current documentation and offers it under NDA without hesitation. The bad answer is documentation permanently "in progress", or a process that exists only in people's heads.

15. What is the technical leadership succession plan?

You are investing in a company, not in two individuals. Good answers show explicit succession planning, knowledge transfer already happening, and technical depth beyond the founders. "We'll figure it out" concentrates the entire technical thesis in one or two resignations.

16. Ask the CEO and the engineers the same question, separately

What does the AI do, and why does it win? A gap between leadership's answer and the engineers' answer is the single most reliable washing tell available to you. Either the deck is ahead of the product, or leadership does not know where its own value lives. When the two answers describe different products, score this zero and treat every other claim with suspicion.

5. Data and governance

The risk here is quieter than the wrapper question but just as capable of moving price: where the data came from, where it goes, and what a regulator would make of the claims.

17. Where did the training data come from, and what rights do you hold?

A good answer documents provenance, names the licences and consents, and can produce the audit trail. The bad answer is scrape-and-hope, or rights that evaporate the moment they are tested against a representation and warranty. Data you cannot lawfully hold is not an asset, whatever the row count says.

18. What flows into third-party models, and under what terms?

Map the data flow: which customer or internal data crosses into a frontier provider's API, and what do the contracts say about it? Good answers show the flow is mapped, the terms are known, and the boundary is enforced technically rather than by policy document. The sharper signal is attitude. The more relaxed the target is about what goes into the system, the bigger the finding, because casualness about data exposure means nobody is imposing the boundary.

19. What happens when a customer wants their data out?

Deletion requests, data subject rights, and tenant separation all get harder once data has shaped a model. A good answer shows they have engineered for it: per-tenant isolation, retention policies, a tested deletion path. Blank looks here mean regulatory exposure is priced into the deal at zero, and it should not be.

20. Which public AI claims would survive a regulator's reading?

Set the target's public and investor-facing AI claims next to the architecture walk from question 2. Where the claims are material to the valuation, the distance between claim and system is disclosed risk, and you want the claim in writing with an indemnity to match. A good answer is a set of claims the engineers would recognise. A bad one gets vaguer the more technical your questions become.

6. Commercial alignment

The final area is the gap between what the technology can do and what the pitch deck promises, because that gap is where overpayment lives.

21. Which capabilities in the revenue plan exist today?

Ask them to sort the plan's revenue lines into shipped, in build, and imagined. A good answer sorts honestly and matches the engineers' account of the roadmap. The bad answer claims production-grade transformation while describing prototype-grade evidence. Calibrate against stage: a narrow, evidenced claim from an early product beats a sweeping claim from any product.

22. How deeply do customers use the AI, and how do you know?

Seats deployed is not adoption. If customers use the product shallowly (summarise this, draft that), expansion revenue is sitting on shelfware, and healthy-looking retention can mask a base quietly reclassifying the product from transformation to cost of doing business. A good answer shows instrumented usage depth inside real workflows. A bad one offers licence counts as evidence.

23. What happens if the provider ships your feature, or raises prices five-fold?

This is existential platform risk for any product built entirely on a third-party API. Good answers include a contingency plan, portability across providers, or a genuinely proprietary layer underneath that survives the model market moving. The bad answer is "that won't happen". A feature that is the whole product today can become a checkbox in the next model release.

24. Who accrues the value the AI creates?

Value created inside a vendor's roadmap accrues to the vendor. If the target's AI advantage is a capability the model providers improve for free, ask what the target owns as the frontier moves. A good answer points at a layer that compounds to the target: data that grows with use, workflow position, distribution. A bad one describes a margin story that belongs to the supplier.

From score to decision

The total puts you in a band; the pattern tells you what to do.

  • Strong across all six areas. Invest, at a price that reflects the verified moat rather than the deck's version of it.
  • Strong technology, weak commercial alignment. Reprice. The system is real but the revenue plan is ahead of it, and the gap is a negotiation input, not a footnote.
  • Real value, concentrated in a few unretained people. Restructure. Retention packages, earnouts, and key-person terms that hold the moat in place past close.
  • Failures clustered in technical reality and defensibility. Pass, or reprice to the software-margin multiple a thin wrapper deserves.
  • A perception gap, a demo that cannot deviate, and no definition of good output. Walk. Each alone is a concern; together they mean the AI claim is unfalsifiable and the valuation is resting on it.

One tell belongs to your side of the table rather than the target's. If the deal team treats diligence as an obstacle to a decision already made, the checklist will not save the deal, because nobody is going to act on a score they did not want.

Take the checklist with you

The twelve-question core of this page is available as a PDF you can take into the room, and the full AI due diligence service exists for the deals where the number turns on whether the AI is real and no one in-house can make that call. This page gives you the questions. Verifying the answers against the code, the data room, and the eval results is the part that takes an independent technical read.

Related insights: The investor's technical due diligence playbook, the wider read this checklist compresses; and Thin wrapper or true AI?, the real-vs-rented question argued in full.

Who stands behind these guides

Dr Colin Kelly holds a PhD in Natural Language Processing from Cambridge and read Mathematics and Computer Science at Oxford, with technical depth in AI that predates LLMs by a decade across IBM, PA Consulting and two VC-backed AI scale-ups. He is fully independent: no implementation upsell, no vendor affiliation, no dog in the fight. Which is why he can say of every engagement -- “I’ll stand behind every finding.”

Weighing a live deal? Agathon provides independent AI due diligence -- a written assessment you can put in the deal file and stand behind.

Or see what an assessment involves.

Talk about your deal

Common questions

What should an AI due diligence checklist cover?

Six areas: technical reality (is it genuine AI or a wrapper), defensibility (what a funded competitor could not replicate), scalability (what breaks at ten times the volume), team (where the knowledge lives and whether it stays), data and governance (provenance, privacy, and what flows into third-party models), and commercial alignment (whether the tech supports the revenue plan). Most published checklists cover legal and procurement risk only.

Is this about using AI to run due diligence?

No. Most pages ranking for AI due diligence sell software that reads a data room faster. This checklist does the other job, assessing whether a target company's AI is real, defensible, and priced correctly. The tool question and the target question share a search term and nothing else.

How do you evaluate an AI startup before investing?

Work layer by layer. The model is almost always bought, so a claim of proprietary AI tells you nothing on its own. Locate the one layer a same-model competitor could not rebuild in a quarter (data, retrieval, or encoded expert judgement), test the unit economics at ten times current volume, and check whether the people who hold the tacit knowledge are retained past close.

Can I run this checklist without a technical expert?

Most of it, yes. The questions are phrased for an investor to put to management, and the good and bad answers are described so you can score them yourself. Where you need independent expertise is verification, confirming that the architecture walk, the eval results, and the data-room evidence match what was said in the room.

What is the single biggest red flag in AI due diligence?

The perception gap. Ask the CEO what the AI does and why it wins, then ask the engineers the same question separately. When the answers describe different products, either the deck is ahead of the product or leadership does not know where its own value lives. Both mean the AI claim driving the valuation is not anchored to what was built.

How long does technical due diligence on an AI company take?

A focused engagement runs two to six weeks, scoped to the deal. The core arc is scope definition, a technical deep-dive (code, data room, and technical interviews), synthesis against the investment thesis, then a written report and live debrief. This checklist compresses the question set from that process into a form you can start using in the next management session.

Get AI insights in your inbox
Practical analysis on AI strategy, products, and technical leadership
No more than one newsletter a month