BackStartups

By sector

AI company diligence

What is proprietary and what is a wrapper, gross margins that depend on someone else's pricing, and evaluation that means something.

Realistic time Eight to fifteen hours

The distinctive difficulty in assessing an AI company is that the demonstration is almost always impressive and almost never decisive. Foundation models make it possible to build something that looks remarkable in a week, which means the demonstration establishes very little about whether there is a defensible business behind it.

The questions that do discriminate are about what the company owns. Proprietary data that others cannot obtain, a workflow that is genuinely embedded in how customers operate, a distribution advantage, or accumulated evaluation and fine-tuning that materially outperforms a general model — these are assets. A prompt and an interface are not.

The second distinctive question is gross margin. Companies whose product calls a third-party model have a cost of revenue set by another company's pricing. That pricing has moved substantially and can move again in either direction, and a business with a thin margin at current prices has a cost structure it does not control.

01

What is actually proprietary

Ask the question plainly: what would a competent team with the same model access need in order to replicate this, and how long would it take. Founders with a real answer give a specific one — a dataset, a set of integrations, an evaluation suite built over two years, a regulatory approval, a customer relationship.

Proprietary data is the most cited advantage and the least often real. The tests are whether the data is genuinely unavailable elsewhere, whether the company has the right to use it for training, and whether it improves performance measurably. Data collected from customers usually comes with contractual limits on how it may be used, and those limits are worth reading.

Workflow embedding is underrated and frequently more durable than a model advantage. A product that sits inside a regulated process, holds the system of record, or is the place where a team does its work is hard to displace regardless of what any model can do.

Check

  • What would a competent team need to replicate this, and how long would it take?
  • Is any dataset genuinely unavailable to competitors?
  • Do customer contracts permit using their data for model training?
  • Is the product embedded in a workflow, or is it a tool alongside one?
  • What accumulates over time that a new entrant would not have?

02

Evaluation and reliability

Ask how the company knows its output is good. Serious teams have an evaluation suite, run it on every change, and can show performance over time. Teams without one are relying on impressions, and impressions do not survive contact with production traffic.

Establish the failure modes and how they are handled. Every model produces wrong outputs; what matters is whether the product detects them, whether a human is in the loop where the stakes require one, and what happens when it is wrong in front of a customer.

In regulated or high-stakes domains, ask about auditability. Can the company explain why a particular output was produced, and retain the records to demonstrate it? This is increasingly a procurement requirement rather than a nicety.

Check

  • Is there an evaluation suite, and can you see performance over time?
  • What are the known failure modes, and how are they handled?
  • Where is a human in the loop, and where is there not one?
  • Can outputs be audited and explained after the fact?
  • What happens when the underlying model is updated by its provider?

03

Unit economics and model dependency

Get the gross margin, and get the cost of inference per unit of value delivered. A product priced per seat with unlimited usage and a per-token cost has an exposure that grows with engagement — the most successful customers are the least profitable.

Ask what happens to margin if model pricing rises, and whether the company could move to a different provider or a self-hosted model. Portability is a real asset and takes work to maintain.

Then ask what the company does that justifies the price relative to the model it calls. A meaningful gap between what a customer pays and what the underlying inference costs must be explained by something the company adds.

Check

  • Gross margin, and inference cost as a proportion of revenue.
  • How does cost scale with customer usage under the current pricing model?
  • What happens to margin if model prices rise materially?
  • How portable is the product between model providers?
  • Are the heaviest users profitable?

04

Legal and data questions

Training data provenance has become a live commercial risk rather than a theoretical one. Ask what the models were trained or fine-tuned on, whether the company has rights to that data, and what indemnities it gives customers and receives from its providers.

Then ask about customer data: whether it is used for training, whether customers have agreed to that, and how it is segregated. Enterprise procurement asks these questions and a company without clear answers will lose deals for reasons unrelated to product quality.

Check

  • Provenance of any training or fine-tuning data, and rights to use it.
  • What indemnities does the company give customers, and receive from providers?
  • Is customer data used for training, and have customers agreed?
  • How is customer data segregated between accounts?
  • Which regulatory regimes apply to this use case, and where?

Stop and think

Red flags

  • An impressive demonstration with no answer to what a competitor would need to replicate it.
  • Proprietary data claimed but not genuinely exclusive, or used without contractual rights.
  • No evaluation suite, and no way to show output quality over time.
  • Gross margins that only work at current model pricing, with no portability.
  • Per-seat pricing with unlimited usage and a per-token cost structure.
  • Customer data used for training without a clear contractual basis.
  • A product that is a thin interface over a general model, described as a platform.

Take these into the room

Questions to ask the founders

  1. What would a good team with the same model access need to replicate this?
  2. How do you know your output is good? Show me the evaluation.
  3. What is your inference cost as a share of revenue?
  4. What happens to your margin if model prices double?
  5. Are your heaviest users profitable?
  6. What data are you training on, and what rights do you have to it?
  7. What breaks when your model provider ships an update?

AI company diligence: common questions

How do I tell a real AI company from a wrapper?
Ask what a competent team with the same model access would need to replicate it, and how long that would take. A specific answer — a dataset, an evaluation suite built over years, deep workflow integration, a regulatory approval — describes an asset. A vague answer about product quality usually describes a wrapper.
Why do AI gross margins matter more than in other software?
Because the cost of revenue is set by a third party and scales with usage. Traditional software has near-zero marginal cost; a product calling a model pays per use, so the most engaged customers can be the least profitable, and a provider's pricing change moves the whole margin structure.
Is proprietary data a genuine moat?
Sometimes, and it is claimed far more often than it is real. The tests are whether the data is genuinely unobtainable elsewhere, whether the company has the contractual right to train on it, and whether it produces a measurable performance advantage. Many companies fail the second test without realising it.