Two vendors quoted for the same AI project. Both were competent. Both proposals were, in hindsight, unbuildable as scoped — not because the vendors were dishonest, but because the scope let them describe success without defining it. It took five rewrites to produce a statement of work that made the promises testable. These are the seven clauses that did the work.
Why AI scopes fail differently from software scopes
In a conventional software project, acceptance is largely self-evident. The checkout flow either processes a payment or it does not. Disagreements happen, but they happen about priorities and timelines, not about whether something works.
AI projects break that assumption. The output is probabilistic, so “the model works” is a judgement rather than an observation. Without a threshold agreed in advance on data nobody could tune against, the last three weeks of the project become a negotiation in which both sides sincerely believe they are right. The vendor demonstrates strong results on the cases it developed against; you test it on the cases you care about and see something worse.
Every clause below exists to remove one of those ambiguities before it becomes expensive.
Clause 1 — Define acceptance on a frozen evaluation set you own
This is the clause that matters most, and it must be built before the engagement starts.
Assemble a sample of real cases — a few hundred is usually enough — and have your own team label the correct answer for each. Keep this set out of the shared workspace entirely. The vendor never sees it during development. The statement of work then reads, in substance: acceptance is measured on the client’s held-out evaluation set, and the target is a stated number.
The reason this works is not distrust. It is that any team optimising against a visible test set will overfit to it, without intending to and often without noticing. A number produced on data the vendor selected proves almost nothing about behaviour on data it has not seen. This is the same discipline as a frozen regression suite in conventional engineering, and it should feel equally unremarkable.
Setting the target honestly is the other half. “95 percent accuracy” sounds professional and is usually a number picked because it sounds professional. Ask instead what the current process achieves and what threshold would make the system useful. Frequently the honest answer is far lower than 95 percent, and stating it protects both sides.
Clause 2 — Separate data preparation from modelling
The most common source of overrun in AI outsourcing is not the model. It is the state of the data, which nobody can assess accurately at quoting time — including you.
Structure the engagement as two priced phases. Phase one is a data assessment: the vendor examines the actual data, reports on completeness, labelling quality, class balance and the volume genuinely usable, and delivers a written finding. Phase two, priced after that finding, is the modelling work.
Vendors sometimes resist because it delays the larger commitment. It is worth insisting, because the alternative is a fixed price quoted against assumptions that turn out to be wrong, followed by a change request that neither side budgeted for. A vendor comfortable with a paid discovery phase is generally a vendor that has been burned by this before, which is a good sign.
Clause 3 — State who owns the model, the weights and the prompts
Standard IP clauses were written for source code. AI engagements produce artefacts that do not obviously fall under that heading, and the ambiguity only surfaces when you try to leave.
Name them individually: fine-tuned model weights; training, validation and evaluation datasets, including anything derived from your data; prompt templates and system prompts; retrieval configurations and index-building code; and any internal tooling created to produce the above.
Prompts deserve particular attention because they are easy to overlook and disproportionately valuable. A well-tuned set of prompts and retrieval settings often represents most of the intellectual work in a modern AI system, while looking like a text file. If your contract lists “source code” and stops there, you may find that the artefact carrying the value is the one nobody assigned.
Clause 4 — Cap and attribute inference costs
An accuracy target with no cost constraint has an obvious optimisation: use the largest model available, call it several times per request, and hit the number. The vendor meets the contract. You inherit a per-transaction cost that makes the system unusable at volume.
Two provisions fix this. State a cost target per unit of work — per document processed, per conversation handled, per transaction classified — measured during acceptance alongside quality. And state clearly who pays for model API usage during development, because unattributed inference spend accumulates quietly across a multi-month engagement.
Expressing the target per unit rather than per month is what makes it durable. Volume will change; the unit economics are what determine whether the system survives contact with production.
Scoping an AI engagement this quarter?
We help Singapore employers define acceptance criteria, evaluate vendors and staff the in-house side of AI delivery.
Get startedClause 5 — Require a documented failure mode analysis
An aggregate accuracy figure hides the information you most need. A system that is 92 percent accurate overall but fails systematically on your highest-value customer segment is worse than one that is 88 percent accurate with evenly distributed errors — and the aggregate number will never tell you which one you have.
Require, as a named deliverable, a written analysis covering: performance broken down by the segments that matter to your business; the categories of error the system makes, with examples; the conditions under which performance degrades; and what the system does when it is uncertain.
That last point is the operational one. A system that fails loudly and hands off to a human is deployable. A system that fails silently and confidently is a liability, regardless of its headline score. Making this a deliverable rather than a request means it gets done before handover rather than discovered in production.
Clause 6 — Set data residency and access terms explicitly
For Singapore organisations, and particularly those in financial services or healthcare, this is a compliance requirement rather than a preference. State where data may be processed and stored, and name the acceptable jurisdictions.
Beyond location, specify: whether production data may be used for development at all, or only anonymised or synthetic derivatives; whether data may be sent to third-party model APIs, and which; whether it may be retained for provider-side training, which usually requires an explicit opt-out; who at the vendor has access, by role; and what happens to every copy at the end of the engagement, with a deletion deadline.
The third-party API question is the one most often left implicit and the one most likely to cause a problem, because it can route your data outside the arrangement you negotiated without anyone deciding to. The same discipline applies to any engagement handling regulated data, whether you are staffing it locally or through a team working on government contracts.
Clause 7 — Define what happens when the target is not met
AI projects can miss their target for legitimate reasons. The data may not support the task. The problem may be harder than either party estimated. A contract that assumes success and provides no path for a good-faith miss forces both sides into an adversarial position when the honest outcome would be a controlled stop.
Write the fork in advance. Define a remediation window — a fixed additional period at the vendor’s cost to close a gap. Define a partial acceptance path, where a lower threshold is agreed at a reduced fee if the result is still useful. And define a clean termination path stating what you receive if the project stops: the data assessment, the code produced, the evaluation results, and the documented reasons.
That last one matters more than it looks. A failed AI project that hands back a rigorous account of why the approach did not work has produced something genuinely valuable — it stops the next team repeating it. A failed project that hands back nothing has burned the budget twice.
What we deliberately left out
Two clauses that appear in most AI outsourcing templates are not in ours, and both omissions were deliberate.
A guaranteed accuracy figure with no evaluation set attached. A vendor who agrees to 95% accuracy without agreeing on the data it is measured against has promised nothing, and both sides will discover this at acceptance. The number is not the commitment; the evaluation set is. Once Clause 1 exists, a headline accuracy figure adds no protection and creates false comfort.
Blanket ownership of everything the vendor produces. It sounds prudent and it prices badly. A vendor asked to assign rights to internal tooling and general-purpose components will either refuse, or raise the price to cover the loss of reuse, or sign and quietly ignore it. Clause 3 works because it is specific about what actually matters to you — the model, the weights, the prompts and the evaluation data — and silent about the vendor’s general toolkit.
The wider principle is that a scope document is not stronger for being more demanding. It is stronger for being enforceable, and every unenforceable clause weakens the ones next to it by teaching both parties that the document is decorative.
What changed after the rewrite
With these clauses in the scope, both vendors revised their proposals. One lowered its accuracy commitment substantially and explained why, which was the more credible response. The other declined to quote against a held-out evaluation set at all.
That second reaction is worth treating as a result rather than a setback. A vendor unwilling to be measured on data it cannot see is telling you something useful for the price of a rewritten document. The five rewrites cost about a week of internal time; they were the cheapest part of the project by a wide margin. For teams building the in-house counterpart to an outsourced effort, the same rigour applies to how you select the vendor in the first place.
Frequently asked questions
Why do AI outsourcing projects go wrong more often than normal software projects?
Because acceptance is ambiguous. In conventional software, a feature either works or it does not, and both sides can see it. In an AI project the output is probabilistic, so “the model works” is a matter of interpretation unless a threshold was agreed in advance on data that nobody could tune against. Without that, the final weeks become a negotiation about whether the result is good enough, and both parties believe they are right.
What is a frozen evaluation set and why does it belong in the contract?
It is a sample of real cases, labelled by your own team, that the vendor never sees during development. It belongs in the contract because it is the only mechanism that makes an accuracy target meaningful. A target measured on data the vendor selected and tuned against proves very little. Build it before the engagement starts, keep it out of the shared workspace, and state in the statement of work that acceptance is measured on it.
Who should own the model and the prompts at the end of an engagement?
You should, and it needs saying explicitly because it is frequently ambiguous. Name the artefacts individually: fine-tuned weights, training and evaluation datasets, prompt templates, retrieval configurations, and any tooling built to produce them. Ambiguity here is common in AI work because these artefacts do not look like conventional deliverables, and disputes about them surface only when you try to move to another vendor.
How should inference costs be handled in an AI outsourcing contract?
With a stated cost target per unit of work, measured during acceptance, and clarity on who pays for model API usage during development. A vendor optimising only for an accuracy number will reach it in the most expensive way available, because nothing in the contract penalises that. Expressing the target as cost per document, per conversation or per transaction keeps the economics visible before you inherit the running bill.
Ready to scope it properly?
We help Singapore teams write enforceable AI statements of work — and staff the in-house engineers who hold the vendor to them.
Get started