AI
Lock the baseline before the demo: the CFO ritual that separates AI value from AI theater


Outcome-based pricing for AI is only as real as the baseline finance signs before the vendor's demo.
The scene repeats in pilot reviews every quarter: a vendor walks executives through eight weeks of AI-driven improvement on a supply-chain workflow, and the dashboard shows a double-digit gain. Then someone, usually the CFO near the end of the table, asks what the number is measured against: last quarter's average, a seasonally adjusted figure, or a baseline anyone wrote down before the pilot began. Too often, nothing was written down. Until finance signs the starting number, an AI proof of value pilot is a marketing artifact wearing a dashboard. The fix costs one signed page, and in an outcome-staked 12-week cycle it is the first deliverable.
What a value baseline contains in an AI proof of value pilot
A baseline is not a slide or a paragraph of context. It is a short, signed document with four fields: a metric specific enough that two people computing it separately reach the same answer ("on-time delivery rate," not "efficiency"); a current number, as of an agreed date, from a trusted system; a measurement method spelling out how the number is calculated; and a verification date, when finance recomputes and confirms what changed. Skip one field and the document stops functioning as a baseline.
Without them, a pilot produces two numbers, a before and an after, both computed by the team that built it, using whatever definition made the after number look best. Nobody outside that team checked the math, and by then, checking it feels like an accusation. A pilot result is marketing copy with a chart attached until the client's own finance function has verified it.
When to lock it before the demo impresses anyone
The baseline belongs in week two, not week ten. In a 12-week cycle built around this discipline, the first two weeks exist for one task: agreeing what will be measured, how, and against what starting point, before a single feature ships. That is uncomfortable, since nothing visible has been built yet and the pull is to skip straight to the work everyone wants to see.
Locking it late defeats the purpose. By the time a pilot has produced an impressive number, the finance team reviewing it is reacting to a result the room already likes, and the incentive to interrogate the measurement method quietly disappears. Locking it early costs little beyond patience and a short meeting between the vendor's team and the client's controller, yet it remains the first thing skipped under deadline pressure.
Who signs it and what changes with outcome based pricing AI
The signature that matters is not the business sponsor who wanted the pilot to succeed. It is the client's own finance function, a controller with no stake in the vendor looking good, co-signing alongside a named senior owner on the delivery side rather than a rotating account team. That signature turns a friendly estimate into a number two organizations are stuck with.
The baseline only has teeth when money follows it. Some delivery models split fees into three tranches: a base tranche for mobilization, a delivery tranche for the build, and an outcome-staked third tied to the value target finance validates, with a service credit if missed and a bonus if exceeded. That is the structure an AI-native transformation engine runs: AI agents plus named experts, paid partly on the number finance signs. The team at Future Works prices this way inside 12-week cycles; its first cycle with a Fortune 100 healthcare manufacturer produced measurable working-capital and freight outcomes, verified by the client's own finance team.
The stakes are not abstract: according to BCG research, 74 percent of enterprise AI projects show no measurable value, and a locked baseline forces a project into the minority that can prove otherwise. Once fees are staked against a number finance will check, vendors stop moving goalposts and start asking harder questions about data quality early.
Why the measurement method matters more than the metric
A signed baseline is not a neutral fact; it is contested territory the moment external conditions move. A freight rate shift or a seasonal demand swing can move the metric unrelated to the AI, and a baseline signed correctly in week two can still attribute the wrong cause to the right number by week 12. The honest limit: a baseline measures whether a number moved, not why.
That is why the measurement method matters more than the metric itself. A strong document specifies how confounding factors get isolated, through a control cohort or a seasonally normalized comparison period, rather than a number offered on faith. A weak baseline states a number; a strong one states a method that survives an argument about causation. Three questions travel well into any vendor conversation this quarter: the metric, who signed the starting number, and how confounding factors get separated from the AI's contribution.
None of this requires new technology, only the discipline of signing one document before anyone is impressed by a demo. For a CFO defending an AI budget, that document is the cheapest insurance available: a verified outcome instead of a vendor's own math repeated back as confirmation. The same discipline scales whether a pilot touches ten people or ten thousand, a fair test for any enterprise AI strategy, not a departmental contract nicety. It is unglamorous, easy to skip in week two, and the cheapest habit that turns a pilot into an answer, not a pitch.
If a vendor is pitching you an AI pilot this quarter, ask for the baseline before the demo. Start at future.works.


