An engine that could price a plan and could not yet be trusted
Kanopi looks like an AI story. It is not. The engine is twelve years old; the AI part is new.
At Tellus, my design-build firm, the biggest source of delay in preconstruction was never design. It was waiting on trade bids. Every trade prices its own work off a metric, room count, linear feet, square feet, and we were sending drawings out to subcontractors and waiting weeks for a number before we could tell a client what their project would cost. Around 2014 I encoded those same per-trade metrics into an estimating system we ran in house, so Tellus could produce a complete estimate without the wait. I put more into estimating, as a function inside the company, than most firms at that revenue would have, because removing that wait was worth more than almost anything else we could fix. Clients got a real number fast instead of a placeholder followed by silence.
In 2026 the same shortage came back in a new shape. LÏEF's bids depended on outside takeoff vendors and scattered rate sheets, both slow and both a source of silent error on anything priced under a deadline. So I rebuilt the 2014 decision with better tools: an engine that reads a plan set, takes off the quantities, prices them from a sourced rate library and assembles a bid, with a set of guards that refuse to ship a number the engine cannot stand behind.
The first end-to-end run, in April, priced a premium spec home from plans in a few minutes, then three more autonomous bids in about the same time. That week proved the pipeline could execute. It did not prove the output was worth anything. The question that actually mattered was not "can it price a plan," it was "can it price a plan it has never seen, and how far off is it when it is wrong."
The people in it: me in the estimating seat, an architect of record who reviewed drawings and answered requests for information, the manufacturer's VP of Operations who tested the engine unsolicited in July, and the outside auditor I gave the engine to in June, which was a frontier model told to find every place the engine was flattering itself.
Which number would I bet on
The pipeline's first runs came back with an in-sample gap of a few percent against the one home whose actual costs we held. That is the number every builder of an estimating tool wants to print. I did not print it, and the reason is the whole engagement.
An in-sample number and a held-out number wear the same unit and measure two different things. In-sample tells you how well a model reproduces answers it was shown. Held-out tells you whether it can price something it was never shown. Both come back as a percentage error on the same kind of bid, so they read as interchangeable to anyone who has not built the tool, and a seller who wants a stronger number has every incentive to publish the first one and call it accuracy. What broke that incentive here is not a rule about statistics. I have spent twenty years signing my name to a price and eating the difference when it was wrong. That turns "which number is stronger" into "which number would I bet a thin-margin bid on."
The options were plain. Publish the pipeline's in-sample accuracy, which was flattering. Or design the test that would embarrass it, hold one bid out of the corpus at a time, price it as if unseen, compare it to the number that actually went out the door, and publish whatever it said. I chose the second, and I barred in-sample numbers from ever appearing as an accuracy claim before I saw the held-out result, so the rule could not bend to the outcome.
Three further decisions followed the same logic. Trust the calibrated numbers, or check every formula against the corpus of real bids one at a time. Keep scaling rendered images of drawings, or measure the drawing's own vectors with a mandatory overlay check by eye. Put the rate library on the public site to look impressive, or publish results only and keep the method private. Each time I took the harder version, because the harder version is the one a buyer can build a decision on.
How I came at this one
The question I asked first was whether the number in front of me was a forecast or a memory, so I built the test to prove myself wrong. That question fit because Kanopi's whole value is a number a stranger can check, and a number tested on the jobs it was tuned on cannot be checked by anyone. The second question was what a takeoff actually is: a claim about a drawing, so I measured the drawing itself. The third was what to keep private: the corpus and the method are the moat, the results are the proof.