The work, one test at a time

- The first runs and the first mistake. What was known: the pipeline could go from plans to a priced bid without a person touching it. The question I asked was where it would fail first, and the answer arrived in May when a quick-service restaurant bid shipped with a markup-stack error: overhead, profit and contingency each compounding on the one before instead of each hitting its own base. I had let it go out because the engine had been right before. I caught it, corrected it in the second version, and converted the failure into the first of the engine's permanent verification gates, a check that fails the build mechanically rather than trusting a reviewer to notice. What it changed: from that bid on, a wrong number was something the engine had to refuse, not something a person had to catch.
- The audit and the held-out test. What was known: the engine had been tuned and scored on the same jobs, so its flattering number was closer to a mirror than a measurement. I wrote an adversarial audit prompt and gave the engine to a frontier model with instructions to find every place it was grading its own homework. It found the mechanism: a large share of the rates in the catalog had been transcribed from the anchor project's own budget, so the engine was reading its answer off the answer key. It sorted the 96 formulas; 84 were classed calibrated, assumed or known wrong (40, 33 and 11) and 12 were left unclassed, and it found that the six-layer verification stack existed in code with zero production callers. I wired the gates in, added 24 hard assertions that exit nonzero when an invariant breaks, and ran the leave-one-out test: one bid removed from the corpus at a time, priced as if unseen, compared to its own final submitted total. Three single-family bids had honest engine-versus-actual pairs; the other corpus rows were excluded as invalid pairs because their engine runs had been fed defaults. The result is the figure on the cover. It was not good enough to sell as finished, so the public posture was set at integration roadmap, not launch. What it changed: every status report and every public page since cites the held-out figure, never the in-sample one.
- The calibration pass. What was known: a held-out error tells you the size of the miss, not where it comes from. On July 6 I crossed the formula audit against the 41-bid corpus of real submitted bids, one formula at a time, with a rule that no formula would be tuned toward a figure a later revision had retracted, and no corpus evidence meant no change. Of the 54 formulas dispositioned, 36 could be calibrated from the corpus and 18 could not; those 18 ship with their debt annotated in the engine itself, so the gap stays visible instead of being papered over. Locating the formulas exposed two engine defects that no single bid review had shown: a premium finish was silently priced at the mid tier because of an orphaned key, and a life-safety trade existed in the catalog with no hook, so its line never emitted at all. Both were fixed before either reached a client bid, and the direct-cost gap on the anchor home closed. What it changed: the corpus, LÏEF's own priced history, became the thing the engine is checked against, and a formula without corpus evidence is now visibly an assumption.
- The manufacturer's blind trial and the rule it produced. In July a building-materials manufacturer's VP of Operations, who had his own in-house estimators and no reason to trust an outside tool, sent two unsolicited tests the same day: a stamped set, and a marketplace "website plan" with no dimension strings, no schedules and no elevations. He wanted linear footage and heights for exterior walls, interior walls and roof, and he wanted to know how long it took. We ran three takeoffs, each under an hour. On the stamped set every overall dimension chain was closed before a quantity was reported, the footprint reproduced the plan's own area table to the digit, and we flagged a conflict inside his own set, structural notes calling a 5-inch minimum slab against plan labels calling 4 inches. On the dimension-less set we calibrated scale off the labeled room dimensions and caught that one file was a basement-option variant mislabeled as a second story. Then the blind second pass, a re-measurement with no access to the first, caught a real error in our own first pass: a wall-height model that would have undersized the shell by roughly a third. The corrected numbers are what shipped, inside his own estimating spreadsheet on his rates and waste factors, so his team could check our quantities in a framework they already trusted. His first read was that he wanted to buy the method outright. I kept the method. What it changed: the blind second pass, made under a prospect's own clock, became the standing rule for every takeoff LÏEF produces, not a flourish for a trial.
- The blind backcast and the proof page. Before anything went on a public page I wanted a test where the engine had no chance to have seen the answer first. I took eleven complete bids from our own history, gave the estimator only the original intake plan set, produced a full takeoff and estimate blind under a logged seal, and only then compared against what had actually been sent. Every miss traced to a pricing convention, scope tier, who supplies material, the furniture boundary, the markup stack; zero misses traced to takeoff quantities, and a 121-unit hotel's building area landed within 1 percent of the set. That is a proof of measurement and scope completeness, not a cost-accuracy figure, and I did not publish a universal percentage next to it. What it changed: the public rule became results only, method private, no rate library, no internals, no named projects.
- The vector overhaul and the campus re-measure. By September the engine had a working history, but sheet reading was still done by rendering pages and reading them the way a person would, careful, but a person, and a person misses things. Modern architectural PDFs carry exact line work as vectors, on the order of 49,000 paths on a single plan sheet: wall runs, door swings, hatch bands, grid scale. I pulled the engine apart and rebuilt the front end on that geometry instead of on a picture of it, with the rule that nothing is trusted until an overlay has been laid back on the original sheet and checked by eye. The first thing I ran it on was a campus we had already estimated in April by hand. The rebuilt engine also caught a real mid-project revision on another set, a 4.9 percent pixel change and a path count that moved from 36,646 to 76,994 between versions, which is what a genuine revision looks like and not what a rounding error looks like. What it changed: quantities now come from the drawing's own numbers, the overlay check is mandatory, and an order-of-magnitude ratio check runs beside it as a second, independent control.
- The architecture arm. The same discipline runs on the design side, where a Revit skills library with a human in the loop turned a stalled as-built into a design-development set on a six-week agentic build. A three-agent adversarial review of that model caught a wrong setback before any person did, the kind of mistake that is cheap on a screen and expensive after concrete. That six-week build cut architectural production time 65 percent against a 50 percent target, measured against history: I know how long these projects take. The same model was later found to carry a footprint problem of its own, traced to another project's plan set, which is exactly why the overlay rule and the blind pass exist and why no self-caught error is treated as the last one.
What it produced
The held-out chart, behind the gate, is the engagement in one picture. The two in-sample bars are what the engine scored on the home whose budget had leaked into its own rate catalog; they are printed there as the contrast the audit exposed, and they are not accuracy claims. The three held-out bars are the same engine priced blind on bids it had never seen, one high and two low, which is the other thing the test taught me: the error is not a steady bias a single correction factor can fix, it is quantity error, and the cure is division-level corpus pairs and a real takeoff, not a fudge on the total. Three bids is a small corpus. I say so beside the number, and the number gets re-run as the corpus grows, because the figure a buyer can check is worth more than the figure that reads well.
Where it is now. Kanopi did not stay a LÏEF tool. The builders and developers on the cover are pricing their own work through it, on projects across the country, and the adoption is still growing. The door is open to any architecture firm or build firm that gets hold of what we are able to do. What did not move with the volume is the rule: the held-out figure is still the only accuracy claim Kanopi makes, and more proposals through the engine is more corpus to re-run it on.

The calibration chart shows what the corpus could and could not settle. Where LÏEF had priced the work many times, finishes, plumbing, the cross-cutting ratio formulas, the corpus calibrated nearly everything. Where it had not, site work, framing, thermal and moisture, the formulas stayed assumptions and now say so inside the engine.

The partition catch is the vector overhaul's first proof and my own five-times overcount. The April desktop estimate had put interior partitions at 9,470 linear feet on a 36,153 SF campus, one linear foot of wall for every 3.8 SF of floor, a wall every four feet. The vector takeoff measured 1,834, about 19.7 SF of floor per linear foot, an ordinary ratio for the building type, and an independent blind pass with no access to the engine's run landed within 2.4 percent of it, at 1,878 gross. The size of the gap is itself a finding. A misread scale bar produces something close to a two-times error, not five; the April figure had come from a density guess applied to the whole floor plate, which the ratio check alone would have flagged. Eleven sheet-level contradictions from the same pass went to the architect as formal requests for information rather than being quietly resolved in house.

The blind backcast, by project type:
| Project type | Plan stage | What was checked blind |
|---|---|---|
| Accessory dwelling units | Concept sketch to permit set | Quantities and scope list |
| Custom homes | Concept to permit set | Quantities, scope, pass-through quotes |
| Additions | Schematic to permit set | Quantities and scope list |
| Restaurants | Permit set | Quantities, scope, the markup stack |
| Commercial shells | Schematic | Wall linear footage and coated area |
| An office interior | Permit set | Quantities, the furniture boundary |
| A youth facility | Schematic | Quantities and scope list |
| The hotel, 121 units | Design development | Area, quantities, window count |
The speed and cost figures, reconciled once. Six numbers that look inconsistent are six different measurements:
| What the number measures | The figure | When and where |
|---|---|---|
| Waiting on trade bids before an estimate could go out | Three weeks | Tellus, about 2014, before the in-house system |
| A detailed estimate from Tellus once the metrics were encoded | Two days | Tellus, 2014 onward |
| The same estimate at LÏEF today, engine behind it | About two hours | LÏEF, 2026 |
| A full quantity takeoff, scale ruler against the engine with its overlay check | About 40 hours to about 40 minutes | LÏEF, July to August 2026; the time pair is the claim, not a percentage |
| The first autonomous run, plans to priced bid, a 6,919 SF spec home | 3 minutes 31 seconds; the next three ran 2 to 4 minutes each | April 2026 |
| Compute per autonomous run | Roughly a dollar | April 2026 |
| A takeoff delivered to a counterparty, blind pass and overlay included | 24 to 48 hours | June to September 2026 |
| Internal cost per takeoff, all in | Under $350 | June to September 2026; there is no public price card |
What I kept, replaced and installed
Kept: the 2014 decision to own estimating instead of renting it from a trade every time, and the per-trade metrics that decision encoded. Kept: the human in the seat. I plan to run the estimating seat myself in year one, because the engine is only as good as the person checking the overlay.
Replaced: outside takeoff vendors, scattered rate sheets, and manual scaling of rendered drawings. The rate sheets were the faulty logic that had to change now. They had been assembled by whoever was estimating on a given day, with no date and no source beside a rate, so two people reading the same plan set drifted apart and nobody could say why. The replacement is one versioned rate library where a rate counts as verified only with a date and a source next to it.
Installed: the leave-one-out gate on any published accuracy claim, with ground truth and corpus size stated beside it; the overlay inspection before any vector-derived quantity enters a bid; the blind second pass on every takeoff, with no access to the first; the order-of-magnitude ratio check beside the geometry; the hard assertions that fail the build; and the public rule, results only, method private. The public accuracy posture is held at integration roadmap through 2026 so that none of it relaxes under commercial pressure.
A model graded on the jobs it was tuned on is grading its own homework.
What it costs to hold the line, and what I watch
Holding the worse number costs speed to market against competitors who are not making the same disclosure, and some polish in a sales conversation where a rounder figure would close faster. It also means saying, in public, that the held-out test is still self-run and has not been re-measured since the vector overhaul, which gives back some of the credibility the disclosure was meant to buy. I would rather under-claim a product that is still proving itself than oversell a brand before the accuracy work is done.
There is a second cost, one I priced at Tellus with my eyes open. Hand someone a fast, thorough estimate and you have handed them your thinking. A client can shop it to a stranger and ask them to beat the number without doing the work. My answer then and now is that those are not the clients I want. The ones worth keeping recognize a real answer delivered fast and pay for the person who can keep delivering it.
What I watch: the corpus count behind the held-out figure, because three single-family bids is a starting point and the number is re-run as division-level pairs come in; whether a bid lands inside its stage band by honest quantities or by compensating errors, since a total that lands by luck does not pass; the sign of the error on each new held-out bid, because a steady bias is fixable and scatter means quantity work; and the day an outside party re-runs the test with no stake in the result, which is the day the figure stops being mine. Kanopi is one brand today, a flagship site with an anonymized proof page and a gated data room, live at kanopibuild.com. The method stays private, and every public result is one a stranger can check.
The result, in short
Held out one bid at a time and priced it blind: mean error 18.6 percent across three single-family bids, worst bid 31.9 percent, the only accuracy claim Kanopi makes. As of September 25, 2026, the engine is in place with eight builders and developers and has priced over $600 million in proposals nationally this year. The public posture stays integration roadmap, not launch, and every page since cites the held-out figure, never the in-sample one.
A slice of the project list
A few related projects.
- Canyon Corporate: takeoff, pricing structure and bid revision, priced through Kanopi (2026)
- Xtrata Consulting Seat: a standing estimating and pricing seat on a structural panel system (2026)
- 301 W Osborn: development strategy and entitlement, LÏEF Development (2024 to present)