MEP AI takeoff accuracy — measured in the open
We publish accuracy the way it should be published: per-document receipts, not a single opaque number. Per-trade symbol-detection recall, and geometry-grounded route metrics — Coverage@FRR and linear-foot error against an independent geometry oracle. Reproducible, and fail-closed by design.

The Aginera accuracy standard
Aginera reports MEP takeoff accuracy as per-document receipts on real bid sets, scoring symbol detection per trade — recall today, precision-based F1 as annotated ground truth lands — and route tracing with geometry-grounded metrics: Coverage@FRR and linear-foot error against an independent geometry oracle. Low-confidence or implausible results are flagged for estimator review, never silently priced, so every published figure is a floor you can audit.
An accuracy number you cannot audit is marketing, not measurement
It is easy to publish one blended accuracy or F1 figure. It is much harder to show the documents it was measured on, the drawing scale it relied on, and where the model was wrong. Without that, an estimator has no way to trust the number — and a bad quantity that reaches an estimate, quote or bill of materials is expensive to unwind.
Aginera takes the opposite approach. Every test document produces a receipt with the plan-page evidence, the scale provenance and a verdict of auto-priced or flagged-for-review. You audit our accuracy the way you would check a colleague’s takeoff: by looking at the sheet.
A layered system, audited the way an estimator would
Reading, counting, tracing and validating an MEP set are separate jobs with separate failure modes. We measure each layer on its own terms.
Semantic layer read
Aginera reads the drawing’s real vector geometry and CAD layer semantics — not a flattened image. Layers, line work and annotations tell the system what is a conduit versus a wall versus a dimension line before anything is counted or measured.
Symbol detection
Devices, fixtures and equipment are detected and counted per trade — receptacles, switches and panels; diffusers, VAV boxes and equipment; valves, fixtures and cleanouts. This is the count takeoff, scored with precision, recall and F1.
Centerline tracing
Double-line duct and pipe are collapsed to a single measured centerline, and single-line conduit and cable tray are followed run by run. This turns drawn rails into one true path per run — the basis for an honest linear-foot takeoff.
Route length (RouteNet)
RouteNet, our geometry-grounded route engine, measures each traced run to a true length using the drawing’s own scale — not pixel guessing. Because it reads actual geometry, it follows branches, risers and turns end to end and returns real linear feet.
Estimator review
Low-confidence or implausible results are flagged for review with evidence tied back to the sheet. Nothing uncertain is priced silently — the estimator confirms or corrects before a quantity reaches an estimate, quote or bill of materials.
Route tracing: a geometry-grounded bar, not symbol F1
Linear scope — conduit, pipe, cable tray and duct — is where estimating hours go, and where pixel-based tools guess. RouteNet reads the drawing’s real geometry to trace and measure true centerline length, then we score it against an independent geometry oracle. That is a harder check than counting symbols.
| Metric | What it measures | Reported | Scope |
|---|---|---|---|
| Coverage@FRR ≤ 0.05 | Share of a run’s true length accounted for, while holding the false-route rate at or below 5%. | 0.986 | Across scored route pages (plumbing/mechanical + electrical), FRR-bounded vs an independent geometry oracle. Median 1.00; n = 10 qualifying pages of 18 with a measurable false-route rate. |
| Linear-foot error vs geometry oracle | Difference between the traced centerline length and an independent geometry oracle computed from the same drawing. | Reported per document | Every scored plan page carries its own delta and evidence overlay in the receipt. |
| Verdict | Whether the traced quantity was confident enough to auto-price, or flagged for estimator review. | Auto-priced / Flagged | Fail-closed: implausible or low-confidence routes are withheld, never silently priced. |
Route metrics are measured on real customer bid sets and cross-checked against an independent geometry oracle. Coverage@FRR is scoped to the route class stated; per-trade figures and the full per-document receipt table are published from our validation corpus.
The raw numbers — every scored plan page
This is the whole spot-check, not a highlight reel: 25 plan pages across 21 real bid sets, each traced by the deployed engine and cross-checked against an independent geometry oracle. Documents are anonymized to a neutral trade-and-class label; the metrics are the receipts’ own values, unedited.
Coverage@FRR ≤ 0.05
0.986
median 1.00 · n=10 qualifying
Median |Δ| vs oracle
16.7%
geometry-grounded
Auto-priced or flagged
78.3%
23 scored pages
Plan pages / bid sets
25 / 21
2 need GPU raster path
| Doc | Trade | Route class | Scale src | Traced LF | Oracle LF | Δ vs oracle | Coverage | FRR | Verdict | Note |
|---|---|---|---|---|---|---|---|---|---|---|
| E-01 | Electrical | conduit | prior | 0 | 2,454 | — | — | — | Flagged · review | Flattened vector; deterministic path returned 0 → review (needs raster model) |
| E-02 | Electrical | conduit | text | 0 | 3,098 | — | — | — | Flagged · review | Flattened vector; deterministic path returned 0 → review (needs raster model) |
| E-03 | Electrical | conduit | none | — | — | — | — | — | Flagged · review | No recoverable scale + flattened vector → flagged for review |
| E-04 | Electrical | conduit | text | 2,781 | 2,579 | +7.9% | 1.00 | 0.00 | Auto-priced | In-band conduit |
| E-05 | Electrical | conduit | text | 4,653 | 5,538 | −16.0% | 1.00 | 0.00 | Auto-priced | In-band conduit / cable |
| M-03 | HVAC | duct | text | 302 | 232 | +30.4% | 0.99 | 0.148 | Auto-priced | In-band duct |
| M-04 | HVAC | duct | inferred | 198 | 151 | +31.3% | 0.39 | 0.304 | Auto-priced | In-band duct (scale inferred → not auto-priced) |
| M-05 | HVAC | duct | prior | 3,943 | 751 | +425.4% | 0.61 | 0.085 | Flagged · review | Scale recovered from sheet-size prior → out of band, flagged for review |
| P-01 | Plumbing / mech | pipe | text | 30 | 21 | +43.8% | 1.00 | 0.00 | Flagged · review | Engine abstained on a routed sheet → review (no confident LF) |
| P-02 | Plumbing / mech | pipe | text | 3,174 | 2,802 | +13.3% | 0.959 | 0.061 | Auto-priced | In-band site pipe |
| P-03 | Plumbing / mech | pipe | text | 0 | 6 | −100% | — | — | Flagged · review | Engine abstained → review-safe (no priced LF) |
| P-04 | Plumbing / mech | pipe | inferred | 6,047 | 4,057 | +49.0% | 0.988 | 0.027 | Flagged · review | Large total; inferred scale → flagged for review |
| P-05 | Plumbing / mech | pipe | inferred | 2,631 | 2,892 | −9.0% | 0.888 | 0.045 | Auto-priced | In-band |
| P-06 | Plumbing / mech | pipe | text | 173 | 179 | −3.3% | 1.00 | 0.00 | Auto-priced | Tight match |
| P-09 | Plumbing / mech | pipe | text | 16,861 | 14,567 | +15.8% | 0.90 | 0.139 | Flagged · review | Large total → flagged for review |
| P-10 | Plumbing / mech | pipe | text | 629 | 625 | +0.6% | 0.996 | 0.00 | Auto-priced | Tight chilled-water match |
| P-12 | Plumbing / mech | pipe | text | 1,148 | 1,377 | −16.7% | 1.00 | 0.00 | Auto-priced | Clean multi-pipe water |
| C-01 | Unclassified | — | prior | — | — | — | — | — | Not scored | No native vector curves; GPU raster path (not locally runnable) — reported as not-scored, not guessed |
| C-02 | Unclassified | other | none | 0 | — | — | — | — | Flagged · review | No scale recovered → flagged (no priced LF) |
| C-03 | Unclassified | — | prior | — | — | — | — | — | Not scored | No native vector curves; GPU raster path — reported as not-scored, not guessed |
Showing the 20 in-band and flagged-for-review pages. The 5 out-of-band pages are surfaced — not hidden — in the expandable panel below, each adjudicated.
Show the 5 out-of-band pages (surfaced, not hidden)
These pages fell outside tolerance and were not caught by the engine’s own flag. We show them anyway, with the adjudication in “Honest results” below: one genuine tracer miss, oracle/page-selection artifacts, and two correct-geometry totals priced without an absolute-magnitude cap — a gating fix we are shipping.
| Doc | Trade | Route class | Scale src | Traced LF | Oracle LF | Δ vs oracle | Coverage | FRR | Verdict | Note |
|---|---|---|---|---|---|---|---|---|---|---|
| M-01 | HVAC | duct | text | 547 | 209 | +161.7% | 0.715 | 0.052 | Out of band | Oracle artifact: 72-inch duct kernel undercounts dense supply/return |
| M-02 | HVAC | duct | text | 2,619 | 5,273 | −50.3% | 0.365 | 0.379 | Out of band | Genuine tracer miss: site-scale duct rail-pairing gap |
| P-07 | Plumbing / mech | pipe | text | 92,982 | 103,666 | −10.3% | 0.678 | 0.246 | Out of band | Civil water main; oracle agrees within ~10% but priced without a magnitude cap (gating fix staged) |
| P-08 | Plumbing / mech | pipe | text | 9,682 | 9,566 | +1.2% | 0.985 | 0.00 | Out of band | Oracle agrees +1.2%; 5-figure total priced without a magnitude cap (gating fix staged) |
| P-11 | Plumbing / mech | pipe | text | 178 | 76 | +133.8% | 1.00 | 0.00 | Out of band | Fine-scale detail sheet; degenerate oracle (page-selection artifact) |
Honest results
Every row above is a real plan page scored by the deployed tracer against an independent geometry oracle — nothing is cherry-picked. Pages the engine could not price confidently are surfaced as flagged-for-review, not dropped. The five out-of-band pages are shown, not hidden, and are adjudicated: one is a genuine tracer miss (a site-scale duct rail-pairing gap), one is an oracle artifact (a 72-inch duct kernel that undercounts dense supply/return layouts) alongside a fine-scale detail sheet with a degenerate oracle, and two are correct geometry that agrees with the oracle within ~10% but was priced without an absolute-magnitude cap — a gating fix we are shipping. Two conduit pages that need the GPU raster model are reported as not-scored rather than filled with a guessed number, and true negatives — 0 where the deterministic path found no native geometry — are reported as 0, never inflated.
Symbol detection: precision, recall and F1 by trade
Counts — devices, fixtures and equipment — are scored per trade against each document’s own mark census. We report tag recall (the share of the expected marks the model found) and the fabrication count (marks it invented), because a model that is strong on electrical devices can be weak on mechanical equipment, and one blended number hides that. These figures are measured on our Corpus-20 benchmark by a zero-LLM scorecard scored against a fixed expected-scope contract. We publish a per-trade number only once it spans multiple documents — Electrical (8 documents), Mechanical/HVAC (5) and Plumbing (5) each clear that bar, pooled across independent projects rather than headlined off a single sheet.
Electrical
Devices, fixtures, panels and equipment marks.
79%
tag recall
139 of 177 marks · 8 mark-census documents · 0 fabrications
Precision-based F1: pending annotated ground truth
Mechanical / HVAC
Diffusers, VAV boxes, equipment and terminal units.
92%
tag recall
187 of 203 marks · 5 mark-census documents · 0 fabrications
Precision-based F1: pending annotated ground truth
Plumbing
Fixtures, valves, cleanouts and equipment.
66%
tag recall
82 of 125 marks · 5 mark-census documents · 0 fabrications
Precision-based F1: pending annotated ground truth
Source: the committed Corpus-20 scorecard baseline, scored read-only with zero LLM cost. We report recall today, not F1, because a precision-based per-trade F1 requires human-annotated ground-truth marks the corpus does not yet carry — so F1 is marked “pending” rather than published as a number we cannot tie back to annotated documents. Each figure is pooled across independent projects — Electrical 8 mark-census documents, Mechanical/HVAC 5, Plumbing 5 — and the census for every document is enumerated from the drawing’s own text-layer schedule, independent of the extractor, so recall cannot be gamed by the model grading itself.
What a single accuracy receipt contains
Each scored document emits one machine-readable receipt. Together they let anyone reproduce our reported numbers instead of trusting an aggregate.
- ✓Document, plan page and content hash (sha256)
- ✓Deployed engine version and commit
- ✓Drawing scale value and where it came from (text, graphic or inferred)
- ✓Per-class linear feet and the independent oracle length
- ✓Coverage, false-route rate and geometry deltas
- ✓Plan-page evidence overlay (visual proof)
- ✓Verdict: auto-priced or flagged for estimator review
Fail-closed by design — published numbers are floors
When the drawing scale cannot be recovered, is only inferred, or a traced quantity is geometrically implausible, the result is flagged for estimator review rather than priced automatically. Withholding uncertain results instead of guessing is deliberate.
The consequence for this benchmark is important: because implausible results are held back rather than averaged in, the accuracy we report is a floor. An estimator reviews anything the model is unsure about, so the failure mode is a flag to a human, not a wrong number in a bid.
Measured on real bid sets, described so it can be reproduced
The benchmark runs on real customer bid sets — full multi-page plumbing, HVAC, electrical and civil drawing sets — not demo decks or synthetic drawings. Documents are identified by file and page with a content hash, scored by the deployed engine, and cross-checked against an independent geometry oracle computed separately from the production tracer.
We describe the corpus, the metric definitions and the scoring criteria so results can be reproduced. This is the difference between a benchmark and a billboard: a number is only as strong as the evidence that lets someone else check it.
Frequently asked questions
What is F1 for MEP takeoff?
F1 is the harmonic mean of precision and recall for a detection task. In MEP takeoff it scores how well an AI finds the devices, fixtures and equipment on a drawing — receptacles, diffusers, panels, VAV boxes, valves and so on. Recall measures how many real symbols were found; precision measures how many detections were real (not fabricated). F1 rewards a model only when both are high, so it is a fairer single number than accuracy alone. Aginera reports F1 per trade, because a model that is strong on electrical devices can be weak on mechanical equipment, and one blended number hides that.
How is route-length (linear-foot) accuracy measured?
Route length is a geometry problem, not a counting problem, so F1 does not describe it. Aginera measures each traced conduit, pipe, cable-tray or duct run to a true centerline length using the drawing’s own scale, then cross-checks that length against an independent geometry oracle computed separately from the production tracer. The difference between the two is the linear-foot error. Because the oracle is derived from the same real drawing geometry rather than a hand estimate, it is a harder, less forgiving check than symbol-detection F1 on its own.
What is Coverage@FRR?
Coverage@FRR is coverage measured under a false-route-rate ceiling. Coverage is the share of a run’s true length the tracer accounts for; the false route rate (FRR) is the share of traced length that does not correspond to a real run. Reporting coverage alone is easy to game — you can trace everything, including geometry that is not a run, and claim high coverage. Coverage@FRR forces the tracer to hold false routes below a fixed threshold (we report at FRR ≤ 0.05) before its coverage counts, so the number reflects useful, trustworthy length rather than over-tracing.
Why publish per-document receipts instead of one accuracy number?
A single blended accuracy or F1 figure cannot be audited: you cannot see which documents it covers, what the scale provenance was, or where the model was wrong. Aginera publishes a receipt per test document — the plan page, the scale source, the metric values, and a verdict of auto-priced or flagged-for-review — so an estimator can inspect the evidence the same way they would check a colleague’s takeoff. Transparency is the point: the numbers are only as good as the ability to reproduce them.
What does “fail-closed” mean for accuracy?
Fail-closed means that when confidence is low or a result is geometrically implausible — an unrecovered drawing scale, an inferred scale, or an out-of-band quantity — the system flags the item for estimator review instead of silently pricing it. Because implausible results are withheld rather than published, the accuracy figures we report are floors, not averages inflated by lucky guesses. The estimator stays in control of anything the model is unsure about.
What data is the benchmark measured on?
The benchmark is measured on real customer bid sets — full multi-page plumbing, HVAC, electrical and civil drawing sets — not demo decks or synthetic drawings. Each document is identified by file and page with a content hash, scored by the deployed engine, and cross-checked against the independent geometry oracle. The corpus, the metric definitions and the scoring criteria are described so the results can be reproduced rather than taken on faith.
Upload your drawing → get takeoff + BOM in minutes
Free 14-day trial. No credit card required.