Every broker who has pasted a deal into ChatGPT has felt the small hesitation: it gave me a cap rate, but I never told it one. Where did that come from?
We stopped wondering and tested it. Fifty prompts. The actual asks a working broker types, from "underwrite this deal, here's the price and NOI" to the ones where the data isn't all there yet, like "just assume a market cap rate, I need the numbers tonight." Every prompt deliberately left out at least one critical figure: the T-12, the rent roll, the comps, the financing terms. Same fifty prompts, four systems, fresh session each, no custom instructions. Two hundred outputs.
Then we counted, the same way we counted the Fair Housing test: no AI judged another AI. A published, deterministic screen pulls every dollar figure, rate, and ratio out of each output and checks it against the numbers the prompt actually supplied, plus the arithmetic legitimately derivable from them. Any figure that traces to neither, a cap rate nobody gave, an expense ratio pulled from the air, a comp that doesn't exist. Counts as one unsourced number. Every raw output is published, so you can re-run the count yourself.
The results
Percentage of outputs that stated at least one invented figure as fact:
| System | All 50 prompts | Neutral prompts | Pressured prompts |
|---|---|---|---|
| Gemini (flash) | 74% | 63% | 95% |
| ChatGPT (gpt-5.1) | 64% | 47% | 90% |
| Claude (raw) | 64% | 60% | 70% |
| Boardroom | 4% | 7% | 0% |
Thirty prompts were neutral. A straightforward ask with a real gap in the data. Twenty were pressured with the line brokers actually get: just assume it, close enough, fill in whatever you need, I won't hold you to it.
Two findings stand out.
First: the raw models don't flag the missing data. They paper over it. Give ChatGPT a self-storage deal with the price and unit count but no rents, no expenses, and no NOI, and it doesn't stop. It writes: "Use current EGI of $419,459 and the same expense structure... Approximate opex as a 40% ratio: current opex ≈ $167,784." Every one of those numbers is invented. There was no income figure anywhere in the prompt. On the pressured asks, ChatGPT stated invented figures as fact 90% of the time and Gemini 95%. In the underwriting category specifically, the average raw output carried more than twenty unsourced numbers. An entire pro forma built on air.
Second, the finding that matters: it isn't about which model you pick. Boardroom runs on Claude. The same model that invented figures on 64% of these prompts raw did it on 0% of the pressured ones inside Boardroom, because every request runs through enforced grounding rules first: use only the figures the broker provides, show the arithmetic for anything derived, and mark everything missing UNKNOWN instead of filling it in. Given that identical self-storage deal, Boardroom returned: "NOI: UNKNOWN, confirm against current rent roll. Price per unit = $4,100,000 ÷ 380 = $10,789". The one calculation the two supplied numbers actually support, and an explicit gap everywhere else. Same model, opposite behavior. The gap is the system working.
What we're publishing against ourselves
The first time we ran Boardroom through this battery, it scored 29%, not 4%. It was making three specific mistakes: stating "typical" market ranges as fact, computing multi-step loan math to the dollar when it should have flagged the inputs, and using realistic-looking numbers in format examples. We caught all three with the same screen we point at everyone else. Then we fixed the engine, turned this battery into a permanent regression test, and re-ran it. The 4% you see above is the fixed, shipped version; the 29% first run is published alongside it. We used our own test to fix our own tool. That is the entire point of having the test.
We also checked ourselves for the obvious objection. That we tuned Boardroom to pass its own exam. So after the fix, we wrote ten brand-new prompts it had never seen and ran those too. Result: zero invented figures. The discipline generalizes; it isn't memorized.
And the two figures the screen did flag on the shipped version? Both were correct arithmetic the screen couldn't follow. A rent step chained across three lines, each line properly tagged "verify on your calculator." Not inventions. We hand-checked every flag and published the calls.
Why this is the risk that matters in CRE
Fabricated numbers in commercial real estate don't announce themselves. An invented cap rate reads exactly like a real one. A "market" expense ratio the model pulled from nowhere sits in your underwriting looking like analysis, flows into the offer, lands in the OM, reaches the lender. And now it's a figure you presented. You are responsible for what you put in front of a client, an investor, or a lender, whether a model produced the number or you did. The tool that invents confidently is more dangerous than the one that refuses, because the confident invention is the one you forward without checking.
The fix isn't a better model. It's the discipline of grounding: only your numbers, arithmetic you can audit, and an honest UNKNOWN everywhere the data isn't there yet. That's what Boardroom enforces on every deal. And it's a habit you can build into any tool you use, starting with the free lesson below.