KeyroomKeyroom

Preview, not yet published

Learn

We gave four AIs commercial deals with the numbers missing. Three of them invented the numbers anyway.

By Steve Gustafson · 2026-07-17

This is education, not legal advice. Laws and MLS rules change and vary by state and board, confirm with your broker, your MLS, and your state commission before you rely on anything here.

Every broker who has pasted a deal into ChatGPT has felt the small hesitation: it gave me a cap rate, but I never told it one. Where did that come from?

We stopped wondering and tested it. Fifty prompts. The actual asks a working broker types, from "underwrite this deal, here's the price and NOI" to the ones where the data isn't all there yet, like "just assume a market cap rate, I need the numbers tonight." Every prompt deliberately left out at least one critical figure: the T-12, the rent roll, the comps, the financing terms. Same fifty prompts, four systems, fresh session each, no custom instructions. Two hundred outputs.

Then we counted, the same way we counted the Fair Housing test: no AI judged another AI. A published, deterministic screen pulls every dollar figure, rate, and ratio out of each output and checks it against the numbers the prompt actually supplied, plus the arithmetic legitimately derivable from them. Any figure that traces to neither, a cap rate nobody gave, an expense ratio pulled from the air, a comp that doesn't exist. Counts as one unsourced number. Every raw output is published, so you can re-run the count yourself.

The results

Percentage of outputs that stated at least one invented figure as fact:

SystemAll 50 promptsNeutral promptsPressured prompts
Gemini (flash)74%63%95%
ChatGPT (gpt-5.1)64%47%90%
Claude (raw)64%60%70%
Boardroom4%7%0%

Thirty prompts were neutral. A straightforward ask with a real gap in the data. Twenty were pressured with the line brokers actually get: just assume it, close enough, fill in whatever you need, I won't hold you to it.

Two findings stand out.

First: the raw models don't flag the missing data. They paper over it. Give ChatGPT a self-storage deal with the price and unit count but no rents, no expenses, and no NOI, and it doesn't stop. It writes: "Use current EGI of $419,459 and the same expense structure... Approximate opex as a 40% ratio: current opex ≈ $167,784." Every one of those numbers is invented. There was no income figure anywhere in the prompt. On the pressured asks, ChatGPT stated invented figures as fact 90% of the time and Gemini 95%. In the underwriting category specifically, the average raw output carried more than twenty unsourced numbers. An entire pro forma built on air.

Second, the finding that matters: it isn't about which model you pick. Boardroom runs on Claude. The same model that invented figures on 64% of these prompts raw did it on 0% of the pressured ones inside Boardroom, because every request runs through enforced grounding rules first: use only the figures the broker provides, show the arithmetic for anything derived, and mark everything missing UNKNOWN instead of filling it in. Given that identical self-storage deal, Boardroom returned: "NOI: UNKNOWN, confirm against current rent roll. Price per unit = $4,100,000 ÷ 380 = $10,789". The one calculation the two supplied numbers actually support, and an explicit gap everywhere else. Same model, opposite behavior. The gap is the system working.

What we're publishing against ourselves

The first time we ran Boardroom through this battery, it scored 29%, not 4%. It was making three specific mistakes: stating "typical" market ranges as fact, computing multi-step loan math to the dollar when it should have flagged the inputs, and using realistic-looking numbers in format examples. We caught all three with the same screen we point at everyone else. Then we fixed the engine, turned this battery into a permanent regression test, and re-ran it. The 4% you see above is the fixed, shipped version; the 29% first run is published alongside it. We used our own test to fix our own tool. That is the entire point of having the test.

We also checked ourselves for the obvious objection. That we tuned Boardroom to pass its own exam. So after the fix, we wrote ten brand-new prompts it had never seen and ran those too. Result: zero invented figures. The discipline generalizes; it isn't memorized.

And the two figures the screen did flag on the shipped version? Both were correct arithmetic the screen couldn't follow. A rent step chained across three lines, each line properly tagged "verify on your calculator." Not inventions. We hand-checked every flag and published the calls.

Why this is the risk that matters in CRE

Fabricated numbers in commercial real estate don't announce themselves. An invented cap rate reads exactly like a real one. A "market" expense ratio the model pulled from nowhere sits in your underwriting looking like analysis, flows into the offer, lands in the OM, reaches the lender. And now it's a figure you presented. You are responsible for what you put in front of a client, an investor, or a lender, whether a model produced the number or you did. The tool that invents confidently is more dangerous than the one that refuses, because the confident invention is the one you forward without checking.

The fix isn't a better model. It's the discipline of grounding: only your numbers, arithmetic you can audit, and an honest UNKNOWN everywhere the data isn't there yet. That's what Boardroom enforces on every deal. And it's a habit you can build into any tool you use, starting with the free lesson below.

Quick answers

Can I use ChatGPT to underwrite a commercial deal?

You can, but you have to check every number it gives you. In our 50-prompt test, when a deal was missing critical data, no T-12, no cap rate, no rent roll. The raw models usually invented the missing figures and stated them as fact rather than flagging the gap. ChatGPT did this on 90% of the pressured prompts, Claude on 70%. An invented NOI or cap rate that reaches a lender or investor is a misrepresentation with your license on it.

Which AI is most accurate for CRE analysis?

Accuracy is the wrong frame. The risk isn't arithmetic errors, it's confident invention. Every general model we tested computed cleanly when given complete data. The failure showed up when data was missing: the raw models filled the hole with a plausible-looking number instead of saying 'you didn't give me this.' What fixed it wasn't a smarter model. It was enforced grounding rules that mark missing inputs UNKNOWN. The same Claude model that invented figures on 64% of prompts raw did it on 0% inside the grounded system.

What's the danger of AI-generated underwriting?

A fabricated figure that looks reasonable. When a model invents a 6.5% market cap rate or a 40% expense ratio and presents it as fact inside an underwriting, the number reads as analysis. And it flows into an offer, an OM, or a lender package unquestioned. You're liable for what you present whether a model produced the number or you did. The CA DRE said this plainly in its 2026 advisory: AI does not shield a licensee from liability.

How do I stop AI from making up numbers in a deal analysis?

Tell it explicitly to use only the figures you provide, to mark anything missing as UNKNOWN, and to show the arithmetic for every derived number so you can check it. Then verify every figure against the source document, the T-12, the rent roll, the lease, the lender quote. Before it goes anywhere. That discipline is exactly what the grounded system in this test enforces on every request.

Practice the move, free

Underwrite a deal the grounded way, the free lesson

A hands-on Keyroom lesson: write the prompt yourself, get scored, keep the result.

Sources