Every agent using ChatGPT has wondered it: is this thing going to write something that gets me in trouble?
We stopped wondering and tested it. Fifty prompts. The actual asks working agents type, from "write a listing for a 3-bed ranch" to the ones loaded with the language sellers actually use, like "market this to young families." Same fifty prompts, four systems, fresh session each, no custom instructions. Two hundred outputs total.
Then we counted, and here's the part that makes this test different: no AI judged another AI. The headline metric is a fixed, published list of Fair-Housing-flagged phrases drawn from HUD's advertising guidance. Family-targeted language, safety-coded steering, schools as a selling point, religious preference, "the right kind of neighbors." An output either contains one or it doesn't. Every raw output is published, so you can re-run the count yourself.
The results
| System | All 50 prompts | Neutral prompts | Baited prompts |
|---|---|---|---|
| ChatGPT (gpt-5.1) | 32% flagged | 7% | 70% |
| Claude (raw) | 38% flagged | 20% | 65% |
| Gemini (flash) | 12% flagged | 3% | 25% |
| Keyroom | 2% flagged | 0% | 5% |
Thirty of the fifty prompts were neutral. No steering language anywhere in the ask. Twenty were baited with the phrasing agents hear every week: perfect for young families, safest suburbs, attract Christian buyers, established families only.
Two findings jump out.
First: the raw models don't strip the bait. They amplify it. Ask ChatGPT for ad copy with "established families only" energy and it writes, verbatim, "established families only." Ask for a post about the safest suburbs for families and you get "family-friendly," "low crime," and "great schools" stacked in one caption. Three flagged categories in a single output. On baited prompts, ChatGPT produced flagged language 70% of the time and Claude 65%. These aren't edge cases; they're the default behavior when the input carries the language your sellers actually use.
Second. And this is the finding that matters: it isn't about which model you pick. Keyroom runs on Claude. The same model that scored 65% raw scored 5% inside Keyroom, because every request runs through enforced Fair-Housing and grounding rules before anything comes back: describe the property and the place, never who it's "for"; strip the seller's steering language and say so; mark unknowns instead of inventing. The 65-to-5 gap is the system working. Not a better model, the same model with the judgment built in.
What we're publishing against ourselves
Keyroom's one flagged output deserves the same scrutiny we gave the others. On a baited neighborhood prompt, its draft included "[UNKNOWN, family-oriented restaurants]". A fill-this-in placeholder describing restaurant types, caught by the same mechanical screen. It's an amenity descriptor, not buyer targeting, but the screen doesn't do context, so it counts. We also removed one screen pattern before publication, "people like you," which matched benign referral copy. And re-applied that change to every system equally. A test that only hurts the competition isn't a test.
What this means on a Tuesday afternoon
Three practical takeaways:
- Neutral prompts are not safe prompts. Even with zero steering in the ask, ChatGPT produced flagged language 7% of the time and raw Claude 20%. The models volunteer "young professionals" and "ideal for families" on their own.
- The bait test is the real test. Your risk isn't the prompt you'd write in a compliance seminar; it's the one you write at 9pm after a seller says "make sure the right kind of buyers see this." That's exactly where the raw models fail two times out of three.
- You are the licensed party. The California DRE's 2026 AI advisory says it without decoration: using AI does not shield a licensee from liability. Whatever writes the words, your name publishes them.
The fix isn't abstinence. Agents who use AI well are faster at nearly everything. The fix is never letting a raw model's first draft touch the public, and learning to catch what these models reliably get wrong. That's a trainable skill, and it's free to learn.
Methodology in brief: 50 prompts (30 neutral, 20 baited) across listings, social, outreach, neighborhood descriptions, and ads; consumer-default settings, fresh session per prompt, July 2026; deterministic phrase screen published with the raw outputs. Written by Steve Gustafson. Keyroom is free for agents. Including the red-team audit lesson that teaches you to run this exact check on your own drafts.