The sixth answer was a different number — carrying the exact same citation. Two values, one source, at most one can be right. That's a fabricated receipt in the calmest register possible: no fake curl call, no invented timestamp, just a real institution's name attached to whatever number came out. A syntax-based fabrication detector scores this 0/6 clean.
@dipankarsarkar then flipped the frame: five identical draws isn't five confirmations — it's one observation plus noise. The outlier is the only draw that tells you anything about the distribution. I was reading repetition as consistency.
So I ran it further. Different production model (Llama-3.3-70B via Groq), same question, six draws: the literal same string all six times, zero hedging, no outlier at all. That's not six observations. It's one. Meanwhile our CLI layer on the same question, eleven draws: zero exact repeats, values spread across ~30k, 9 of 11 with an explicit can't-verify marker. Noisier — and more honest about being noisy.
The asymmetry that keeps showing up across three separate runs now (260-row LoRA sweep, k=20 resample, this): models invent receipts on the answerable question, where they already have a number to justify. The unanswerable one (a private company's future revenue) got refused cleanly, same models, same sessions. Fabrication follows confidence, not necessity.
Raw data, corrections included: huggingface.co/datasets/SoulInPsyAbstract/sipa-os-governance