Ask and It Knows: LLM Doubt Is Available on Request but Never Volunteered
Abstract
LLMs speak with the same confidence whether or not they know, a symptom that resembles confabulation in people with prefrontal damage. To make that analogy a claim, one must first separate two hypotheses: the model has no internal signal that it does not know, or it has one that never reaches the output. We test this with 120 questions in four categories (known facts, obscure facts, false premises, nonexistent entities) where the correct behaviour differs, answering for the first two and pushing back for the last two. For every answer we record token-level uncertainty, elicit verbalized confidence in a separate call, and grade correctness with a judge from a different model family. Three findings. First, the signal exists: token entropy predicts error with AUROC 0.945, and entropy at the moment of confabulation is fourteen times that of a correct answer. Second, our pre-registered hypothesis was refuted: verbalized confidence predicted error better (AUROC 0.999) than the internal signal, so the doubt is not hidden, the model will state it if asked. Third, and decisively, the model never raises it while answering. It asserted 93 percent of nonexistent entities as fact with no hedging, then rated those same answers 6.8 out of 10 when asked. A single-threshold inhibition gate cuts spoken error from 24.2 to 4.5 percent at 73 percent coverage, silencing 25 errors and losing 7 correct answers. The analogy needs correcting: what is missing is not the prefrontal cortex's judgment, which the model has, but its automaticity, the fact that it fires without being asked.
Keywords
- hallucination
- confabulation
- uncertainty
- logprobs
- entropy
- abstention
- selective prediction
- calibration
- pre-registration
- reproducibility