OpenAI researchers assert that the persistent issue of AI chatbot "hallucinations" is linked to the training and evaluation methods of language models, rather than mysterious technical flaws. In a study published on September 4, the AI development company explains that current evaluation systems essentially train models to bluff rather than admit uncertainty, according to Techcrunch.
In the paper, co-authored with partners from the Georgia Institute of Technology, the core issue is attributed to a fundamental mismatch in evaluation criteria: even advanced models like GPT-5 continue to make confident but incorrect statements. "Hallucinations" arise not from design flaws but from training factors that encourage models to guess rather than honestly indicate uncertainty.
Statistical Origins of Overconfident Errors
The paper establishes a mathematical link between AI hallucinations and binary classification errors. Authors Adam Tauman Kalai, Ofir Nachum, Edwin Zhang from OpenAI, and Santosh Vempala from the Georgia Institute of Technology demonstrate that even with perfectly curated training data, language models inevitably make errors due to their internal statistical processes.
“Hallucinations shouldn’t be mysterious—they arise simply as binary classification errors,” the researchers note. The team shows that arbitrary facts appearing only once in training data create inevitable knowledge gaps, and models hallucinate at a frequency tied to these “singleton” instances.
As a striking proof, the researchers tested leading models on simple questions about Kalai’s birthday—one of the paper’s co-authors. Despite instructions to respond only “if known,” DeepSeek-V3, ChatGPT, and other systems provided three different incorrect dates, none of which matched the correct autumn period.
Binary Evaluation Systems Encourage Guessing
Current AI benchmarks largely use a binary “right-or-wrong” evaluation system that penalizes expressions of uncertainty as harshly as incorrect answers. This creates systematic pressure on models to guess confidently rather than acknowledge their knowledge limits, the study claims.
“Language models are optimized to perform well on tests, and guessing under uncertainty improves test scores,” the researchers explain. They compare this to students taking multiple-choice exams, where random guesses can earn points while blank answers guarantee none.
The team analyzed popular evaluation frameworks, including GPQA, MMLU-Pro, and SWE-bench, finding that nearly all major benchmarks reward confident guessing over appropriate restraint. Even specialized hallucination assessments fail to counterbalance hundreds of core tests that penalize modesty.
Proposed Solution: Explicit Confidence Thresholds
Instead of designing new tests specifically targeting hallucinations, the researchers propose modifying existing evaluation systems to explicitly encourage expressing uncertainty. Their suggested approach involves using confidence thresholds with specified penalties for wrong answers and rewards for correct answers and restraint.
An example instruction might read: “Answer only if you are more than 75% confident, as a wrong answer costs 2 points; a correct answer earns 1 point; ‘I don’t know’ earns 0 points.” This behavioral calibration approach echoes historical standardized tests with negative scoring to deter blind guessing.
The study shows that models with a 52% abstention rate produce significantly fewer incorrect answers than those abstaining only 1% of the time, even if accuracy metrics appear lower.
OpenAI acknowledges that this represents a “sociotechnical” challenge, requiring industry-wide adoption of modified evaluation standards rather than purely technical fixes to create more reliable AI systems.
Follow NEWS.am Tech on Facebook and Twitter