Researchers have developed a technique to expose when AI models produce reasoning that appears logical but doesn't genuinely rely on the information they claim to use.
The method, called "interventional grounding audits," tests whether large language models truly depend on their stated premises when generating chain-of-thought reasoning. Published in a paper accepted at ICLR 2026's Workshop on Logical Reasoning, the technique works by substituting key terms in premises with random symbols and checking if the model's reasoning steps change accordingly.
Testing GPT-4o's reasoning reliability
Researcher Hironao Nakamura evaluated the method on 50 problems from ProntoQA, a synthetic reasoning benchmark where correct logical dependencies are known. When applied to OpenAI's GPT-4o, the audit achieved an F1 score of 0.806 in detecting genuine proof-tree dependencies.
The technique significantly outperformed existing self-consistency baselines, which only managed an F1 score of 0.343. The audit method achieved 100% recall in identifying predicate-determining dependencies, with an F1 score of 0.885 for this subset.
A concerning finding emerged from the analysis: 66% of problems that GPT-4o solved correctly contained at least one reasoning step that appeared aligned but was actually insensitive to direct proof-tree dependencies. This "right answer, wrong reasoning" phenomenon was particularly prevalent in steps involving entity-introduction premises.
The black-box nature of the audit makes it applicable to any language model without requiring access to internal weights or activations. The method works by intervening on single premises, substituting target predicates with fresh symbols, then re-running the model to check whether reasoning steps change as expected.
All audit certificates, raw model outputs, and reproduction scripts have been made available in a public GitHub repository. The researchers acknowledge the method's current limitations beyond formal, parsable benchmarks but suggest it could help identify when AI reasoning is more superficial than it appears.
💬 Discussion
Sign in to join the discussion.
Sign in →No comments yet — be the first.