Researchers have developed a new framework to audit question-order effects in large language models using the QQ (quantum question) equality, a parameter-free prediction method originally applied to human survey responses.
The study, published on arXiv by Pilsung Kang, adapts quantum measurement theory to evaluate how the sequence of questions influences LLM judgments. The QQ equality framework separates order sensitivity, QQ imbalance, and residual contextuality rather than treating them as equivalent phenomena.
The research introduces a "committed multi-turn forced-branch protocol" that reconstructs order-conditioned joint distributions from next-token log-probabilities. This method uses counterbalanced label mappings and pre-specified health gates to ensure measurement validity.
Testing reveals measurement saturation problem
Pilot testing on an open-weight instruction-tuned model revealed a critical measurement issue. Binary-conditioned distributions were near-deterministic for 17 of 18 item pairs under direct evaluation and 7 of 8 under persona framing.
Label assignment materially changed several mapping-specific QQ verdicts, and no item was certified as residually contextual. The findings suggest that observed QQ outcomes did not uniquely identify response mechanisms when measurement interfaces were saturated and label-sensitive.
The study characterizes which mechanism families satisfy QQ robustly, showing that marginal-independent kernels satisfy QQ only when all four mismatch transition rates coincide. Classical repetition can reproduce the equality exactly under specific conditions.
The research combines QQ with the rank-2 Contextuality-by-Default criterion through a mathematical relationship that bounds QQ imbalance by order-sensitivity scores.
"Next-token probabilities should not be interpreted as survey-response distributions without first establishing adequate dispersion," the paper concludes.
The findings have immediate implications for LLM evaluation methodology. The researchers argue that saturation screening and label counterbalancing should precede structural interpretation in distribution-level audits of LLM judgments.
This work addresses growing concerns about bias and consistency in AI systems as they become more widely deployed in decision-making contexts. The quantum-inspired approach offers a mathematically rigorous framework for detecting subtle ordering effects that traditional evaluation methods might miss.
💬 Discussion
Sign in to join the discussion.
Sign in →No comments yet — be the first.