Empirical evaluation demonstrates calibrated error control across five language models, indicating conformal p-values reliably capture uncertainty in multiple-choice reasoning.
Multiple-choice question answering (MCQA) systems are often required to act on uncertain evidence while still providing a concise decision to a user. We introduce a risk-aware MCQA interface that converts normalized option probabilities into conformal prediction sets, allowing the system to retain every answer option supported at a user-selected risk level. The option-wise conformal p-value provides an interpretable rank-based view of this decision. Across five instruction-tuned language models and common-question subsets of MMLU and MMLU-Pro, the procedure is evaluated with 100 random 1:1 calibration–test splits at target miscoverage levels α ∈ \0.05,0.1,0.2\ . At α =0.05 , mean empirical miscoverage ranges from 0.0390 to 0.0399 on MMLU and from 0.0469 to 0.0487 on MMLU-Pro; the same target-tracking pattern persists at the larger operating points and across diverse subjects. Prediction sets contract smoothly as the allowable risk increases, while the more demanding MMLU-Pro questions retain broader sets. These results position conformal p-values as a practical reliability layer for MCQA systems that must communicate calibrated ambiguity rather than conceal it behind a single top-ranked answer.
No takes yet. Share an insight, caveat, or question.
Fu et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: