In this paper, we study how different methods for classifying correct answers to open questions impact calibration scores in language models. We compare seven different techniques to perform this task. In this setting, we find that even though answer classification techniques have up to 21% differences between each other, calibration scores are not affected significantly. We find these results show evidence of unreliability on commonly used metrics in this field.
No takes yet. Share an insight, caveat, or question.
Ivetta et al. (2024) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: