Agreement, calibration, and failure of three large language models as high-stakes multimodal ospe graders: a comparative psychometric analysis | Synapse