Assessing the reliability and validity of large language models in automatic essay scoring | Synapse