Authors
Loading...
Experiments showed that DualReward improves distractor generation in cloze tests, suggesting adaptive scaling enhances model performance.
Huang et al. (2025) studied this question.