Synapse
⌘+K
Synapse
PulseExploreJournal ClubResearchersJournals
Instagram
HomeJournal ClubExplore
October 15, 2025Open Access

DualReward: A Dynamic Reinforcement Learning Framework for Cloze Tests Distractor Generation

View Full Paper
Ask AI
Bookmark
Share

Authors

THTianyou HuangXCXinglu ChenJZJingshen Zhang

Discussion

Loading...

Member takes

Overview

Experiments showed that DualReward improves distractor generation in cloze tests, suggesting adaptive scaling enhances model performance.

Key Points

  • DualReward employs a unique dual reward structure that adapts based on performance and confidence.
  • Consistent improvements were observed over baseline methods, with up to 3.86% gains in P@1 for diverse datasets.
  • The framework balances learning from human examples while generating novel distractors for testing.
  • Its benefits were notable across both passage-level and sentence-level cloze test datasets.

Cite This Study

Huang et al. (2025) studied this question.

synapsesocial.com/papers/68ef858cc6a308ba063553bahttps://doi.org/10.48550/arxiv.2507.11875
View Full Paper
Ask AI
Bookmark
Share