Authors
Loading...
Computational study demonstrates state-of-the-art reward alignment in diffusion models, highlighting an effective tuning-free method to prevent reward hacking.
Tang et al. (2024) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: