On-Policy Self-Distillation for Reinforcement Learning in LLM Post-Training: A Unified Framework and Survey | Synapse