Enhancing Self-supervised Video Representation Learning via Multi-level Feature Optimization | Synapse