ViT-AE++: Improving Vision Transformer Autoencoder for Self-supervised Medical Image Representations | Synapse