ViTE-MobileNetV2-ResNet101: Fusion Vision Transformer Encoder and CNNs Based on Spatial Detail Enhancement for Early Diagnosis Skin Cancer | Synapse