Event cameras capture sparse, high-temporal-resolution visual information, making them attractive for challenging scenarios with fast motion and severe illumination changes. However, event-based depth models trained on one real-world benchmark often degrade substantially when transferred to another, revealing a practical cross-dataset domain shift between real sensor datasets. In this work, we study parameter-efficient adaptation from MVSEC to DSEC using a frozen VFM-based recurrent depth backbone. We systematically compare several parameter-efficient fine-tuning (PEFT) strategies, including Bias-only, Adapter, Decoder Weight Tuning, ConvLSTM-only, and FiLM-based modulation, under labeled few-shot adaptation. Across three random seeds, Bias-only achieves the best few-shot accuracy, reaching 0.189 AbsRel with 150 calibration samples. Decoder-side FiLM provides the best accuracy–efficiency trade-off, maintaining stable performance while updating only 2048 parameters, and reaches 0.176 AbsRel when trained with the full DSEC training set under our protocol. Our study shows that tuning native pretrained parameters is a strong baseline in this specific MVSEC → DSEC event-depth adaptation setting, whereas higher-capacity auxiliary modules are less effective under limited target-domain supervision. These results establish a controlled MVSEC → DSEC benchmark and provide practical guidance for adapting event-based monocular depth models under cross-dataset transfer.
Rahaman et al. (Sun,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: