We present a Query-by-Vocal Imitation (QBV) system that ensembles four pretrained audio encoders (MobileNetV3,PANNs, PaSST, BEATs), each contrastively fine-tuned on VimSketch pairs. Two inference fusions are compared:System 1, a novel class-aware mixture-of-experts using per-class Mean Reciprocal Rank (MRR) for soft weights;and System 2, global fusion (weighted average) without class information. System 1 attains the best MRR (0.3191) andclass-wise metrics, outperforming System 2 and single-model baselines. Code is available at: https://github.com/RP335/qvim-challenge-aalto
Peter et al. (Tue,) studied this question.