• Introduces A 2 D 2 C: attention-driven dynamic convolution for local adaptation. • Uses multi-point random sampling to route and fuse k base kernels efficiently. • Presents A 2 D 2 C + that fuses kernels once, cutting redundancy and MAdds at parity. • Shows consistent gains on ImageNet, CIFAR-100 and COCO with statistical reports.
Zhang et al. (2026) studied this question.