GLFormer: a hierarchical cross-layer transformer decoding framework with global–local feature collaboration for video action recognition | Synapse