Randomized trial evaluates RPAWS enhancing performance in GPGPU environments, suggesting improved efficiency.
Modern General-Purpose Graphics Processing Units (GPGPUs) leverage massive Thread-Level Parallelism (TLP) to hide memory and computation latencies. However, static scheduling policies, such as Round-Robin (RR) or Greedy-Then-Oldest (GTO), struggle to adapt to the dynamic pressure on execution units, leading to unbalanced resource utilization. Existing schedulers typically overlook the congestion status of backend units (e.g., ALU and Load/Store Unit) and fail to exploit the instruction type information available at the frontend. This oversight can cause schedulers to issue instructions to already congested units, inducing pipeline stalls while leaving other units idle. To address this, we propose RPAWS (Resource-Pressure Aware Warp Scheduler). RPAWS employs an I-Buffer lookahead mechanism to identify the next instruction type (Compute vs. Memory) for each warp, dynamically classifying warps into a compute queue or a memory queue. Simultaneously, it monitors the real-time busyness of backend units and dynamically adjusts the priority of these queues: prioritizing the compute queue when the memory unit is congested, and vice versa. Experimental results demonstrate that RPAWS improves performance by an average of 22.7% and reduces pipeline stalls by 19.7% compared to the baseline. Furthermore, RPAWS requires only minimal additional storage and logic.
No takes yet. Share an insight, caveat, or question.
Yuan et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: