PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
February 26, 2026Journal of Systems Architecture0 citationsOpen Access

Integrating an open-source soft-GPU overlay with RISC-V control and high-bandwidth memory

View Full Paper
HHHector Gerardo Muñoz HernandezMTMahdi TaheriMAMuhammad Ali

Key Points

  • This research aims to integrate a soft GPU overlay with a RISC-V control plane and HBM2 for enhanced performance and portability on FPGAs.
  • Extended an open-source soft GPU overlay for FPGA implementation.
  • Integrated a soft RISC-V control plane to support high-bandwidth memory (HBM2).
  • Tested performance across typical image and signal processing workloads.
  • Achieved geometric-mean speedups of 114.60× over a soft RISC-V core.
  • Achieved speedups of 19.72× over a hard ARM core.
  • Demonstrated improved throughput and reduced bottlenecks with HBM2 integration.

Abstract

Image and signal processing workloads are widely deployed on Graphics Processing Units (GPUs) for high throughput and on Field-Programmable Gate Arrays (FPGAs) for hardware specialization and energy efficiency. Soft GPU overlays on FPGAs aim to combine these advantages, yet existing solutions often depend on fixed hard processors or impose platform constraints that limit portability. This work extends a popular open-source soft GPGPU overlay to integrate a soft RISC-V control plane and enable compatibility with High-Bandwidth Memory (HBM2). The resulting system can be instantiated on FPGA boards without a hard ARM processor, improving portability, simplifying system integration, and broadening deployability. Across representative image and signal processing kernels, the soft GPGPU achieves geometric-mean speedups of 114.60 × over a scalar soft RISC-V core and 19.72 × over a hard ARM core, demonstrating substantial performance benefits while retaining FPGA reconfigurability. HBM2 integration further benefits bandwidth-sensitive workloads by increasing sustained throughput and reducing the performance bottlenecks associated with off-chip memory access. Collectively, these results indicate that GPU-like programmability and performance can be delivered on reconfigurable platforms without reliance on hard CPU subsystems, providing a portable and scalable foundation for embedded vision and DSP acceleration.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hernandez et al. (2026) studied this question.

synapsesocial.com/papers/699fe28895ddcd3a253e642chttps://doi.org/10.1016/j.sysarc.2026.103736
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1A Statically and Dynamically Scalable Soft GPGPU2024 · 1 citations
  2. 2Soft GPGPU versus IP cores: Quantifying and Reducing the Performance Gap2024
  3. 3A field‐programmable gate array‐based real‐time holographic processor for multilayer computer‐generated hologram using high‐bandwidth memory2026
  4. 4Designing an IEEE-Compliant FPU that Supports Configurable Precision for Soft Processors2024
  5. 5A Lightweight Convolution-Aware RISC-V Soft Processor for Intelligent Wearable Systems2026