AMD AI Engines (AIEs) extend the design space and open up new options for coarse-grained processing in re-configurable accelerators. Pure FPGA designs for machine learning often struggle to compete with the high clock frequencies of GPUs for data-intensive workloads with only limited control flow. Having AIEs available on-chip with an FPGA fabric allows for low-latency co-processing and permits parts of an application to be placed on the most suitable kind of processing unit. Many data-heavy workloads, particularly in the AI domain, benefit from data streaming. With TaPaSCo-AIE, we present a framework for heterogeneous systems centered around data streams. Our framework focuses on AMD Versal devices and incorporates AI Engines and 100G network. We demonstrate the efficient use of TaPaSCo-AIE in a real-world evaluation based on a neural network, and achieve significant performance improvements over CPUs, and even exceed the performance of an A100 GPU.
No takes yet. Share an insight, caveat, or question.
Heinz et al. (2024) studied this question.
Synapse has enriched 3 closely related papers on similar clinical questions. Consider them for comparative context: