PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
September 12, 2025Systems3 citationsOpen Access

Performance and Efficiency Gains of NPU-Based Servers over GPUs for AI Model Inference

View Full Paper
YHYoung‐Pyo HongDKDongsoo Kim

Key Points

  • NPU servers demonstrate equivalent or superior throughput compared to GPU servers, reducing power usage by 35–70%.
  • Performance metrics, including latency and energy efficiency, were systematically analyzed in various AI inference scenarios.
  • Optimizing NPUs with the vLLM library increased tokens-per-second and power efficiency significantly by 92%.
  • The findings suggest NPU-based architectures present a sustainable alternative to traditional GPU systems.

Abstract

The exponential growth of AI applications has intensified the demand for efficient inference hardware capable of delivering low-latency, high-throughput, and energy-efficient performance. This study presents a systematic, empirical comparison of GPU- and NPU-based server platforms across key AI inference domains: text-to-text, text-to-image, multimodal understanding, and object detection. We configure representative models—LLama-family for text generation, Stable Diffusion variants for image synthesis, LLaVA-NeXT for multimodal tasks, and YOLO11 series for object detection—on a dual NVIDIA A100 GPU server and an eight-chip RBLN-CA12 NPU server. Performance metrics including latency, throughput, power consumption, and energy efficiency are measured under realistic workloads. Results demonstrate that NPUs match or exceed GPU throughput in many inference scenarios while consuming 35–70% less power. Moreover, optimization with the vLLM library on NPUs nearly doubles the tokens-per-second and yields a 92% increase in power efficiency. Our findings validate the potential of NPU-based inference architectures to reduce operational costs and energy footprints, offering a viable alternative to the prevailing GPU-dominated paradigm.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hong et al. (2025) studied this question.

synapsesocial.com/papers/68d44a3731b076d99fa5375ehttps://doi.org/10.3390/systems13090797
Ask AI
Helpful
Bookmark
Share
View Full Paper