PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
July 14, 2026Artificial Intelligence Review0 citationsOpen Access

Evaluating large language model compression: a comparative analysis on state-of-the-art models across diverse hardware platforms

DHDominik HildebrandBKBenjamin KieferAZAndreas Zell

Key Points

  • This work systematically compares compression techniques for large language models to assess their performance across different hardware platforms.
  • Empirical comparison of quantization, pruning, and parameter-efficient fine-tuning across various model families.
  • Evaluation used benchmarks, deployment metrics, and settings to capture trade-offs in real-world scenarios.
  • Models evaluated included open-source families like Llama, Mistral, Phi, and Qwen at scales from 1.7 to 70 billion parameters.
  • Quantization provided optimal deployment feasibility but needed careful tuning to prevent throughput declines.
  • Pruning lowered parameter counts significantly but often led to substantial performance loss beyond moderate sparsity levels.
  • PEFT methods allowed for competitive performance, matching larger models on benchmarks while reducing storage and overhead.

Abstract

Abstract This work presents a systematic, empirical comparison of contemporary compression techniques for large language models (LLMs), namely quantization, pruning, and parameter-efficient fine-tuning (PEFT) using a representative set of open-source model families (Llama, Mistral, Phi and Qwen) and model scales (1.7 Billion to 70 Billion). Evaluation combined benchmarks (MMLU, SQuAD v2, TinyBenchmarks and WikiText), deployment metrics (peak memory, time-to-first-token, tokens/sec and maximum sequence lengths) and settings (multi-GPU clusters, single-GPU PC, laptop, and smartphone) to capture real-world trade-offs. Quantization often delivered the best wins for deployment feasibility—enabling single-device and mobile inference—but required careful per-model tuning and backend support to avoid throughput regressions. Pruning reduced parameter counts substantially but frequently incured large, even catastrophic, performance loss beyond moderate sparsity levels. Retraining partially mitigated this but did not uniformly close the gap to quantization. Finally, PEFT methods enabled models to match or outperform models with up to 18 times the parameters on SQuAD v2 while reducing storage as well as optimizer overhead and often improved task performance even when full fine-tuning failed.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Hildebrand et al. (2026) studied this question.

synapsesocial.com/papers/6a55d11a5aafca87247f8383https://doi.org/10.1007/s10462-026-11614-6
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Evaluating Quantized Large Language Models2024 · 9 citations
  2. 2Efficient Compression of Large Language Models: A Case Study on Llama 2 with 13B Parameters2024 · 18 citations
  3. 3A Quantization Approach for the Reduced Size of Large Language Models2024 · 4 citations
  4. 4A Comprehensive Evaluation of Quantization Strategies for Large Language Models2024 · 5 citations
  5. 5Contemporary Model Compression on Large Language Models Inference2024 · 5 citations