Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
October 13, 2025Open Access

AMQ: Enabling AutoML for Mixed-precision Weight-Only Quantization of Large Language Models

View Full Paper
Ask AI
Bookmark
Share

Authors

SLS. H. LeeSWSeung-taek WooJJJungyu Jin

Discussion

Loading...

Member takes

Overview

Automated Mixed-Precision Weight-Only Quantization optimizes model quality and memory usage, suggesting new pathways for deploying LLMs under constraints.

Key Points

  • AMQ achieves high-performance large language models by balancing memory usage and model quality efficiently.
  • The framework explores over 10^{100} configurations, employing search space pruning to enhance performance.
  • Innovative components like quantization proxy and quality predictor minimize computational overhead during optimization.
  • Iterative search-and-update strategies ensure fast convergence towards the Pareto frontier of quality and efficiency.

Cite This Study

Lee et al. (2025) studied this question.

synapsesocial.com/papers/68ed1896f29694dd1da78ac3https://doi.org/10.48550/arxiv.2509.12019
View Full Paper
Ask AI
Bookmark
Share