PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
June 13, 20125,675 citationsOpen Access

Practical Bayesian Optimization of Machine Learning Algorithms

JSJasper SnoekTwitter (United States)HLHugo LarochelleMila - Quebec Artificial Intelligence InstituteRARyan P. AdamsPrinceton University

Key Points

  • To develop a practical, automated Bayesian optimization framework using Gaussian processes to tune hyperparameters in machine learning algorithms more effectively than human experts.
  • Modeled generalization performance using Gaussian process priors and tractable posterior distributions to guide iterative hyperparameter selection.
  • Introduced novel Bayesian optimization algorithms that account for variable evaluation runtimes and support parallelization across multiple processing cores.
  • Evaluated the optimization framework across multiple complex models, including latent Dirichlet allocation, structured support vector machines, and convolutional neural networks.
  • Demonstrated that careful choices of Gaussian process priors and inference procedures significantly improve optimization efficacy, reaching or exceeding human expert-level tuning.
  • Showed that the proposed cost-aware and parallelized algorithms outperform previous automated tuning methods across all evaluated machine learning architectures.

Abstract

Machine learning algorithms frequently require careful tuning of model hyperparameters, regularization terms, and optimization parameters. Unfortunately, this tuning is often a "black art" that requires expert experience, unwritten rules of thumb, or sometimes brute-force search. Much more appealing is the idea of developing automatic approaches which can optimize the performance of a given learning algorithm to the task at hand. In this work, we consider the automatic tuning problem within the framework of Bayesian optimization, in which a learning algorithm's generalization performance is modeled as a sample from a Gaussian process (GP). The tractable posterior distribution induced by the GP leads to efficient use of the information gathered by previous experiments, enabling optimal choices about what parameters to try next. Here we show how the effects of the Gaussian process prior and the associated inference procedure can have a large impact on the success or failure of Bayesian optimization. We show that thoughtful choices can lead to results that exceed expert-level performance in tuning machine learning algorithms. We also describe new algorithms that take into account the variable cost (duration) of learning experiments and that can leverage the presence of multiple cores for parallel experimentation. We show that these proposed algorithms improve on previous automatic procedures and can reach or surpass human expert-level optimization on a diverse set of contemporary algorithms including latent Dirichlet allocation, structured SVMs and convolutional neural networks.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Snoek et al. (2012) studied this question.

synapsesocial.com/papers/69dc870d7873f5f05b1334afhttps://doi.org/10.48550/arxiv.1206.2944
Ask AI
Helpful
Bookmark
Share
View Full Paper