PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 17, 201541 citationsOpen Access

Unsupervised Feature Selection on Data Streams

HHHao HuangSYShinjae YooSKShiva Prasad Kasiviswanathan

Key Points

Key points are not available for this paper at this time.

Abstract

Massive data streams are continuously being generated from sources such as social media, broadcast news, etc., and typically these datapoints lie in high-dimensional spaces (such as the vocabulary space of a language). Timely and accurate feature subset selection in these massive data streams has important applications in model interpretation, computational/storage cost reduction, and generalization enhancement. In this paper, we introduce a novel unsupervised feature selection approach on data streams that selects important features by making only one pass over the data while utilizing limited storage. The proposed algorithm uses ideas from matrix sketching to efficiently maintain a low-rank approximation of the observed data and applies regularized regression on this approximation to identify the important features. We theoretically prove that our algorithm is close to an expensive offline approach based on global singular value decompositions. The experimental results on a variety of text and image datasets demonstrate the excellent ability of our approach to identify important features even in presence of concept drifts and also its efficiency over other popular scalable feature selection algorithms.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Huang et al. (2015) studied this question.

synapsesocial.com/papers/6a10f548cfa01e990d9fe2f7https://doi.org/10.1145/2806416.2806521
Ask AI
Helpful
Bookmark
Share
View Full Paper