Synapse
⌘+K
Synapse
PulseExploreClubsResearchersJournals
Instagram
HomeClubsExplore
August 30, 2026Machine LearningOpen Access

Less is More: Unlabeled Data Selection for Efficient Tabular Self-supervised Learning

View Full Paper
Ask AI
Bookmark
Share

Authors

SSSintija StevanoskaCVChristian L. Camacho VillalónSDSašo Džeroski

Discussion

Loading...

Member takes

Overview

Empirical study demonstrates that subsampling unlabeled data maintains downstream accuracy in tabular self-supervised learning, indicating substantial computational savings are achievable.

Key Points

  • To determine whether subsampling unlabeled examples using uncertainty-, diversity-, and transport-based selection criteria reduces computational costs while preserving representation quality in tabular self-supervised learning.
  • Benchmarked four self-supervised learning architectures across 25 tabular datasets under varying amounts of labeled data and label selection regimes (biased and unbiased).
  • Evaluated multiple unlabeled data subset selection strategies based on uncertainty, diversity, and transport criteria.
  • Conducted a meta-analysis linking intrinsic dataset characteristics to observed downstream performance shifts.
  • Subsampling the unlabeled data pool yielded substantial reductions in computational pretraining costs.
  • Downstream task performance was maintained or improved relative to pretraining on the complete unlabeled dataset across the evaluated benchmarks.

Cite This Study

Stevanoska et al. (2026) studied this question.

synapsesocial.com/papers/6a93f1626c1a8fb52e79e3a9https://doi.org/10.1007/s10994-026-07133-8
View Full Paper
Ask AI
Bookmark
Share