PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 16, 20221,043 citationsOpen Access

LAION-5B: An open large-scale dataset for training next generation image-text models

CSChristoph SchuhmannRBRomain BeaumontRVRichard Vencu

Key Points

  • To create and openly distribute a multi-billion-scale dataset of image-text pairs to democratize the training and analysis of large-scale multimodal models.
  • Extracted and filtered 5.85 billion image-text pairs using CLIP embeddings, including 2.32 billion English-language pairs and multilingual subsets.
  • Computed nearest neighbor search indices alongside safety and quality detection scores for watermarks, NSFW content, and toxicity.
  • Validated the utility of the dataset by replicating and fine-tuning foundational vision-language architectures including CLIP, GLIDE, and Stable Diffusion.
  • Constructed the largest openly available multimodal dataset comprising 5.85 billion filtered pairs, breaking prior exclusivity of private multi-billion-scale training sets.
  • Demonstrated effective downstream transfer, zero-shot classification robustness, and text-guided image synthesis using models trained directly on the dataset.

Abstract

Groundbreaking language-vision architectures like CLIP and DALL-E proved the utility of training on large amounts of noisy image-text data, without relying on expensive accurate labels used in standard vision unimodal supervised learning. The resulting models showed capabilities of strong text-guided image generation and transfer to downstream tasks, while performing remarkably at zero-shot classification with noteworthy out-of-distribution robustness. Since then, large-scale language-vision models like ALIGN, BASIC, GLIDE, Flamingo and Imagen made further improvements. Studying the training and capabilities of such models requires datasets containing billions of image-text pairs. Until now, no datasets of this size have been made openly available for the broader research community. To address this problem and democratize research on large-scale multi-modal models, we present LAION-5B - a dataset consisting of 5.85 billion CLIP-filtered image-text pairs, of which 2.32B contain English language. We show successful replication and fine-tuning of foundational models like CLIP, GLIDE and Stable Diffusion using the dataset, and discuss further experiments enabled with an openly available dataset of this scale. Additionally we provide several nearest neighbor indices, an improved web-interface for dataset exploration and subset generation, and detection scores for watermark, NSFW, and toxic content detection. Announcement page https://laion.ai/laion-5b-a-new-era-of-open-large-scale-multi-modal-datasets/

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Schuhmann et al. (2022) studied this question.

synapsesocial.com/papers/6a7cea3b537c44c1bef02e6chttps://doi.org/10.48550/arxiv.2210.08402
Ask AI
Helpful
Bookmark
Share
View Full Paper