PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 1, 201794 citationsOpen Access

TL;DR: Mining Reddit to Learn Automatic Summarization

MVMichael VölskeMPMartin PotthastSSShahbaz Syed

Key Points

Key points are not available for this paper at this time.

Abstract

Recent advances in automatic text summarization have used deep neural networks to generate high-quality abstractive summaries, but the performance of these models strongly depends on large amounts of suitable training data. We propose a new method for mining social media for author-provided summaries, taking advantage of the common practice of appending a "TL;DR" to long posts. A case study using a large Reddit crawl yields the Webis-TLDR-17 corpus, complementing existing corpora primarily from the news genre. Our technique is likely applicable to other social media sites and general web crawls.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Völske et al. (2017) studied this question.

synapsesocial.com/papers/6a1043e92badbc352affa82chttps://doi.org/10.18653/v1/w17-4508
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Teaching Machines to Read and Comprehend2015 · 1,939 citations
  2. 2Teaching Machines to Read and Comprehend2015 · 1,519 citations
  3. 3Automatic generic document summarization based on non-negative matrix factorization2008 · 154 citations
  4. 4Applying regression models to query-focused multi-document summarization2010 · 166 citations
  5. 5LCSTS: A Large Scale Chinese Short Text Summarization Dataset2015 · 327 citations