PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
August 3, 2021Proceedings of the International AAAI Conference on Web and Social Media565 citationsOpen Access

What Yelp Fake Review Filter Might Be Doing?

AMArjun MukherjeeVVVivek V. VenkataramanBLBing Liu

Key Points

  • This work aims to uncover the mechanisms behind Yelp's fake review filtering by analyzing its filtered data.
  • Analyzed Yelp's filtered reviews using a supervised learning approach.
  • Utilized both linguistic and behavioral features for evaluation.
  • Conducted information theoretic analysis to compare crowdsourced and commercial fake reviews.
  • Found that behavioral features significantly outperformed linguistic features in identifying fake reviews.
  • Identified a correlation between Yelp's filtering algorithm and abnormal spamming behaviors.
  • Achieved around 90% accuracy with some approaches, suggesting effectiveness of Yelp's filter.

Abstract

Online reviews have become a valuable resource for decision making. However, its usefulness brings forth a curse ‒ deceptive opinion spam. In recent years, fake review detection has attracted significant attention. However, most review sites still do not publicly filter fake reviews. Yelp is an exception which has been filtering reviews over the past few years. However, Yelp’s algorithm is trade secret. In this work, we attempt to find out what Yelp might be doing by analyzing its filtered reviews. The results will be useful to other review hosting sites in their filtering effort. There are two main approaches to filtering: supervised and unsupervised learning. In terms of features used, there are also roughly two types: linguistic features and behavioral features. In this work, we will take a supervised approach as we can make use of Yelp’s filtered reviews for training. Existing approaches based on supervised learning are all based on pseudo fake reviews rather than fake reviews filtered by a commercial Web site. Recently, supervised learning using linguistic n-gram features has been shown to perform extremely well (attaining around 90% accuracy) in detecting crowdsourced fake reviews generated using Amazon Mechanical Turk (AMT). We put these existing research methods to the test and evaluate performance on the real-life Yelp data. To our surprise, the behavioral features perform very well, but the linguistic features are not as effective. To investigate, a novel information theoretic analysis is proposed to uncover the precise psycholinguistic difference between AMT reviews and Yelp reviews (crowdsourced vs. commercial fake reviews). We find something quite interesting. This analysis and experimental results allow us to postulate that Yelp’s filtering is reasonable and its filtering algorithm seems to be correlated with abnormal spamming behaviors.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Mukherjee et al. (2021) studied this question.

synapsesocial.com/papers/6a0ede12b7cc3b883f22d32chttps://doi.org/10.1609/icwsm.v7i1.14389
Ask AI
Helpful
Bookmark
Share
View Full Paper

Also Consider

Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context:

  1. 1Fake Reviews Detection using Supervised Machine Learning Algorithms2024
  2. 2Machine Learning-Based Fake Online Review Comment Detection2024
  3. 3Unmasking the Deception: A Focused Survey of Machine Learning Techniques for Fake Review Detection2024
  4. 4Using Machine Learning Techniques to Identify False Opinion Spam in Online Reviews2026
  5. 5CyberSentinel: Fake Product Review Detection Using Machine Learning2026