PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
January 1, 202044 citationsOpen Access

Domain adaptation challenges of BERT in tokenization and sub-word representations of Out-of-Vocabulary words

ANAnmol NayakHTHariprasad TimmapathiniKPKarthikeyan Ponnalagu

Key Points

Key points are not available for this paper at this time.

Abstract

BERT model However, it still has several research challenges which are not tackled well for domain specific corpus found in industries. In this paper, we have highlighted these problems through detailed experiments involving analysis of the attention scores and dynamic word embeddings with the BERT-Base-Uncased model. Our experiments have lead to interesting findings that showed: 1) Largest substring from the left that is found in the vocabulary (in-vocab) is always chosen at every sub-word unit that can lead to suboptimal tokenization choices, 2) Semantic meaning of a vocabulary word deteriorates when found as a substring in an Out-Of-Vocabulary (OOV) word, and 3) Minor misspellings in words are inadequately handled. We believe that if these challenges are tackled, it will significantly help the domain adaptation aspect of BERT.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Nayak et al. (2020) studied this question.

synapsesocial.com/papers/6a1b90720ea968f653abf9dfhttps://doi.org/10.18653/v1/2020.insights-1.1
Ask AI
Helpful
Bookmark
Share
View Full Paper