Randomized trial evaluates new deep learning approach for classifying web texts, suggesting improved methods for semantic representation.
With the development of internet, vast amount of web text data in the form of news articles, blogs, social network statuses, online reports and so on have been accumulated. Automatic classification of web text accurately is a very important task in the area of Natural Language Processing and Artificial Intelligence. It enables many applications, such as web information retrieval, automatic recommendation, web filtering, sentiment analysis and so forth. The problem is that the complicated semantic structures and long distance relationships of the web text make it difficult to classify web text correctly. Existing word representations, such as Word2Vec and GloVe, generate fixed word embeddings for each word.To address the issues mentioned above, this paper proposes a more advanced deep learning architecture named BERT, BGCA, which is composed of BERT(Bidirectional Encoder Representations from Transformers), BiGRU(Bidirectional Gated Recurrent Unit), CNN(Convolutional Neural Network) and Attention in a web text classification problem. BERT is applied to extract the contextual encoding of words in the web text, which enables the model to learn the syntax and semantics between words using bidirectional transformer architecture. The BiGRU layer is used to analyse contexts in both forward and backward directions simultaneously. CNN layer responsible for the capturing of crucial information and n-gram features from the text. and the Attention is focused on capturing the importance of core words.Experiments on a large enough news dataset named THUCNews were conducted to verify the efficiency of the proposed model, namely BERT, BGCA. Experiments were conducted between traditional embedding techniques (Word2Vec, GloVe) and BERT based embedding. The comparison results indicated that the BERT resulted in a significant improvement over the static embedding approaches. The high accuracy, F1 score(95.21%, 94.36%) and F measure of 95.21% and 94.36% respectively, on the 20 categories THUCNews news dataset, show that the BERT, BGCA, compared with other deep leaning models like TextCNN, can reach a better classification results and be conducive to semantic representation. Therefore, BGCA can be concluded as an effective, scalable and practical method for large scale web text classification.
No takes yet. Share an insight, caveat, or question.
LAKSHMI et al. (2026) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: