Digital academic platforms increasingly shape how researchers share ideas, communicate results, and sustain collaborations. However, computational studies of scientific discourse remain concentrated on high-resource languages and Western contexts, leaving emerging research communities underexplored. This study investigates sentiment analysis in digital scientific communications produced in the Democratic Republic of the Congo. We collected 26,480 raw texts from Google Scholar, ResearchGate, LinkedIn, X, and public WhatsApp research groups. Because manual annotation was not feasible at this scale, we used a weak-supervision strategy to generate sentiment labels and retained a high-confidence subset of 747 samples for model training and evaluation. This subset should be understood as an initial benchmark rather than a fully representative sample of Congolese scientific discourse. The retained corpus is predominantly French (95%), with limited Swahili and Lingala content. We evaluated term frequency-inverse document frequency with logistic regression, convolutional neural networks, long short-term memory networks, and transformer-based models, while also incorporating a contextual uncertainty signal to improve prediction interpretability. Among the evaluated approaches, BERT (bidirectional encoder representations from transformers) achieved the best overall performance, with a Matthews correlation coefficient of 0.831 and a Macro-F1 score of 0.605, outperforming the classical and neural baselines. These findings indicate that transformer-based models are promising for sentiment analysis in French-dominant Congolese scientific communication, while the moderate Macro-F1 score and limited multilingual coverage highlight the need for larger, more balanced, and externally validated datasets.
Kanduki et al. (Fri,) studied this question.
Synapse has enriched 5 closely related papers on similar clinical questions. Consider them for comparative context: