PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
October 3, 20241 citations

LLM Survey Analysis Using Random Forest

View Full Paper
PSPriyanka Singla

Key Points

Key points are not available for this paper at this time.

Abstract

This project investigates the application of a Random Forest Classifier for analyzing metadata from survey papers on large language models (LLMs), a rapidly growing area within AI. The goal is to assist new researchers by providing insights into the trends and patterns in LLM survey publications. Through a structured workflow—comprising data loading, exploration, manipulation, and visualization—key attributes such as release dates, categories, and taxonomies were analyzed. Techniques like TF-IDF vectorization, one-hot encoding, and feature scaling were employed to construct a robust feature matrix. Hyperparameter tuning using grid search optimized the classifier’s performance. Although the model achieved perfect training accuracy, a lower test accuracy (0.39) indicated overfitting, likely caused by dataset imbalance. With a best cross-validation score of 0.26, future improvements will focus on addressing data imbalance, enhancing feature engineering, and exploring alternative models to boost performance. The project highlights trends in LLM research and suggests paths for enhancing model accuracy.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

Priyanka Singla (2024) studied this question.

synapsesocial.com/papers/68e55ee7e2b3180350efbe39https://doi.org/10.31224/3956
Ask AI
Helpful
Bookmark
Share
View Full Paper