PulseExploreJournal ClubDebatesTrendingResearchersJournals
Instagram
HomeExploreJournal ClubTrending
Synapse
⌘+K
Synapse
April 8, 2026Data3 citationsOpen Access

A Survey Dataset on Student Retention in Higher Education: A Colombian Public University Case

ELErika María López-LópezOBOsnamir Elias Bru-CorderoCACristian David Correa Alvarez

Key Points

  • The research aims to identify factors affecting student retention at a Colombian public university.
  • Structured cross-sectional survey conducted between August and December 2025
  • Data collected via an online questionnaire from 333 students
  • Included variables on demographics, socioeconomic conditions, and academic experiences
  • Dataset cleaned to remove duplicates and personal information
  • Provided a machine-readable codebook and documented preprocessing steps
  • Dataset includes 33 variables related to student demographics and experiences
  • Highlighting perceived dropout intention across various domains
  • Encourages responsible data reuse and supports reproducible retention research
  • Facilitates equity-oriented analyses and benchmarking predictive models

Abstract

Student attrition remains a persistent challenge in higher education and is shaped by interacting socioeconomic, academic, institutional, and wellbeing-related mechanisms. Although learning analytics and educational data mining increasingly support early-warning and intervention workflows, dataset reuse is often limited by incomplete documentation and inconsistent variable definitions. This Data Descriptor presents a structured cross-sectional survey dataset on factors influencing student persistence at a Colombian public university campus (La Paz). Data were collected between August and December 2025 through an online questionnaire and subsequently cleaned to remove duplicate entries and personally identifiable information. The released dataset contains 333 student records and 33 variables covering demographics (e.g., age, gender, first-generation status), socioeconomic conditions (e.g., residential stratum, housing, financial aid), academic experience and satisfaction (multiple 1–5 Likert items), perceived dropout intention across personal/socioeconomic/academic domains, thematically coded open-ended items describing challenges and motives, and a self-allocation of 0–100 weights across three dropout-factor domains. We provide a machine-readable codebook, a transparent preprocessing description, and technical validation checks (value ranges, category consistency, and composite-score integrity). The dataset is intended to support reproducible retention research, equity-oriented analyses, and benchmarking of predictive models, while encouraging responsible reuse through privacy-preserving release practices and FAIR-aligned metadata, repository deposition, and versioning.

Ask AI
Helpful
Bookmark
Share
View Full Paper

Cite This Study

López-López et al. (2026) studied this question.

synapsesocial.com/papers/69d5f11e74eaea4b11a7ab21https://doi.org/10.3390/data11040075
Ask AI
Helpful
Bookmark
Share
View Full Paper