Student attrition remains a persistent challenge in higher education and is shaped by interacting socioeconomic, academic, institutional, and wellbeing-related mechanisms. Although learning analytics and educational data mining increasingly support early-warning and intervention workflows, dataset reuse is often limited by incomplete documentation and inconsistent variable definitions. This Data Descriptor presents a structured cross-sectional survey dataset on factors influencing student persistence at a Colombian public university campus (La Paz). Data were collected between August and December 2025 through an online questionnaire and subsequently cleaned to remove duplicate entries and personally identifiable information. The released dataset contains 333 student records and 33 variables covering demographics (e.g., age, gender, first-generation status), socioeconomic conditions (e.g., residential stratum, housing, financial aid), academic experience and satisfaction (multiple 1–5 Likert items), perceived dropout intention across personal/socioeconomic/academic domains, thematically coded open-ended items describing challenges and motives, and a self-allocation of 0–100 weights across three dropout-factor domains. We provide a machine-readable codebook, a transparent preprocessing description, and technical validation checks (value ranges, category consistency, and composite-score integrity). The dataset is intended to support reproducible retention research, equity-oriented analyses, and benchmarking of predictive models, while encouraging responsible reuse through privacy-preserving release practices and FAIR-aligned metadata, repository deposition, and versioning.
López-López et al. (2026) studied this question.