Full text
Corresponding author: Adedoyin S. Adebanjo Copyright © 2025 Author(s) retain the copyright of this article. This article is published under the terms of the Creative Commons Attribution License 4.0. A logistics regression-based student performance prediction system Adedoyin Samuel Adebanjo 1, *, Chiamaka G. Anyanwu 2, Babajide E. Adeoti 1, Emmanuel Mgbeahuruike 1 and Emmanuel I. Oyerinde 3 1 Department of Software Engineering, Babcock University, Ilishan, Nigeria. 2 Department of Computer Science, Babcock University, Ilishan, Nigeria. 3 Department of Information Technology, Babcock University, Ilishan, Nigeria. Global Journal of Engineering and Technology Advances, 2025, 25(01), 193-200 Publication history: Received on 18 September 2025; revised on 25 October 2025; accepted on 27 October 2025 Article DOI: https://doi.org/10.30574/gjeta.2025.25.1.0311 Abstract Predicting student performance has become an important focus in educational data mining. Schools are using datadriven insights to find learners who may be struggling and to improve academic success. This study uses the Logistic Regression model on student performance data to examine how demographic, behavioral, and academic factors affect learning outcomes. Logistic Regression predicts outcomes like pass/fail and high/low performance with strong accuracy, due to its probabilistic framework. Besides predicting, the model also helps explain the importance of different factors. This allows educators to create informed and targeted interventions. The results show that Logistic Regression is an effective model that strikes a good balance between accuracy and clarity, making it a useful tool for early warning systems and data-based decision-making in higher education. Keywords: Logistics Regression; Machine Learning; Student Performance Prediction; Supervised Learning 1. Introduction In recent years, the education sector in Nigeria has seen a significant decline in student performance in both internal and external exams. Poor academic results have been reported at all levels of education [1]. A recent report from the Joint Admission and Matriculation Board (JAMB) showed that about 1.4 million candidates scored below 200 in the 2024 Unified Tertiary Matriculation Examination (UTME). This highlights the seriousness of the problem [2]. While this number pertains to Nigeria, student dropout is a global issue. Dropout rates exceed 40% in some European and Latin American countries, and around 30% of first-year students in U.S. Baccalaureate Institutions do not return for their second year [3]. The United Nations has pointed out that poverty, poor infrastructure, limited funding, and low quality of life in urban slums are major factors leading to poor educational outcomes in developing countries [4]. These challenges often result in delayed graduation, low self-esteem, and, in extreme cases, withdrawal from academic programs [3]. This trend highlights the urgent need for effective solutions that can improve both student retention and performance. One promising method is using predictive analytics and data mining techniques to identify at-risk students. This allows for proactive academic support [5]. Student Performance Prediction (SPP) is becoming an important area of research in Educational Data Mining (EDM) and Machine Learning. It goes beyond grades and looks at the skills and knowledge needed for academic and societal success [6]. By using historical and behavioral data, SPP offers useful insights for everyone involved. Students can plan their learning paths, instructors can adjust their teaching methods, and institutions can improve retention with targeted support [6].
Global Journal of Engineering and Technology Advances, 2025, 25(01), 193-200 194 Several studies have looked into the potential of SPP using various machine learning methods. Al Husaini and Shukor [7] conducted a review of literature from 2014 to 2020. They identified a wide range of internal factors, like entry grades and family support, as well as external factors, such as socioeconomic background and e-learning engagement, that influence academic outcomes. They noted that female students generally show more persistence. Similarly, Tjandra et al. [8] analyzed over 250 studies on SPP. They found that most of these studies focused on monitoring learning activities (67.2%). Fewer studies (9.2%) addressed dropout prevention. Feng et al. [9] demonstrated the use of machine learning to predict academic performance, along with standardized tests, teacher ratings, and classroom observations. With more schools using digital technologies and having access to a lot of student data, Educational Data Mining (EDM) offers new chances to gather knowledge and improve learning outcomes. This study adds to this expanding area by using supervised machine learning algorithms, specifically Logistic Regression, to predict student performance in Nigerian schools. The insights gained are meant to help with curriculum design, guide academic interventions, and improve overall student success in both internal and external examinations. The aim of the study is to develop a student performance prediction system using Logistic Regression model that predicts student performance and student Cumulative Grade Point Average (CGPA) respectively. Objectives The specific objectives of the study are • To evaluate the relationship between student performance and their grades. • To identify critical features that contribute to student performance. • To design and develop a Logistic Regression model to predict student performance. • Evaluate the model accuracy and performance. 2. Literature review 2.1. Educational Data Mining (EDM) and Learning Analytics (LA) Educational Data Mining (EDM) and Learning Analytics (LA) have become important methods for tackling issues like student retention, dropout rates, and overall performance in higher education. EDM uses computational and statistical techniques to find meaningful patterns in large datasets created by students. Meanwhile, LA focuses on interpreting this data to generate actionable insights [10], [11]. Romero and Ventura [12] examined how EDM can predict academic performance. They emphasized the use of classification, clustering, and association rules. These methods help educational institutions identify risks of failure, suggest personalized learning paths, and support adaptable curricula. Similarly, Siemens and Long [11] explained that LA allows institutions to monitor students in real time, which enables early interventions to enhance outcomes. Both EDM and LA are now critical in developing predictive models for student outcomes, enhancing resource allocation, and designing interventions for retention and success [10]-[14]. Frameworks in EDM typically emphasize predictive modeling using algorithms such as Support Vector Machines (SVM), Random Forests (RF), and Gradient Boosting (GB), while LA frameworks focus on systematic data collection and reporting to inform pedagogical decisions [10], [13]-[15]. Hybrid approaches that integrate EDM and LA are increasingly popular, providing comprehensive pipelines that include preprocessing, feature selection, model validation, and deployment for predictive and recommendation systems [15]. The relationship between EDM and LA lays the groundwork for predictive analytics in education. This combination equips stakeholders with data-driven strategies aimed at improving learning outcomes and the effectiveness of institutions. 2.2. Predictive Modeling in Education Predictive modeling techniques are increasingly used to estimate student performance indicators like grade point average (GPA), exam scores, and dropout risk. Earlier models mostly relied on regression analysis, but recent studies show a growing preference for machine learning methods, including random forests, gradient boosting, and support vector machines [16]-[21]. These algorithms effectively capture nonlinear relationships, making them a good fit for complex educational data.
Global Journal of Engineering and Technology Advances, 2025, 25(01), 193-200 195 However, "black-box" models have a drawback: they are hard to interpret. Because of this, regression-based approaches still hold value, especially in academic settings where transparency and explanation matter as much as accuracy [13], [15], [22]-[25]. Logistic regression, in particular, strikes a balance between predictive performance and interpretability. It helps stakeholders pinpoint the relative importance of factors like academic preparation, socioeconomic background, and learning engagement in determining student success. 2.3. Logistic Regression in Student Performance Prediction Logistic regression (LR) is commonly used to predict binary outcomes like pass/fail, success/dropout, or retention/nonretention. The logistic function that forms the basis of LR is expressed as ℎ𝜃(𝑥)=1 1+𝑒−𝜃𝑇𝑥 (I) where 𝜃 represents the parameter vector, and 𝑥 is the input feature vector [24], [26]. The model parameters are improved by reducing the cost function 𝐽(𝜃)= − 1 𝑚∑[𝑦(𝑖)log(ℎ𝜃(𝑥(𝑖)))+(1−𝑦(𝑖))log(1−ℎ𝜃(𝑥(𝑖)))] 𝑚 𝑖=1 +𝜆 2𝑚∑𝜃𝑗2𝑛 𝑗=1 (II) Here, the inclusion of the regularization term 𝜆 2𝑚∑𝜃𝑗2𝑛 𝑗=1 helps prevent overfitting, which is common in highdimensional educational datasets [25], [27]. Logistic regression is a common method in educational data mining for binary classification tasks, like predicting whether a student will succeed or fail. It calculates the probability of an input belonging to a specific class, usually by applying a logistic function to a linear combination of features. The output can be seen as a probability, and by setting a threshold, often at 0.5, instances are sorted into one of two categories. The model is simple and easy to understand. It can also include regularization techniques like Lasso (L1) or Ridge (L2), making it effective for dealing with noisy or small datasets often found in educational research. These benefits keep logistic regression a popular option for tasks such as predicting student performance and identifying at-risk students early [24], [27]. Logistic regression is popular in educational research because it can handle categorical variables and provide clear model parameters. This helps educators make decisions based on data. The model is also affordable to run, offers probabilities, and is flexible for both binary and multiclass situations. This makes it useful for predicting student engagement and timely submissions. However, it has some drawbacks. It can overfit, especially with small datasets, and it struggles to capture complex, nonlinear relationships found in educational data. While logistic regression is transparent, this can sometimes lower predictive accuracy compared to more complex models like deep learning or random forests. Additionally, problems with data quality and the model’s limited assumptions may lessen its effectiveness in various educational settings [13], [25], [28]. 3. Methodology This study followed a structured approach to create and use a predictive model for student performance through Logistic Regression. The process included the following stages: data collection, data pre-processing, model development, model evaluation, and deployment. Each stage is explained in detail below. 3.1. Data Collection A reliable dataset is crucial for effective model training. This study used the GTZAN dataset as the main source because it's widely recognized in music genre classification. GTZAN has 1,000 audio clips evenly spread across 10 genres, such as jazz, rock, classical, and blues. This provides a balanced representation for evaluation. To address the limited size of labeled data and improve model robustness, data augmentation techniques were applied. These include • Pitch Shifting: Adjusting the pitch without changing the tempo. • Time Stretching: Modifying the playback speed while maintaining pitch. • Adding Noise: Introducing synthetic noise to simulate real-world recording conditions. These techniques create synthetic variations of the dataset, increasing diversity and enhancing generalization.
Global Journal of Engineering and Technology Advances, 2025, 25(01), 193-200 196 3.2. Data Pre-processing Feature extraction bridges the gap between raw audio and machine learning models by converting audio signals into structured representations. Using Python's Librosa library, the following features were extracted: • Data cleaning: Addressing missing values, fixing inconsistencies, and removing irrelevant attributes. • Feature selection: Finding which variables are most relevant for predicting performance. • Transformation: Encoding categorical variables and normalizing data when needed to improve model stability. 3.3. Data Splitting To ensure a fair evaluation of the predictive model, the dataset was divided using the 75/25 ratio into two parts: a training set for fitting the model and a testing set for checking its performance. This split is crucial to how well a model generalizes to new, unseen data and avoid overfitting, which occurs when a model memorizes the training data instead of learning general patterns. Table 1 Description of part of the dataset # COLUMN TYPE DESCRIPTION 0 Age int Student’s Age 1 Gender object Student’s Gender 2 Family_income object Family’s Income 3 Parent_education_level object Parent’s Education Level 4 School object Student’s Faculty 5 Level int Student’s Level 6 100_level_cgpa int Student’s CGPA at 100 level 7 Student’s CGPA at 100 level int Student’s CGPA at current level 3.4. Model Development The Logistic Regression model was created using the Scikit-learn Python library. The training aimed to reduce prediction error by adjusting model parameters through repeated optimization. The logistic function and cost function, with a regularization term, guided this process to balance accuracy and generalization. 3.5. Model Evaluation After training, the model was tested on the test set, with accuracy as the main performance measure. Additional evaluations looked at true positives, false positives, true negatives, and false negatives to give a full performance assessment. Functions from Scikit-learn were used to calculate these metrics and ensure reliable interpretation of the results. 3.6. Data Visualization Exploratory data visualization was conducted using Seaborn in Python. Pair plots were created to show the relationships between academic and lifestyle features. These visualizations helped to identify correlations, trends, and potential predictor variables that could impact student performance. 3.7. Model Deployment The final phase involved deploying the trained model with Django, a Python web framework. The Logistic Regression model was serialized using Pickle to allow it to be stored as a binary file and reused within the web interface. This integration lets users enter student attributes and receive real-time predictions about academic risks, which supports timely interventions.
Global Journal of Engineering and Technology Advances, 2025, 25(01), 193-200 197 Figure 1 System Architecture Figures 2 and 3 presents the results of the model. Fig 2 shows the confusion matrix for the Logistics Regression algorithm, which describes the performance of the algorithm by summarizing the correct and incorrect predictions for each category. Fig. 3 demonstrates the feature importance tool which was added to provide insights on features which have significant impact on the prediction model. The model was evaluated to determine how well the model will perform on new, unseen data. The evaluation metrics used in this project is accuracy, precision, recall and F1-score. These metrics help evaluate the model by analysing the number of true positives, false positives and false negatives. With the use of functions from Scikit-learn Python package, the metrics results are as follows: • Accuracy: 0.95 • F1 Score: 0.91 • Precison:0.96 • Recall: 0.89
Global Journal of Engineering and Technology Advances, 2025, 25(01), 193-200 198 Figure 2 Confusion Matrix for Logistics Regression Figure 3 Feature Importance for Logistics Regression When compared to other classification algorithms such as Random Forest, Support Vector Machine and Gradient Boosting Regresso, Logistic Regression showed a better degree of accuracy of 95% making it the best choice model for predicting student performance. This makes the system a veritable tool for educators to not just recognize students in danger of failing but also understand patterns and variables that affect their performance. 4. Conclusion This study successfully developed a Student Performance Prediction System that applies machine learning algorithms and techniques to predict student’s grade based on previous academic performance, behavioural attributes, family relationship and demographic information. Logistic Regression was used to predict overall student performance based on selected features and the overall performance rating of students. It has a high accuracy of 95% making it a tool for
Global Journal of Engineering and Technology Advances, 2025, 25(01), 193-200 199 educators in identifying students who are at danger of failing and providing them with required support to enhance academic support. The results demonstrated that machine learning algorithms can effectively classify student into different performance categories, helping educators and administrators make decisions. Compliance with ethical standards Disclosure of conflict of interest No conflict of interest to be disclosed. References [1] S. Tete and O. M. Wizoma, “Education in Nigeria: Challenges and Way Forward,” vol. 8, no. 1, pp. 42–48, 2020. [2] D. Tolu-Kolawole, “1.4 million UTME candidates scored below 200 – JAMB,” Punch, Apr. 29, 2024. Accessed: Sep. 6, 2025. [Online]. Available: https://punchng.com/1-4-million-utme-candidates-scored-below-200-jamb/. [3] F. Ofori, E. Maina, and R. Gitonga, “Using machine learning algorithms to predict students’ performance and improve learning outcome: A literature-based review,” Journal of Information and Technology, vol. 4, no. 1, pp. 33–55, 2020. [4] P. Sokkhey and T. Okazaki, “Developing web-based support systems for predicting poor-performing students using educational data mining techniques,” Int. J. Adv. Comput. Sci. Appl., vol. 11, no. 7, 2020, doi: 10.14569/IJACSA.2020.0110704. [5] A. F. Núñez-Naranjo, “Analysis of the determinant factors in university dropout: a case study of Ecuador,” Frontiers in Education, vol. 9, Art. no. 1444534, 2024, doi: 10.3389/feduc.2024.1444534. [6] R. Cariaga, “What is student performance?” J. Uniq. Crazy Ideas, vol. 1, no. 1, pp. 42–46, Aug. 2024, doi: 10.5281/zenodo.13410584. [7] Y. Al Husaini and N. S. Ahmad Shukor, “Factors affecting students' academic performance: A review,” Res Militaris: European Journal of Military Studies, vol. 12, pp. 284–294, 2023 [8] E. Tjandra, S. S. Kusumawardani, and R. Ferdiana, “Student performance prediction in higher education: A comprehensive review,” AIP Conference Proceedings, vol. 2468, Art. no. 050005, 2022, doi: 10.1063/5.0080187. [9] G. Feng, M. Fan, and Y. Chen, “Analysis and prediction of students’ academic performance based on educational data mining,” IEEE Access, vol. 10, pp. 19558–19571, 2022, doi: 10.1109/ACCESS.2022.3151652. [10] R. H. Ali, “Educational data mining for predicting academic student performance using active classification,” Iraqi Journal of Science, vol. 63, no. 9, pp. 3954–3965, 2022. [11] G. Siemens and P. Long, “Penetrating the fog: Analytics in learning and education,” EDUCAUSE Review, vol. 46, no. 5, pp. 30–40, 2011. [12] C. Romero and S. Ventura, “Educational data mining: A review of the state of the art,” IEEE Transactions on Systems, Man, and Cybernetics, Part C (Applications and Reviews), vol. 40, no. 6, pp. 601–618, 2010. [13] F. E. El Habti, M. Hiri, M. Chrayah, A. Bouzidi, and N. Aknin, “Enhancing student performance prediction in elearning ecosystems using machine learning techniques,” International Journal of Information and Education Technology, vol. 15, no. 2, pp. 301–311, 2025. [14] M. Kumar, V. Bhardwaj, D. Thalkral, A. Rashid, and M. T. B. Othman, “Ensemble learning-based model for student performance prediction,” Ingénierie des Systèmes d’Information, vol. 29, no. 5, pp. 1925–1935, 2024, doi: 10.18280/isi.290524. [15] R. Bertolini, S. J. Finch, and R. H. Nehm, “Enhancing data pipelines for forecasting student performance: Integrating feature selection with cross-validation,” International Journal of Educational Technology in Higher Education, vol. 18, no. 1, Art. no. 45, 2021, doi: 10.1186/s41239-021-00279-6. [16] P. Chakrapani and D. Chitradevi, “Academic performance prediction using machine learning: A comprehensive and systematic review,” in Proc. 2022 Int. Conf. Electron. Syst. Intell. Comput. (ICESIC), 2022, pp. 335–340, doi: 10.1109/ICESIC53714.2022.9783512.
Global Journal of Engineering and Technology Advances, 2025, 25(01), 193-200 200 [17] H. Goss, “Student learning outcomes assessment in higher education and in academic libraries: A review of the literature,” J. Acad. Librariansh., vol. 48, no. 2, Art. no. 102485, 2022, doi: 10.1016/j.acalib.2021.102485. [18] S. F. A. Hossain, Z. Xi, M. Nurunnabi, and B. Anwar, “Sustainable academic performance in higher education: A mixed method approach,” Interact. Learn. Environ., vol. 30, no. 4, pp. 707–720, 2019, doi: 10.1080/10494820.2019.1680392. [19] B. Sekeroglu, R. Abiyev, A. Ilhan, M. Arslan, and J. B. Idoko, “Systematic literature review on machine learning and student performance prediction: Critical gaps and possible remedies,” Appl. Sci., vol. 11, no. 22, Art. no. 10907, 2021, doi: 10.3390/app112110907. [20] H. Sahlaoui, E. A. A. Alaoui, A. Nayyar, S. Agoujil, and M. M. Jaber, “Predicting and interpreting student performance using ensemble models and Shapley additive explanations,” IEEE Access, vol. 9, pp. 152688– 152703, 2021, doi: 10.1109/ACCESS.2021.3124270. [21] Y. Qu, F. Li, L. Li, X. Dou, and H. Wang, “Can we predict student performance based on tabular and textual data?,” IEEE Access, vol. 10, pp. 86008–86019, 2022, doi: 10.1109/ACCESS.2022.3198682. [22] F. Afrin, M. Hamilton, and C. Thevathyan, “On the explanation of AI-based student success prediction,” in Proc. Int. Conf. Comput. Sci., 2022, pp. 252–258, doi: 10.1007/978-3-031-08754-7_34. [23] P. Balaji, S. Alelyani, A. Qahmash, and M. Mohana, “Contributions of machine learning models towards student academic performance prediction: A systematic review,” Appl. Sci., vol. 11, no. 21, Art. no. 10007, 2021, doi: 10.3390/app112110007. [24] A. Kord, A. Aboelfetouh, and S. Shohieb, “Academic course planning recommendation and student performance prediction multi-modal based on educational data mining techniques,” J. Comput. Higher Educ., 2025, doi: 10.1007/s12528-024-09426-0. [25] H. Farhood, I. Joudah, A. Beheshti, and S. Muller, “Evaluating and enhancing artificial intelligence models for predicting student learning outcomes,” Informatics, vol. 11, no. 3, Art. no. 46, pp. 1–17, 2024, doi: 10.3390/informatics11030046. [26] M. Yağcı, “Educational data mining: Prediction of students' academic performance using machine learning algorithms,” Smart Learn. Environ., vol. 9, no. 11, Art. no. 192, 2022, doi: 10.1186/s40561-022-00192-z. [27] N. Alruwais and M. Zakariah, “Evaluating student knowledge assessment using machine learning techniques,” Sustainability, vol. 15, no. 7, Art. no. 6229, pp. 1–25, 2023, doi: 10.3390/su15076229. [28] A. Merchant, N. Shenoy, A. Bharali, and M. A. Kumar, "Predicting students' academic performance in virtual learning environment using machine learning," 2022 Second International Conference on Power, Control and Computing Technologies (ICPC2T), 2022, pp. 1-6, doi: 10.1109/ICPC2T53885.2022.9777008.