scieee AI-readable full text Open interactive document viewer

Predictive Modelling of Student Outcomes Using Ensemble Regression and Classification Methods

Mary Teresa; et. al.

Abstract

Accurate prediction of student academic outcomes is vital for developing data-driven interventions in education. This study proposes a robust ensemble learning framework based on the HistGradientBoostingClassifier (HGB) to classify student grades using behavioral and academic features such as self-study hours, attendance, class participation, and total performance scores. Leveraging a large-scale synthetic dataset of 1,000,000 student records, we benchmarked the proposed HGB model against widely used ensemble classifiers including XGBoost, LightGBM, CatBoost, and Random Forest. Comprehensive experiments demonstrated that HGB consistently outperformed all baselines, achieving a testing accuracy of 99.6%, with macro-averaged precision, recall, and F1-score of 0.99. The model also showed strong generalization across both majority and minority grade categories, as confirmed by confusion matrix analysis. These results highlight the effectiveness of histogram-based boosting in educational data mining and support its application in real-time academic performance monitoring and intervention systems.

Full text

ISSN (Online): 2583-5696 Int. Jr. of Hum Comp. & Int. © The Author(s) 2026. Published by Milestone Research Publications. This article is licensed under the Creative Commons Attribution 4.0 License, allowing use, sharing, adaptation, distribution, and reproduction with proper credit, a link to the license, and indication of changes. Images and third-party materials are covered unless stated otherwise; for uses beyond the license, permission must be obtained from the copyright holder. Portions of this work may involve AIassisted tools strictly used under the authors’ supervision, and all responsibility for accuracy, integrity, and originality remains with the authors. License: http://creativecommons.org/licenses/by/4.0/ 664 RESEARCH ARTICLE OPEN ACCESS Predictive Modelling of Student Outcomes Using Ensemble Regression and Classification Methods Mary Teresa1,* . Sukerthi Sutraya2 . Y Vijaya Sambhavi3 . Saritha Dasari4. J K Neelima5 1,* Department of CSE-AIML, Guru Nanak Institutions Technical Campus, Hyderabad, India 2Department of Computer Science and Engineering (Data Science), G. Narayanamma Institute of Technology and Science, Hyderabad, India. 3Department of EEE, Annamacharya Institute of Technology and Sciences (Autonomous), Tirupati, India. 4Department of Computer Science and Engineering (Data Science), G. Narayanamma Institute of Technology and Science, Hyderabad, India. 5Department of E.C.E, Narayana Engineering College, Nellore, India. DOI: 10.5281/zenodo.17599070 Received: 12 October 2025 / Revised: 04 November 2025 / Accepted: 13 November 2025 *Corresponding Author: [email protected] ©Milestone Research Publications, Part of CLOCKSS archiving Abstract –Accurate prediction of student academic outcomes is vital for developing datadriven interventions in education. This study proposes a robust ensemble learning framework based on the HistGradientBoostingClassifier (HGB) to classify student grades using behavioral and academic features such as self-study hours, attendance, class participation, and total performance scores. Leveraging a large-scale synthetic dataset of 1,000,000 student records, we benchmarked the proposed HGB model against widely used ensemble classifiers including XGBoost, LightGBM, CatBoost, and Random Forest. Comprehensive experiments demonstrated that HGB consistently outperformed all baselines, achieving a testing accuracy of 99.6%, with macro-averaged precision, recall, and F1-score of 0.99. The model also showed strong generalization across both majority and minority grade categories, as confirmed by confusion matrix analysis. These results highlight the effectiveness of histogram-based boosting in educational data mining and support its application in real-time academic performance monitoring and intervention systems. Index Terms – Student Performance Prediction, Ensemble Learning, HistGradientBoostingClassifier, Educational Data Mining, Multiclass Classification, Machine Learning. 665 I. INTRODUCTION Institutions in higher education are increasingly adopting data-driven approaches to enhance learning outcomes and promote student success, driven by the rapid growth of learning analytics and educational data mining (EDM). Central to these efforts is the use of predictive modeling, which enables the early identification of at-risk learners and facilitates the design of timely interventions to maximize academic impact. Early works in this field, such as Kabakchieva [1], used test results and first-year admission scores as measures of student performance to emphasize the importance of data mining methods. Building on this framework, Sekeroglu et al. [2] highlight the importance of data refinement and preprocessing as essential prerequisites for accurate prediction results. Later research has broadened the modeling field. The resilience of Support Vector Machines (SVMs) in forecasting secondary school outcomes, as demonstrated by Ng et al. [3], further confirms that past academic success is a reliable indicator of future performance. In their use of AutoML systems, Zeineddine et al. [4] furthered this line of research, demonstrating how automation may improve preadmission performance prediction and reduce dropout risks. More recently, Alshamaila et al. [5] employed imbalance treatment techniques and deep learning frameworks to address the issues of data imbalance in real-world educational settings. Alsariera et al. [6] provided a thorough assessment of the area to integrate these disparate advancements, describing the efficacy of algorithms such as SVM, ANN, and DT. Reiterating the multifaceted nature of forecasting student achievement, their research also highlighted the predictive usefulness of academic, demographic, and family-related characteristics. Despite such advancements, core concerns still need to be tackled. Most of the currently existing approaches either focus on classification (categorical outputs) or regression (numeric outputs) without trying to merge both schools of thought into providing a balanced image of student performance. Furthermore, issues of data imbalance, heterogeneity of features, and non-generalizability across institutions are still challenges to the use of predictive models in real learning environments. Our motivation in this work is to address these gaps by building a robust ensemble system that combines regression and classification problems to optimize student outcome predictive modeling more effectively. By combining the strength of ensemble learning and leveraging multiple sets of features, our approach aims to generate more accurate, explainable, and scalable predictions, eventually informing student success interventions in a timely manner. The primary contributions of this paper are the following: ● We introduce an ensemble prediction modeling framework using HistGradientBoostingClassifier (HGB), which includes histogram-based feature binning to facilitate acceleration enhancement and improved scalability. ● Our strategy is a combination of regression (continuous performance measures) and classification (A–F grades) so that overall student performance is reflected. ● The HGB model is evaluated using a simulated 1,000,000 student records dataset and compared to XGBoost, LightGBM, CatBoost, and Random Forest. ● Experimental outcomes presented that HGB strongly outperforms baselines with 99.6% accuracy for macro-averaged precision, recall, and F1-score of 0.99. 666 II. LITERATURE SURVEY To support early intervention efforts, improve learning outcomes, and inform institutional decision-making, predicting students' academic performance has been a significant focus of educational data mining (EDM) and machine learning research. Numerous studies have examined a wide range of methods, from sophisticated hybrid and ensemble techniques to conventional supervised learning algorithms. These articles discuss the importance of feature selection, dataset heterogeneity, and innovative approaches in producing reliable results, as well as the necessity of sufficient predictive modeling. With particular reference to hybrid models, ensemble approaches, neural networks, and optimization techniques for enhancing predictability and interpretability, the chapter discusses groundbreaking advances in the field. Aulakh et al. [7] introduced a hybrid ensemble model employing Extreme Learning Machine (ELM) and Random Forest (RF) for the prediction of student academic performance. In the model, the ELM component was used because of its fast learning and effectiveness at large-scale input feature mapping, while RF was used because of its high accuracy in classification as well as its capability in modeling intricate interactions between features. This complementary E-RF approach enhanced predictive accuracy as well as computational efficiency, as they demonstrated through their experiment study. Zafari et al. [8] suggested a machine learning system for examining the performance of high school students and predicting their classification into four labels: very good, good, medium, and bad. They considered a dataset collected through questionnaires and teacher evaluations of 459 students with demographic, behavior, and grade features. They employed Random Forest (RF), Support Vector Machines (SVM), Logistic Regression (LR), and Artificial Neural Networks (ANN), feature selection by Boruta algorithm. Results showed that ANN was best for the original feature set but following elimination of low-importance features, SVM was best with an accuracy of 0.78. Student's grades and class attention were found to be most important, and feature set reduction also increased the efficiency of the models further and made them feasible for real-world deployment. Feng et al. [9] presented a hybrid educational data mining method that combines clustering and deep learning for student academic performance modeling and prediction. They extended the traditional K-means clustering algorithm with an objective statistic for determining the optimal number of clusters, improving the stability of clustering results. The clusters of data were then used as category labels to train a Convolutional Neural Network (CNN) for making predictions. Experimental results showed that the approach had greater accuracy compared to traditional score-based analysis and facilitated early academic warnings for at-risk students. Their framework highlights how the combination of unsupervised (clustering) and supervised (CNN) methods can lead to a more objective and effective way of evaluating and predicting student performance. Asselman et al. [10] proposed a enhanced Performance Factors Analysis (PFA) model for the prediction of student performance by incorporating ensemble learning methods, namely Random Forest, AdaBoost, and XGBoost. They focused on the technical aspect of prediction rather than purely pedagogical, evaluating the models on three different datasets. The results showed that XGBoost performed better than traditional PFA and other ensemble models with high scalability and predictability. This research demonstrates the benefit of using advanced ensemble techniques to advance knowledge tracing and adaptive educational systems. 667 Agrawal and Mavani [11] employed Neural Networks to foresee student performance and also juxtaposed its usefulness against Bayesian classification. They emphasized the importance of various features—viz. grades, medium of instruction, and family background—influencing performance, and determined that prior academic performance has a strong correlation with later performance. Through experimentation on engineering student semester data, they demonstrated neural networks to be more predictive than Bayesian classification, particularly on large datasets. Their work highlighted both the need for proper feature selection and the promise of neural networks for capturing non-linear relationships in student performance prediction. Hashim et al. [12] suggested a supervised machine learning student performance prediction method that utilized demographic, academic, and behavior features as input parameters. They evaluated a collection of algorithms—Decision Tree, Naïve Bayes, Logistic Regression, SVM, K-Nearest Neighbors, Sequential Minimal Optimization, and Neural Networks—on the dataset of University of Basra undergraduate students. Among them, Logistic Regression worked best with very good precision for passed (68.7%) and failed (88.8%) students in final examinations. The results highlighted the necessity of feature-based supervised approaches in predicting academic performance and established logistic regression as a baseline classifier for performance prediction. Ahmed [13] proposed a machine learning approach to predict student performance by integrating K-means clustering, Random Forest feature selection, and supervised learning algorithms (SVM, Decision Trees, KNN, Naïve Bayes) with hyper parameter tuning and repeated cross-validation on 32,005 Wollo University students' data. Results were SVM most accurate (96%), followed by Decision Trees (93.4%), KNN (87.4%), and Naïve Bayes (83.3%), with tuning greatly improving performance. Demographic imbalances were also revealed via the study, such as female and regional influences on outcomes. It shows overall how the combination of clustering, feature engineering, and optimized ML models improves the trustworthiness of predictions, but results remain institution-specific. Bhutto et al. [14] suggested a supervised machine learning solution to student academic performance prediction based on Support Vector Machines (SVM) and Logistic Regression, which was tested on academic datasets. Experiments showed that the Sequential Minimal Optimization (SMO)–based SVM performed better than Logistic Regression in terms of higher accuracy in classifying students as good or poor performers. The study also showed emphasis on quantifying determinants such as teachers' performance and students' motivation, with a suggestion on how predictive modelling can be implemented to reduce college dropout rates and guide focused intervention in college. A machine learning-based predictive model for students' academic performance was proposed by [15] to facilitate early intervention by advisors and teachers. Their technique was designed to anticipate final test scores so that students who were most likely to fail may receive individualized academic help and course recommendations from experts. They claimed that their method was 94.88% accurate in its predictions, demonstrating how both students and teachers could utilize it to enhance academic planning and reduce failure rates. From the literature reviewed, it emerges that predictive modelling in education has extended to the application of more than the traditional classification algorithms, and ensemble methods, hybrid models, [16]clustering-based methods, and deep learning models have also been used. These have improved accuracy, scalability, and have provided more stable results in predicting academic performance. Feature selection, pre-processing, and optimization techniques emerge as crucial steps with a consistent effect across various datasets. Furthermore, the incorporation of behavioural, institutional, and demographic variables alongside academic metrics increases the range of predictions and practical 668 usefulness. Overall, the literature supports the fact that machine learning holds tremendous potential for enhancing early warning systems, enabling data-informed educational planning, as well as advancing students' success. III. METHODS & MATERIALS In this section, we briefly describe the dataset description, data preprocessing and feature engineering for reliable experimental results. We also mention the overall methodology employed for our research. Figure 1 illustrates the overall research approach followed in this work. Fig. 1: Graphical representation of the overall research methodology A. Dataset Description For this experiment, we used a Kaggle dataset named Student Performance Dataset, which is a simulated but realistic dataset designed to facilitate Machine Learning (ML) novices to practice fundamental principles of predictive modeling. The dataset contains 1,000,000 rows, where each row represents one student. The dataset offers a clean and structured environment to practice regression, classification, and model evaluation techniques. The data set includes six important variables: student_id, weekly_self_study_hours, attendance_percentage, class_participation, total_score, and grade. The features are typical scholastic variables that affect student performance. The total_score is a continuous target variable obtained from weekly_self_study_hours by adding random noise to represent natural variations in study habits and personal performance. The grade is a categorical label (A to F) obtained from the total score using some thresholds, enabling classification tasks. The weekly_self_study_hours variable is produced on a standard normal curve with a mean of 15 hours per week, and attendance_percentage and class_participation are produced independently with useful variance. These variables create a more useful dataset for multivariate regression models and allow students to observe how inputs of different types assist in achieving academic performance. It is specifically designed to be newbie-friendly with understandable, but not simplistic, input-output 669 relationships. It allows for a variety of ML experiments like basic and multiple linear regression, grade classification, and testing using MAE, RMSE, and R². Both continuous and categorical targets are supported, allowing users to try a vast range of supervised learning features in a controlled and understandable framework. A detailed analysis of all dataset features is presented in Table 1 below. Table. 1: Summary of Features in the Student Performance Dataset Column Name Description student_id Unique identifier for each student (numeric) weekly_self_study_hours Weekly self-study hours (0–40, normal distribution, mean ≈ 15) attendance_percentage Attendance rate in percentage (50–100, normal distribution, mean ≈ 85) class_participation Class activity score (0–10, mean ≈ 6) total_score Final performance score (0–100, function of study hours + noise) grade Categorical grade label derived from total_score (A, B, C, D, F) B. Data Preprocessing and Feature Engineering To ready the dataset for predictive modeling and to maintain consistency in classification outcomes, an official data preprocessing and feature engineering pipeline was implemented on the Student Performance Dataset. This dataset includes 1,000,000 student entries with six characteristics that measure behavioral and academic factors, such as study effort, attendance, participation, and academic success. The preprocessing involved cleaning, transforming, encoding, and conducting exploratory analysis to prepare the data for classification. 1. Missing Value Analysis: The initial check using. isnull() revealed that the data contains no missing values in any of the features (weekly_self_study_hours, attendance_percentage, class_participation, total_score, grade). This allowed use of the entire data without imputation, preserving the integrity and natural distribution of all the features. 2. Feature Selection and Target Definition: The student_id column was recognized as an uninformative unique identifier and was dropped from modeling to prevent any data leakage or perplexing variance. The remaining features were classified as follows: Numerical Predictors: a. weekly_self_study_hours (float): Average self-study hours on a weekly basis b. attendance_percentage (float): Average attendance rate across all classes c. class_participation (float): Classroom activity participation score d. total_score (float): Continuous performance score mapped between 0 and 100 Categorical Target Variable: e. grade (object): Academic grade (A, B, C, D, F), utilized as the class label when performing classification 670 2. Label Encoding of Target Variable: To prepare the categorical target variable grade for classification algorithms, label encoding was applied. Each letter grade (A–F) was mapped to a corresponding integer as follows: A → 0, B → 1, C → 2, D → 3, and F → 4. This transformation ensured the variable was suitable for tree-based models such as XGBoost and LightGBM, which require numeric inputs. 3. Train-Test Split for Model Validation: To ensure unbiased evaluation of the classification models, the dataset was partitioned using an 80:20 train-test split. The random state was fixed to 42 for reproducibility. The training set was used to build the model, while the test set provided an unbiased estimate of predictive performance. Let D be the full dataset, X the feature matrix after dropping student_id and grade, and y the encoded grade target: (𝑋𝑡𝑟𝑎𝑖𝑛, 𝑋𝑡𝑒𝑠𝑡, 𝑦𝑡𝑟𝑎𝑖𝑛, 𝑦𝑡𝑒𝑠𝑡) = 𝑆𝑝𝑙𝑖𝑡( 𝑋, 𝑦, 𝑡𝑒𝑠𝑡−𝑠𝑖𝑧𝑒 = 0.2, 𝑟𝑎𝑛𝑑𝑜𝑚−𝑠𝑡𝑎𝑡𝑒 = 42) 4. Exploratory Feature Distribution and Correlation Analysis: To better understand the relationships among features and their influence on the target label, the following visual analyses were conducted: ● A scatterplot revealed clustering of grades based on combinations of weekly_self_study_hours and attendance_percentage, highlighting their discriminative power. ● Histogram plots showed the distribution of attendance_percentage, which was roughly normally distributed with a mean near 85%. ● Bar plots and heatmaps indicated a positive correlation between weekly_self_study_hours and total_score (Pearson r ≈ 0.78), confirming that self-study hours are a strong predictor of academic performance. C. Methodology To identify students' academic performance on behavioral traits such as study time, attendance, and participation in classes, we employed an ensemble of machine learning algorithms. We chose the models based on their established success in handling high-dimensional, tabular data with non-linear relationships and class distributions that are imbalanced. We sought to validate each model's predictive capability, generalizability, and stability across different bands of performance (grades A through F). 1. Baseline Models: As a foundational benchmark, we implemented a diverse suite of ensemble machine learning algorithms to evaluate baseline performance and understand data separability using traditional methods. Each model was selected based on its proven utility for structured, tabular data with imbalanced multiclass targets. ● XGBoost Classifier (XGB): XGBoost is a highly optimized boosting algorithm that sequentially builds trees to correct the residuals of prior models: 𝐹 𝑚(𝑥) = 𝐹𝑚−1 (𝑥) + 𝜂ℎ𝑚(𝑥) 671 XGB achieved a training accuracy of 99.84% and a test accuracy of 99.72%, but exhibited slight overfitting, particularly for dominant classes like A and B. While it performed well on the majority of classes, its recall and precision dropped for underrepresented grades, such as D and F. Despite its speed and configurability, XGBoost required careful regularization and hyperparameter tuning to avoid variance issues. ● CatBoost Classifier: CatBoost, a gradient boosting framework designed for categorical features, was evaluated on this numerical dataset for its boosting strengths. It achieved a test accuracy of 91.3%. Although it handled imbalanced class distributions slightly better than Random Forest, its performance on edge grades (particularly F) was still limited. Since the dataset did not contain raw categorical features, CatBoost’s main advantage—ordered boosting with categorical encoding—was underutilized. The model showed better class separation than RF but fell behind in overall accuracy and consistency. ● Random Forest (RF): Random Forest is a bagging-based ensemble that builds multiple decision trees and combines their predictions through majority voting: 𝑦 = 𝑚𝑜𝑑𝑒 (𝑇1(𝑥), 𝑇2(𝑥),.............,𝑇𝐾(𝑥)) The RF model achieved a test accuracy of 90.1%, demonstrating acceptable but clearly suboptimal performance compared to boosting techniques. Its relatively shallow understanding of non-linear class boundaries led to frequent misclassifications in borderline grades like C and D. Moreover, it suffered from underfitting, as evidenced by its lower recall for minority classes. Despite being robust to overfitting and simple to implement, RF lacked the precision needed for a fine-grained, multiclass academic prediction task. ● LightGBM Classifier (LGBM): LightGBM is known for its histogram-based optimization and leaf-wise tree growth. It achieved training and testing accuracies of 99.84% and 99.72%, respectively—similar to XGBoost. However, it demonstrated higher variance in cross-validation and a slightly higher error rate on misclassified samples from lower grade categories. While efficient and fast, LightGBM was more sensitive to data imbalance and required parameter tuning for improved recall in minority classes. 2. Proposed Model: HistGradientBoostingClassifier (HGB): To effectively classify student academic grades based on behavioral metrics such as self-study hours, attendance, and participation, we propose the use of the HistGradientBoostingClassifier (HGB) — a histogram-based gradient boosting framework optimized for large-scale tabular data. Unlike traditional boosting algorithms that rely on exact greedy splitting, HGB discretizes continuous features into bins, allowing for faster training and reduced memory usage without compromising predictive performance. 672 ● Core Principle of HGB HistGradientBoosting builds an additive ensemble of decision trees in a forward stage-wise fashion. At each iteration m, the model fits a new tree ℎ𝑚(𝑥) to the negative gradients (also called pseudo-residuals) of the loss function from the previous iteration: 𝐹 𝑚(𝑥) = 𝐹𝑚−1 (𝑥) + 𝜂ℎ𝑚(𝑥) Where: ● 𝐹 𝑚(𝑥) : the boosted prediction at stage mmm ● η: the learning rate ● ℎ𝑚(𝑥) : the weak learner fitted to residuals The histogram-based strategy partitions continuous features into discrete bins (e.g., 255 by default), which reduces computational complexity from O(n.logn) to O(b.log⁡b) where b≪n. This makes HGB highly scalable on large datasets, such as our 1-million-row student performance dataset. ● Model Configuration and Training: The proposed model was trained using the following settings: ● Model: HistGradientBoostingClassifier ● Loss Function: Multinomial deviance (for multiclass classification) ● Class Balancing: class_weight='balanced' to handle grade imbalance ● Random Seed: random_state=42 to ensure reproducibility ● Training Data Size: 80% of the dataset (800,000 samples) ● Test Data Size: 20% of the dataset (200,000 samples) During training, the model learned complex, non-linear relationships between the input features and the multiclass target variable (grades A to F). Class balancing ensured equitable learning across underrepresented categories like grade D and grade F, which are typically challenging in skewed datasets. ● Performance Evaluation The proposed HGB model demonstrated exceptional performance, outperforming all baseline models in both accuracy and class-wise metrics. Key results are summarized below: ● Training Accuracy: 99.65% ● Testing Accuracy: 99.61% ● Macro-Averaged F1 Score: 0.99 ● Cross-Validation Accuracy (Mean): 99.48% ● Cross-Validation Std Dev: 0.0000196 The confusion matrix confirmed that HGB maintained high recall and precision across all five grade categories (A–F), with particularly strong performance even in minority classes such as grade 'F'.