scieee AI-readable full text Open interactive document viewer

PRIVITOR: A Privacy-Preserving Intelligent Proctoring Framework for Online Exams

Khan, R.; Suleiman, B.; Yaqub, W.; Sun, J.

Abstract

The rise of online learning has been accelerated by several factors, including technological advancements, the need for lifelong education, and global events such as the COVID-19 pandemic. As a result, numerous universities have significantly expanded their online education offerings. This shift has created a growing demand for effective online examination methods, particularly for software engineering courses that require coding tests. The widespread adoption of online exams has increased the need for online exam proctoring. However, this transition has raised significant concerns regarding potential privacy violations and intrusive surveillance practices. Despite the existence of frameworks attempting to address these issues, there remains a pressing need for a system that effectively preserves student privacy without compromising the integrity and fairness of online assessments. We propose PRIVITOR, a comprehensive proctoring framework addressing these issues. Our contributions include: (1) a novel approach to collect data for training proctoringspecific machine learning models, (2) an efficient anomaly detection classifier with an associated cheating detection algorithm, and (3) an innovative facial masking technique for privacy-preserving proctor-student interaction. Results show that our anomaly detection classifier achieves high accuracy while processing videos approximately ten times faster than existing eye-tracking algorithms. The facial masking technique effectively balances privacy protection and invigilation capabilities.

Full text

Research Paper Recommended citation: Khan, R., Suleiman, B., Yaqub, W., & Sun, J. (2025). PRIVITOR: A Privacy-Preserving Intelligent Proctoring Framework for Online Exams. In Kangaslampi, R., Langie, G., Järvinen, H.-M., & Nagy, B. (Eds.), SEFI 53rd Annual Conference. European Society for Engineering Education (SEFI), Tampere, Finland. DOI: 10.5281/zenodo.17631836. This Conference Paper is brought to you for open access by the 53rd Annual Conference of the European Society for Engineering Education (SEFI) at Tampere University in Tampere, Finland. This work is licensed under a Creative Commons Attribution-NonCommercial-Share Alike 4.0 International License. PRIVITOR: A PRIVACY-PRESERVING INTELLIGENT PROCTORING FRAMEWORK FOR ONLINE EXAMS Rasim Khana , Basem Suleiman a,1, Waheeb Yaquba, Jinglin Sunb a University of Sydney, Australia, Sydney, Australia b University of New South Wales, Sydney, Australia Conference Key Areas: Digital tools and AI in engineering education, Engineering ethics education Keywords: Privacy-preservation, Proctoring, Cheating Detection, Exams ABSTRACT The rise of online learning has been accelerated by several factors, including technological advancements, the need for lifelong education, and global events such as the COVID-19 pandemic. As a result, numerous universities have significantly expanded their online education offerings. This shift has created a growing demand for effective online examination methods, particularly for software engineering courses that require coding tests. The widespread adoption of online exams has increased the need for online exam proctoring. However, this transition has raised significant concerns regarding potential privacy violations and intrusive surveillance practices. Despite the existence of frameworks attempting to address these issues, there remains a pressing need for a system that effectively preserves student privacy without compromising the integrity and fairness of online assessments. We propose PRIVITOR, a comprehensive proctoring framework addressing these issues. Our contributions include: (1) a novel approach to collect data for training proctoringspecific machine learning models, (2) an efficient anomaly detection classifier with an associated cheating detection algorithm, and (3) an innovative facial masking technique for privacy-preserving proctor-student interaction. Results show that our anomaly detection classifier achieves high accuracy while processing videos approximately ten times faster than existing eye-tracking algorithms. The facial masking technique effectively balances privacy protection and invigilation capabilities. 1 Corresponding Author B. Suleiman [email protected] 1 INTRODUCTION The COVID-19 pandemic accelerated the shift toward online education, making remote learning increasingly common across global higher education institutions. While this transition has expanded educational access and reduced geographical barriers (Selwyn et al., 2021), it presents significant challenges for maintaining academic integrity during examinations. Online environments create opportunities for misconduct through digital collaboration, secondary device usage, impersonation, and content sharing—particularly in software engineering courses, where students need computing environments to write and test code. A fundamental tension exists between academic integrity and student privacy in online examination environments. Studies show that over 75% of students find traditional online proctoring intrusive (Balash et al., 2021), expressing discomfort with the recording of their screens and personal spaces. Multiple studies confirm these privacy risks associated with collecting and storing personal information (Nigam et al., 2021; Coghlan et al., 2021; Milone et al., 2017). Though eliminating supervision might address privacy concerns, research by Dendir and Maxwell (2020) demonstrates that unproctored exams yield significantly higher scores, suggesting potential academic integrity issues. Automated systems offer a partial solution but cannot consistently detect misconduct without human oversight (Nigam et al., 2021; Singh et al., 2021). Lee and Fanguy (2022) argue that current systems may compromise educational practices, highlighting the urgent need for solutions protecting both integrity and privacy. To address these challenges, we propose PRIVITOR, a privacy-preserving framework for online examination proctoring that combines efficient automated detection with privacy-aware human verification. The framework integrates high-speed anomaly detection with privacy-preserved monitoring, where automated screening identifies potential irregularities and human proctors review only flagged segments through privacy-protecting facial masks. The key contributions of this paper are: (1) A web-based framework for collecting training data, enabling systematic recording of facial expressions and eye movements; (2) A two-stage monitoring approach combining an anomaly detection classifier with behaviour pattern analysis, achieving 85.40% accuracy while processing videos ten times faster than existing methods; and (3) A privacy-focused review process applying facial masking to flagged video segments, retaining essential behavioural indicators while obscuring personal identifiers. 2 RELATED WORK Online proctoring systems typically employ either automated AI-based monitoring or human supervision (Arnò et al., 2021). While automated systems better protect student privacy, they struggle with accuracy, leading to false cheating accusations (Singh et al., 2021; Holden et al., 2021; Nigam et al., 2021). Hybrid approaches combining both methods show promise but raise privacy concerns. Yaqub et al. (2022) proposed a privacy-preserving framework where webcam recordings are automatically analysed by comparing frame hashes against anchors, with anomalous frames flagged for human review. Their approach masks students' faces except for the eyes, which presents two critical limitations: eyes themselves contain sensitive biometric information (Liebling & Preibusch, 2014; Steil et al., 2019), and important facial expressions that help detect deceptive behavior are obscured (Ozdamli et al., 2022; Shen et al., 2021; Zanette et al., 2016). Atoum et al. (2017) moved beyond anomaly detection by developing a classifier-based approach that extracts high-level features like gaze estimation from exam recordings. Their Support Vector Machine classifier, though innovative, relied on staged cheating scenarios for training data, raising questions about its effectiveness in authentic exam environments (Holden et al., 2021). Head pose analysis has become fundamental in detecting potential academic dishonesty. While Irfan et al. (2021) implemented a complex FSA-net solution for webcam-based detection, Cote et al. (2016) demonstrated that simplifying head movements into five basic directions (front, up, down, left, right) is sufficient for proctoring purposes. Although recent research has advanced precise head pose estimation using sophisticated neural networks (Bulat & Tzimiropoulos, 2017; Gupta et al., 2019; Wu et al., 2021; Yang et al., 2015; Zhou & Gregson, 2020), such granularity exceeds proctoring requirements where students typically face webcams directly (Yaqub et al., 2022). Eye gaze estimation provides crucial insights into student attention and intent (Cheng et al., 2021). While specialized hardware solutions have limitations in detecting subtle cheating behaviors (Atoum et al., 2017), appearancebased gaze estimation offers more comprehensive monitoring capabilities. However, most implementations require calibration steps (Chen & Shi, 2023), creating practical barriers for widespread adoption. Despite available datasets like MPIIGaze (Zhang et al., 2019) and Gaze360 (Kellnhofer et al., 2019), a simplified approach without calibration requirements, trained specifically on proctoring contexts, would better serve online exam monitoring systems. 3 METHODOLOGY This section presents PRIVITOR, our privacy-preserving online proctoring framework that balances examination integrity with student privacy. Our approach integrates three key components: a novel data collection methodology yielding high-quality training data, a two-stage detection system combining anomaly classification with specialized cheating detection, and a facial masking technique that enables human review while safeguarding student privacy. 3.1 Data Collection We developed a web-based framework to facilitate remote data collection from participants. The platform presents users with a selection of facial or eye movements to perform, providing both written instructions and example animations for clarity. Each recording session captures 7 seconds of video: 2 seconds of the participant maintaining a central position (establishing a baseline), followed by 5 seconds performing the specified movement. Expanding on Cote et al. (2016), we classified movements into eight directional categories for both facial and eye movements: Left, Right, Up, Down, Upper Left, Upper Right, Lower Left, and Lower Right. To distinguish between intentional movements and incidental head displacements, we implemented a standardisation process that creates scale-invariant representations by computing normalised coordinates within a facial bounding box. To ensure the quality of training labels during data collection, we implemented a twostep verification system to confirm that recorded videos contained the specific behaviours requested from participants. This process used MediaPipe to track 468 facial landmarks and established a baseline reference frame showing each participant in a neutral, forward-facing position. Head movement verification focused on tracking nose tip coordinates, as any head rotation or tilt produces measurable nose displacement. We identified the eight frames showing the greatest nose displacement from the baseline and calculated directional offset values. Samples were validated when movement in the horizontal or vertical axis exceeded predetermined thresholds and aligned with the prompted direction (e.g., "turn left" or "look up"). Non-compliant samples were automatically discarded. Eye movement verification employed a more comprehensive landmark analysis across four distinct facial regions: left eye, right eye, left iris, and right iris. We calculated centre coordinates for each region by averaging their constituent landmark positions, then measured displacement from the baseline frame. Total eye movement was determined by summing displacement measurements across all regions. Samples qualified as valid only when displacement was both statistically significant and directionally consistent with the intended gaze instruction. Following verification, samples were classified into two categories: anomalous behaviours (containing the specified movements) and normal behaviours (maintaining the central forward-facing position). 3.2 Two-Stage Cheating Detection The PRIVITOR framework implements a two-stage approach for identifying potential academic dishonesty: the anomaly detection classifier and the cheating detection algorithm. 3.2.1 Anomaly Detection Classifier Our anomaly detection classifier identifies unusual head and eye orientations in exam recordings. Trained on a dataset containing 468 facial landmarks per sample with 8 samples per movement type from each participant, the classifier processes 128 anomalous and 16 non-anomalous samples per participant. The anchor frame labeled "Centre" represents non-anomalous movement. The classifier employs a Support Vector Machine (SVM), which determines the optimal decision boundary to separate normal and anomalous behaviours in the high-dimensional landmark space. By leveraging the comprehensive facial positioning data, the SVM effectively distinguishes between acceptable exam behaviours and suspicious movements that may indicate academic misconduct. 3.2.2 Cheating Detection Algorithm The cheating detection algorithm serves as a secondary filter that analyses output from the Anomaly Detection Classifier to identify potential academic misconduct. This algorithm focuses on patterns of recurring directional movements that may indicate attempts to access unauthorized materials during exams. We introduced a threshold parameter (ɑ), set to 0.125 in our implementation, which flags behaviours occurring at higher-than-random frequency among detected anomalous movements. This Algorithm 1: Cheating Detection Algorithm 1ω→0.125; 2direction tracker →[]; 3for sample ↑anomalous samples do 4direction tracker[sample.direction]→direction tracker[sample.direction]+1; 5for direction ↑direction tracker do 6direction ratio →direction tracker[direction]/length(anomalous samples); 7if direction ratio < ωthen 8remove direction from direction tracker; 9cheating samples →[]; 10 for sample ↑anomalous samples do 11 if sample.direction ↑direction tracker then 12 add sample to cheating samples; 13 direction tracker[sample.direction]→direction tracker[sample.direction]+1; parameter can be adjusted to match examination requirements—increasing it creates a more lenient review process. This two-stage approach significantly reduces the volume of video requiring manual review, though human verification remains necessary. Rather than eliminating the need for proctors, the system directs their attention to the most suspicious behaviors, improving both efficiency and effectiveness of the proctoring process. 3.3 Privacy-Preserving Facial Mask Fig. 1. Face Mask Comparison While automated cheating detection reduces manual review requirements, human proctor verification remains necessary, raising privacy concerns that previous approaches like Yaqub et al.'s (2022) white mask inadequately addressed by either concealing crucial expressions or exposing sensitive eye regions. Our novel facial mask (Figure 1, Featured Mask) balances privacy protection with effective proctoring by retaining essential facial expressions while masking the eyes, overlaying annotations on masked eye regions using MediaPipe-detected facial landmarks superimposed on a full-face white mask. 3.4 Online Proctoring framework The PRIVITOR framework processes student exam recordings through a sequential pipeline. First, the directed anomaly detection classifier isolates and labels anomalous segments, which the cheating detection algorithm then refines to identify suspicious behaviours. These segments undergo privacy-preserving facial masking before being presented to human proctors for final evaluation, creating an efficient system that maintains exam integrity while protecting student privacy. 4 EXPERIMENTS AND RESULTS This section examines the performance of our directed anomaly classifier and evaluates facial mask comparison responses. Currently, 10 test participants have contributed data through our website, with plans to expand to 500 participants. 4.1 Performance of Anomaly Detection Classifier Our anomaly detection system was developed using 1,368 images and an SVM classifier. We employed 3-fold cross-validation with grid search to identify the optimal kernel for our SVM classifier, testing linear, polynomial, and radial basis function (RBF) kernels. Figure 2 shows that the linear kernel achieved the highest cross-validation accuracy of 82.16%, outperforming both polynomial and RBF alternatives. This outcome indicates that our feature space is largely linearly separable (Varatharajah et al., 2019). Using an 80-20 stratified train-test split to maintain class balance, the final model achieved 85.40% accuracy on the test dataset. Figure 3 illustrates the classification performance across movement categories, revealing that face-related movements demonstrated superior precision and recall compared to eye movements due to their larger facial landmark displacements and greater linear separability. Fig. 2. Grid Search Results for SVM Kernel Fig. 3. Heatmap of Classification Report for SVM Model Fig. 4. Class Prediction Error of SVM Model Fig. 5. Confusion Matrix of SVM Model Fig. 6. ROC Curve of Model The model exhibited systematic error patterns that reveal important limitations in anomaly detection performance. Figures 4 and 5 provide detailed insights into these error patterns through class prediction errors and confusion matrix analysis, showing a tendency to misclassify anomalous behaviours as normal centre positions. This pattern particularly affected precision for the centre label while maintaining relatively higher recall, suggesting the model tends to identify genuine anomalous behaviour as normal. This classification bias is problematic since missing actual cheating instances carries higher costs than flagging additional clips for review, and likely stems from the imbalanced dataset containing twice as many centre-labeled instances as directed anomalous instances. Despite these challenges, Figure 6 presents the ROC curve analysis that demonstrates robust overall performance across all movement categories. The curves rise steeply toward the upper-left corner, indicating high sensitivity with low false positive rates, and remain consistently above the diagonal line representing random chance. Most curves cluster near the upper boundary, suggesting consistent performance across different anomaly types. This analysis confirms that the SVM classifier maintains effective discrimination capability and validates both our feature space design and dataset quality, even in the presence of class imbalance.4.2 Efficiency of the Anomaly Detection Classifier Table 1. Comparison of Inference Times for SVM and ITracker Models on 7-Second Videos ITracker(s) SVM(s) 66.23 6.13 54.54 5.51 54.29 6.12 To evaluate our approach's efficiency, we compared our SVM-based anomaly detection classifier with ITracker (Sun et al., 2019), a prominent eye-tracking model, by testing inference speed on three 7-second videos (Table 1). Our SVM model achieved nearly ten-fold speedup compared to ITracker, attributable to our simplified approach that reduces head pose and eye gaze estimation to less granular labels. This substantial speed improvement makes our model viable for online proctoring contexts where detecting anomalous behaviors takes precedence over precise eye tracking. Furthermore, our model enhances the detection process by incorporating face movement analysis, contributing to more comprehensive identification of abnormal behaviors during examinations. The model's simplicity and efficacy align with the requirements of large-scale online examination settings where rapid processing of multiple video feeds is essential. 4.3 Face Masking Fig. 7. User Preferences for Privacy Protection: Comparative Analysis of Facial Masking Techniques Fig. 8. Perceived Efficacy of Masking Techniques in Cheating Detection: User Confidence Assessment To evaluate our privacy-preserving proctoring framework, we conducted a user study comparing three conditions: unmasked monitoring, plain white masking, and our proposed feature-reconstructed masking. Participants assessed each approach based on two criteria: their comfort level being monitored and the perceived effectiveness for identifying misconduct. To elicit authentic responses, participants viewed themselves under each condition rather than observing others. When evaluating comfort with monitoring(Figure 7), participants unexpectedly favoured the white mask, which reveals only the eyes while concealing other facial features. The featured mask, which reconstructs facial features artificially, was consistently ranked second. No masking received the lowest preference, indicating participants' desire for some level of privacy protection. The preference for the white mask over the featured mask suggests potential discomfort with having facial features artificially reconstructed. Regarding the ability to detect academic misconduct(Figure 8), participants expressed the highest confidence in unmasked monitoring. The featured mask was predominantly ranked second, while the white mask was considered the least effective for misconduct detection. This pattern aligns with the intuitive understanding that more visible facial information enables better assessment of behaviour. The consistent second-place ranking of the featured mask in both evaluations positions it as a balanced compromise that addresses both privacy concerns and monitoring effectiveness in online proctoring environments. 5. Discussion PRIVITOR overcomes critical limitations in existing proctoring systems by providing enhanced privacy protection and practical deployment capabilities. For instance, prior systems by Yaqub et al. use either broad pose changes detected through image hashing (2022) or off-screen glances identified by gaze tracking (2023). In contrast, our system delivers precise behavioral analysis by classifying eight directional head and eye movements using facial landmarks. Although Atoum et al. (2017) achieve high accuracy through multimodal integration, their approach requires extensive hardware setup that limits scalability, whereas PRIVITOR operates with a single camera while maintaining comparable performance. Our comprehensive privacy controls mask both facial and eye regions, addressing biometric concerns present in methods like Yaqub et al. (2022, 2023), whose masking preserves the eye area. However, the framework faces limitations including a small sample size for facial masking evaluation and inability to detect complex temporal patterns that advanced techniques such as LSTM time-series analysis might identify. Future work should focus on expanding the dataset to include participants from diverse academic disciplines and cultural backgrounds, ensuring fair performance across different student populations. Additionally, testing under varied environmental conditions—including different lighting, complex backgrounds, and camera qualities—will be essential. These improvements will ensure robust performance in real-world remote examination environment. 6. Conclusion This study introduces an integrated online proctoring framework that balances student privacy with exam integrity through three major contributions: a novel web-based approach for collecting proctoring-specific training data, an efficient anomaly classifier with cheating detection algorithm that processes videos ten times faster than existing solutions with 85.40% accuracy, and a privacy-preserving facial masking technique that maintains behavioural cues while obscuring identifying features. The framework offers significant benefits for software engineering education by enabling students to take coding exams from home using their own computers, eliminating the need for expensive on-campus computer labs and proving especially cost-effective for large student populations. Future work should focus on expanding the dataset with diverse participants to avoid overfitting, exploring advanced machine learning techniques to improve detection accuracy, and integrating multiple data sources for more robust cheating detection.