Full text
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 10 OCTOBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 81 ANALYSIS OF EXISTING METHODOLOGIES FOR ASSESSING STUDENTS' KNOWLEDGE AND SKILLS USING ARTIFICIAL INTELLIGENCE ALGORITHMS O.O. Jamolov Namangan State University, Uzbekistan https://doi.org/10.5281/zenodo.17445637 Abstract. This study analyzes contemporary methodologies of Automated Essay Scoring (AES) systems based on artificial intelligence (AI) algorithms and their pedagogical foundations. The research examines the limitations of traditional assessment approaches, including subjectivity, excessive time consumption, and insufficient individualization. The effectiveness of Natural Language Processing (NLP), machine learning, and deep learning technologies integrated into AI-based systems is evaluated. A comparative analysis of the architectural and algorithmic characteristics of e-rater, Criterion, WriteToLearn, and Turnitin Feedback Studio systems is conducted. The capabilities and constraints of these systems are assessed through the lens of psychometric reliability, validity, and fairness principles. Theoretical frameworks and practical recommendations for implementing AI-based assessment in Uzbekistan's higher education system are developed. Results demonstrate that AI systems achieve correlation coefficients of 0.88–0.92 with human raters and reduce assessment time by 40–60 fold. Keywords: artificial intelligence, automated essay scoring, natural language processing, machine learning, pedagogical assessment, psychometrics, validity, reliability, higher education, educational technology. INTRODUCTION In contemporary higher education systems, the assessment of students' knowledge and competencies constitutes a critical component in ensuring educational quality and effectiveness. However, traditional assessment methods encounter several significant challenges: increasing workload for educators, the presence of subjectivity in evaluation, difficulties in providing timely and constructive feedback, and limited capacity to accommodate individual learner characteristics [7]. According to the Global Education Monitoring Report, educators allocate 40-50% of their working time to grading student assignments, which negatively impacts the core pedagogical function of teaching [15]. Concurrently, research has demonstrated that subjective errors attributed to human factors in assessment may reach approximately 15-25% [13]. Artificial intelligence technologies offer substantial potential for addressing these challenges. Although Automated Essay Scoring (AES) systems have been developing since the 1960s, recent advancements in Natural Language Processing (NLP), Machine Learning (ML), and Deep Neural Networks (DNN) have elevated this field to a new level over the past decade [11]. LITERATURE ANALYSIS AND METHODOLOGY The theory of pedagogical assessment evolved along several key directions during the twentieth century. Classical Test Theory (CTT), Item Response Theory (IRT), and contemporary psychometrics established the scientific foundations for evaluation practices [4]. Pedagogical assessment encompasses four primary functions: the diagnostic function, which identifies students'
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 10 OCTOBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 82 initial knowledge levels; the formative function, which provides continuous monitoring and feedback throughout the learning process; the summative function, which evaluates final outcomes and certifies achievement; and the motivational function, which encourages and directs learning efforts. Contemporary pedagogy places particular emphasis on formative assessment. Research by Black and Wiliam (1998) demonstrates that effective formative assessment can enhance students' learning outcomes by 0.4-0.7 standard deviations, representing a substantial effect size [2]. Limitations of Traditional Assessment Methods. Contemporary research identifies several critical issues with traditional assessment approaches. Regarding subjectivity and reliability concerns, different instructors may assign varying scores to identical work. A study by Hamp-Lyons (1995) revealed that correlation coefficients between two independent raters typically range from 0.6 to 0.75, indicating substantial inconsistency [6]. In terms of time expenditure, an individual instructor can evaluate approximately 20-30 essays per day. For a cohort of 100-150 students, this requires 3-5 full working days [14]. Students typically wait 1-2 weeks after submission before receiving feedback, creating a delayed response problem. Pedagogical research demonstrates that feedback effectiveness correlates directly with promptness of delivery [8]. The challenge of insufficient individualization manifests in the difficulty of identifying each student's specific errors and providing personalized recommendations within large groups. DISCUSSION AND RESULTS Figure 1. Taxonomic Classification of Artificial Intelligence Assessment Methodologies Taxonomy of Contemporary AI Algorithms. Current AES systems employ the following primary methodological approaches [9]: A) Feature Engineering (Hand-crafted Feature Extraction). In this approach, developers predetermine the set of features to be extracted from text. These include: lexical indicators—word count, average length, Type-Token Ratio; syntactic markers—sentence structure, grammatical
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 10 OCTOBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 83 accuracy levels, part-of-speech distribution; semantic attributes—topical relevance and content precision; discourse characteristics—representing overall text organization, argument sequencing, and coherence. B) Machine Learning Models. Classical algorithms comprise Linear Regression, Support Vector Machines, Random Forest, and Gradient Boosting implementations (XGBoost, LightGBM). Deep Learning encompasses Recurrent Neural Networks (RNN, LSTM, GRU), Convolutional Neural Networks adapted for text processing, and Attention mechanisms. Transformer architectures include BERT (Bidirectional Encoder Representations from Transformers), GPT (Generative Pre-trained Transformer), RoBERTa, ALBERT, and T5 [5]. C) Hybrid Approaches. Contemporary systems frequently integrate multiple methodologies: Feature-based combined with Deep Learning, Multi-model ensembles, and Rulebased approaches merged with ML components. Figure 1 depicts the complete taxonomic framework of AI-driven assessment methodologies. The classification encompasses three principal approaches: Feature-based methodologies demonstrating accuracy ranges of 0.75-0.82, Deep Learning techniques achieving accuracy levels of 0.85-0.92, and Hybrid approaches attaining the highest accuracy of 0.88-0.92. Table 1. Characteristics and Performance of Major AES Systems System Developer Year Technology QWK Application Domain E-rater Educational Testing Service 1999 Feature-based + ML 0.87 GRE, TOEFL Criterion Educational Testing Service 2003 Hybrid Architecture 0.850.90 K-12, Higher Education WriteToLearn Pearson 2000 Latent Semantic Analysis 0.780.83 K-12 Education Turnitin Feedback Studio Turnitin LLC 2020 Hybrid + NLP 0.820.88 Global Education Note: QWK = Quadratic Weighted Kappa (correlation with human raters) Comparative Analysis of Prominent AES Systems E-rater (Educational Testing Service) Developed in 1999, the e-rater system is employed in standardized assessments including GRE, TOEFL, and GMAT. The system analyzes over 12 linguistic features: grammar and mechanics verification, lexical complexity evaluation, topical relevance assessment, discourse structure analysis, and organizational development [1]. According to research by Attali and Burstein (2006), the correlation coefficient between e-rater and human raters reaches 0.87. Criterion (ETS). Criterion represents a comprehensive instructional platform built upon the e-rater foundation, providing not only scoring but also detailed formative feedback [1]. Key features include real-time scoring, grammatical error detection with explanations, developmental recommendations, plagiarism detection, and instructor monitoring dashboards. Performance metrics demonstrate scoring speed of less than 1 second, teacher-system agreement of 0.85-0.90 (QWK), and student satisfaction rates of 78-82% [14]. WriteToLearn (Pearson). Pearson's system utilizes Intelligent Essay Assessor (IEA) technology based on Latent Semantic Analysis (LSA) methodology [10]. LSA analyzes text
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 10 OCTOBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 84 semantics in vector space, enabling comparison with exemplar essays, separate evaluation of content and structure, and provision of formative feedback. Turnitin Feedback Studio. Originally recognized primarily for plagiarism detection, Turnitin evolved into a comprehensive assessment system from 2020 [12]. Components include Similarity Report (plagiarism verification), Grademark (scoring functionality), Originality Check (AI-generated content detection), and Feedback Studio (grammar and style analysis). Distinctive features encompass support for over 70 languages, Learning Management System integration (Moodle, Canvas), and substantial instructor control [17]. Architectural Framework of AI Assessment Systems Figure 2. General Architecture of AI-Based Assessment System Figure 2 illustrates the comprehensive architecture of contemporary AI-driven assessment systems. The system comprises the following core components: Input Layer: The system receives two primary inputs—the student essay (text requiring evaluation) and assessment rubrics (criteria and standards). Preprocessing Stage: Text undergoes sequential processing steps—tokenization (segmentation into words), cleaning and normalization (removal of extraneous characters), and Part-of-Speech (POS) tagging (identification of word categories). Natural Language Processing: This critical phase encompasses five fundamental analysis types—lexical analysis (vocabulary richness), syntactic analysis (sentence structure), semantic analysis (topical relevance), discourse analysis (text organization), and pragmatic analysis (argument strength). Machine Learning Layer: Following NLP analysis, two parallel pathways emerge: Feature Engineering (extraction of 150+ features, vectorization) and Deep Learning (pre-trained models such as BERT/GPT, transfer learning, fine-tuning). Model Ensemble: integration of multiple models through voting or averaging mechanisms to generate final predictions. Output and Feedback: The system generates two primary outputs—scoring (rubric-based and overall scores) and detailed feedback (error identification, recommendations, improvement pathways). Teacher Moderation: AI-generated results undergo instructor review—verification of AI outputs, modifications when necessary, supplementary annotations, particularly examination of student appeals, and final approval.
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 10 OCTOBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 85 Continuous Improvement: The system undergoes ongoing refinement—feedback data collection, model retraining, accuracy enhancement, and parameter optimization. Psychometric Evaluation: Validity and Reliability. The following validity types are critical in assessing AES systems [16]: Construct Validity: Does the system genuinely measure the competencies it intends to assess? Research by Williamson (2009) demonstrates that numerous AES systems evaluate surface-level features but fail to comprehensively assess deep reasoning and creative writing. Criterion Validity: To what extent do AES scores align with professional rater scores? Measurements include Pearson correlation (r), Quadratic Weighted Kappa (QWK), and Exact plus Adjacent Agreement. Standard benchmarks: QWK > 0.75 indicates satisfactory performance, QWK > 0.85 represents good performance, and QWK > 0.90 denotes excellent performance. Consequential Validity: What educational impacts result from AES implementation? Positive outcomes include time savings for educators, increased practice opportunities for students, and prompt feedback delivery. Potential risks encompass teaching to the algorithm, diminished creativity, and excessive reliance on technology. Reliability: Inter-rater Reliability: Traditional assessment (two human raters)—r = 0.700.85; AES (system self-consistency)—r = 0.95-0.99; AES versus human—r = 0.80-0.92. Test-retest Reliability: AES systems should consistently evaluate identical texts uniformly. Contemporary systems achieve 99.9% consistency in this dimension. Fairness and Bias: AES systems may occasionally demonstrate bias toward specific groups: linguistic factors (non-native English speakers receiving lower scores) [3], gender-based variations (stylistic differences between male and female writing), and ethnic considerations (cultural context and discourse pattern differences). Modern AES development incorporates Differential Item Functioning (DIF) testing protocols, which facilitate fairness assurance. Table 2. Psychometric Evaluation Results of AES Systems Criterion Human-Human E-rater BERT-based Standard Validity (QWK) 0.75-0.85 0.82-0.87 0.85-0.90 >0.75 Reliability 0.70-0.85 0.95-0.99 0.96-0.99 >0.80 Gender DIF Neutral ±0.05 ±0.03 <±0.10 Accuracy (%) 35-45 40-50 48-58 >50 Note: QWK = Quadratic Weighted Kappa; DIF = Differential Item Functioning Comparative Analysis: Traditional versus AI-Based Assessment. Figure 3 presents a comprehensive comparison of traditional and AI-based assessment methodologies. Key differentiating factors include: time expenditure (traditional: 10-15 minutes per essay; AI: <1 second), reliability measured by QWK (traditional: 0.70-0.85; AI: 0.80-0.92), feedback delivery speed (traditional: 7-14 days; AI: immediate), consistency (traditional: 70%; AI: 99.9%), and scalability (traditional: 20-30 essays per day; AI: 10,000+ essays per day). Table 3. Comparative Performance Analysis: Traditional, AI-Based, and Hybrid Assessment Methodologies Performance Metric Traditional Assessment AI-Based Assessment Hybrid Approach Processing Time (per essay) 10-15 minutes <1 second 2-3 minutes Inter-rater Reliability (QWK) 0.70-0.85 0.82-0.92 0.88-0.94 Feedback Turnaround Time 3-14 days Immediate 1-2 days
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 10 OCTOBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 86 Scoring Consistency Rate 60-75% 99%+ 90-95% Note: QWK = Quadratic Weighted Kappa Advantages of AI-Based Assessment: Speed and scalability (thousands of essays processed within minutes), consistency (uniform criteria application), cost-effectiveness (reduced instructor labor costs), immediate feedback delivery (students receive results promptly), objectivity (absence of personal bias), data analytics capability (simplified monitoring of student progress), and continuous availability (24/7 service provision). Limitations and Challenges: Inability to evaluate creativity (difficulty comprehending unconventional approaches), constrained contextual understanding (struggles with irony, metaphor, and cultural references), incomplete assessment of deep reasoning (argument strength and originality), gaming the system phenomenon (students may learn to write for algorithmic preferences), technical errors (occasional misgrading instances), data security concerns (risks associated with storing personal information), and diminished teacherstudent relationship (loss of human interaction). Hybrid Model: AI + Human. The contemporary approach integrates AI and human evaluators through a three-stage process: Stage 1—AI conducts initial assessment and provides feedback; Stage 2—instructors review only borderline cases (5-10%); Stage 3—dual verification for critical final evaluations. This methodology conserves 60-70% of instructor time, maintains high quality standards, and optimizes the student experience. Recommendations for Uzbekistan's Higher Education. Uzbekistan's higher education landscape comprises 212 institutions, 1,553,552+ students (2025), 45,000+ faculty members, with an average student-faculty ratio of 1:34.4. Primary requirements include standardizing assessment processes, reducing instructor workload, ensuring prompt student feedback, and developing digital educational infrastructure [21]. Table 4. Implementation Phases and Resource Requirements for Uzbekistan Phase Timeline Primary Objectives Preparation 6-12 months Uzbek NLP development, corpus collection, pilot design Pilot Testing 12-18 months 2-3 universities, 500-1000 students Expansion 18-36 months 5-10 universities, 5,000-10,000 students National Scale 36+ months All higher education institutions, standardization Phased Implementation Strategy: Phase 1 (Preparation, 6-12 months): Development of Uzbek language-adapted NLP tools, corpus compilation (10,000+ essays), rubric design, selection of 2-3 pilot universities, instructor training program implementation. Phase 2 (Pilot Testing, 12-18 months): Deployment in 2-3 universities, application across 3-5 academic disciplines, experimentation with 500-1,000 students, psychometric system evaluation, collection of instructor and student feedback. Phase 3 (Expansion, 18-36 months): Distribution to 5-10 universities, expansion to 10-15 disciplines, engagement of 5,000-10,000 students, system optimization, Learning Management System integration (Hemis, Moodle). Phase 4 (National Scale, 36+ months): System-wide implementation across all higher education institutions, development of national standards, continuous monitoring and updating, regional collaboration (CIS, Asia), international certification. Technical Constraints and Solutions for Uzbek Language: NLP challenges for Uzbek include: morphological richness (agglutinative language structure, single words generating 100+ forms), corpus insufficiency (limited annotated texts), absence of NLP tools (lacking or inadequate
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 10 OCTOBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 87 POS taggers, Named Entity Recognition, parsers), and unavailability of pre-trained models (no BERT model specifically developed for Uzbek language). Table 5. Natural Language Processing Challenges and Proposed Solutions for Uzbek Language Implementation Technical Challenge Problem Description Proposed Solution Technological Approach Morphological Richness Agglutinative language structure, generating 100+ word forms Morphological Analysis Development UzMorphAnalyzer Framework Limited Corpus Availability Fewer than 10,000 annotated essay samples Systematic Data Collection Initiative Inter-university Collaborative Effort Tokenization Complexity Language-specific Grammatical Rules Uzbek-adapted Tokenization System SentencePiece, BytePair Encoding (BPE) Part-of-Speech Tagging Accuracy Current accuracy approximately 85% Advanced Tagger Development Bidirectional LSTMCRF Architecture Word Embedding Quality Absence of High-quality Pre-trained Models Training on Largescale Corpus FastText, Multilingual BERT (mBERT) Recommendation: Initial utilization of multilingual models (mBERT, XLM-RoBERTa), followed by fine-tuning for Uzbek language adaptation. Technical Requirements and Infrastructure: Software Components: Python 3.8+, TensorFlow/PyTorch frameworks, Hugging Face Transformers library, Uzbek language-specific NLP libraries, Learning Management System integration (Hemis, Moodle). Hardware Resources: Cloud computing platforms (AWS, Azure, local data center), GPU servers (NVIDIA A100/V100), database management systems (PostgreSQL, MongoDB), security infrastructure. Expert Personnel: AI/ML specialists (5-10 individuals), NLP experts (3-5 individuals), pedagogical designers (5-7 individuals), technical support staff (3-5 individuals). Ethical and Legal Considerations: Data Security: Personal information protection, anonymous data processing, consent acquisition, compliance with GDPR and Uzbekistan legislation. Transparency: System operation explanation, scoring rationale disclosure, personalized dashboards for instructors and students, explainable AI principles implementation. Fairness: Bias testing across demographic groups, continuous system monitoring, appeals mechanism for contested scores, Differential Item Functioning (DIF) analysis. Human Oversight: Dual verification for critical assessments, AI positioned as assistive tool (not replacement), human-in-the-loop principle adherence. Anticipated Outcomes and Benefits: Quantitative Indicators: 60-70% reduction in instructor time allocation, 80-90% decrease in assessment costs, 3-5 fold increase in student practice opportunities, 20-30 fold acceleration in feedback delivery. Quality Enhancement: Assessment consistency improvement from 70% to 95%, subjectivity reduction from 15-25% to 5-8%, student satisfaction increase from 65% to 80%+, 1015% improvement in academic performance.
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 10 OCTOBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 88 Institutional Impact: International accreditation opportunities, accelerated digital transformation, educational quality monitoring systems, data-driven decision-making capabilities. Summary of Research Findings. This analysis demonstrates that AI-based assessment systems constitute pedagogically sound, technologically mature, and economically efficient solutions. Contemporary AES systems have achieved the following milestones: 1. High Accuracy: Leading systems (BERT-based, GPT architectures) achieve correlation coefficients of 0.88-0.92 with human raters. 2. Consistency: Greater than 99% reliability, eliminating subjective errors. 3. Efficiency: Assessment speed increased 600-900 fold compared to traditional methods. 4. Scalability: Capability to evaluate thousands of students simultaneously. 5. Economic Benefits: Cost reduction of 80-95%. Key Challenges: Technological Limitations: Difficulty assessing creative and unconventional approaches, incomplete measurement of deep reasoning and originality, constrained comprehension of contextual and cultural nuances. Pedagogical Concerns: Risk of teaching to the algorithm phenomenon, diminished teacherstudent human interaction, potential decline in creativity and critical thinking development. Technical Constraints: Insufficient NLP tools for Uzbek language, corpus and data scarcity, elevated infrastructure costs. Necessity of Hybrid Approach: Analysis results indicate that the optimal solution constitutes a hybrid model. AI Role: Initial automated assessment (90-95% of cases), prompt feedback and recommendations, scoring based on objective metrics. Human (Instructor) Role: Review of complex and borderline cases (5-10%), evaluation of creative and deep thinking, final approval and pedagogical oversight, personalized student interaction. This approach balances technological efficiency with humanistic values, integrates effectiveness with quality, and fully aligns with pedagogical principles. International Experience and Uzbekistan Context: International practice demonstrates successful implementations: United States (e-rater successfully deployed in GRE/TOEFL for 20+ years), China (iTest.cn serves 100+ million students), South Korea (WriteToLearn widely adopted in K-12 education). Uzbekistan-Specific Considerations: Necessity for Uzbek language-specific NLP development, accommodation of cultural context and discourse traditions, phased implementation approach (avoiding abrupt transitions), instructor training and capacity building. CONCLUSION Automated assessment systems founded upon artificial intelligence algorithms possess substantial potential for generating transformative changes in contemporary education. Through the utilization of natural language processing, machine learning, and deep neural network technologies, these systems deliver high precision (QWK > 0.85), instantaneous feedback, and scalability. The present research has arrived at the following principal findings: 1. Technological capabilities: Contemporary AI systems have approximated human evaluator performance levels (correlation coefficients ranging from 0.88 to 0.92), achieved assessment speed enhancements of 600-900 fold, and realized cost reductions of 80-95%. 2. Methodological classification: Feature-based approaches demonstrate rapid processing and interpretability, albeit with certain constraints (QWK: 0.75-0.82). Deep Learning architectures
SCIENCE AND INNOVATION INTERNATIONAL SCIENTIFIC JOURNAL VOLUME 4 ISSUE 10 OCTOBER 2025 ISSN: 2181-3337 | SCIENTISTS.UZ 89 attain superior accuracy while exhibiting "black box" characteristics (QWK: 0.85-0.92). Hybrid models represent the most effective solution for professional systems (QWK: 0.88-0.92). 3. Psychometric quality: Validity and reliability measures achieve elevated standards, while fairness challenges persist (though remediation efforts are underway), necessitating continuous monitoring and refinement. 4. Prospects for Uzbekistan: A phased implementation strategy (comprising 4 stages over 3-4 years) is recommended, with essential development of NLP tools adapted to the Uzbek language. 5. Pedagogical considerations: AI assessment functions as a supplementary rather than replacement instrument, teacher oversight remains critical, integration of formative and summative assessment methodologies is necessary, and students require prompt and constructive feedback. Nevertheless, AI assessment systems should be conceptualized as complementary instruments to traditional approaches rather than complete substitutes. The human evaluator's role maintains paramount significance in assessing complex competencies such as creative thinking, profound analysis, and contextual comprehension. The implementation of AI-based assessment methodologies for Uzbekistan's higher education generates the following opportunities: 1. Enhanced efficiency: Conservation of 60-70% of instructors' time allocation 2. Quality improvement: Consistency and objectivity in evaluation processes 3. Enriched educational experience: Rapid and targeted feedback for students 4. Data-informed decision-making: Comprehensive monitoring of student progression 5. International integration: Educational system alignment with global standards Successful implementation necessitates a graduated approach, technologies adapted to the Uzbek language, instructor preparation, and adherence to ethical principles. In final analysis, it merits emphasis that AI assessment systems represent not merely technological innovation, but rather a pedagogical paradigm shift. When appropriately deployed, these systems enable education to become increasingly personalized, efficient, and equitable. For Uzbekistan's higher education system, this constitutes a significant milestone in digital transformation and the attainment of international standards. REFERENCES 1. Attali Y., Burstein J. Automated essay scoring with e-rater V.2.0. The Journal of Technology, Learning and Assessment, 2006, Vol. 4(3), pp. 1-30. 2. Black P., Wiliam D. Assessment and classroom learning. Assessment in Education: Principles, Policy & Practice, 1998, Vol. 5(1), pp. 7-74. 3. Bridgeman B., Trapani C., Attali Y. Comparison of human and machine scoring of essays: Differences by gender, ethnicity, and country. Applied Measurement in Education, 2012, Vol. 25(1), pp. 27-40. 4. Crocker L., Algina J. Introduction to Classical and Modern Test Theory. Mason, OH: Cengage Learning, 2006. 527 p. 5. Devlin J., Chang M., Lee K., Toutanova K. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. Proceedings of NAACL-HLT, 2019, pp. 41714186. 6. Hamp-Lyons L. Rating nonnative writing: The trouble with holistic scoring. TESOL Quarterly, 1995, Vol. 29(4), pp. 759-762.