Machine Learning (AutoML)-Driven Wheat Yield Prediction for European Varieties: Enhanced Accuracy Using Multispectral UAV Data
Abstract
Research article entitled "Machine Learning (AutoML)-Driven Wheat Yield Prediction for European Varieties: Enhanced Accuracy Using Multispectral UAV Data” was published by Agriculture (MDPI)
Full text
Academic Editors: Francesco Marinello, Chen Zhang and Haoteng Zhao Received: 5 May 2025 Revised: 26 June 2025 Accepted: 11 July 2025 Published: 16 July 2025 Citation: Kešelj, K.; Stamenkovi´c, Z.; Kosti´c, M.; A´cin, V.; Teki´c, D.; Novakovi´c, T.; Ivaniševi´c, M.; Ivezi´c, A.; Magazin, N. Machine Learning (AutoML)-Driven Wheat Yield Prediction for European Varieties: Enhanced Accuracy Using Multispectral UAV Data. Agriculture 2025,15, 1534. https://doi.org/ 10.3390/agriculture15141534 Copyright: © 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license (https://creativecommons.org/ licenses/by/4.0/). Article Machine Learning (AutoML)-Driven Wheat Yield Prediction for European Varieties: Enhanced Accuracy Using Multispectral UAV Data Krstan Kešelj 1, Zoran Stamenkovi´c 1,* , Marko Kosti´c 1, Vladimir A´cin 2, Dragana Teki´c 1, Tihomir Novakovi´c 1, Mladen Ivaniševi´c 1, Aleksandar Ivezi´c 3and Nenad Magazin 1 1Faculty of Agriculture, University of Novi Sad, Trg Dositeja Obradovi´ca 8, 21000 Novi Sad, Serbia; [email protected] (K.K.); [email protected] (M.K.); [email protected] (D.T.); tihomir[email protected] (T.N.); [email protected] (M.I.); [email protected] (N.M.) 2Institute of Field and Vegetable Crops, Maksima Gorkog 30, 21000 Novi Sad, Serbia; vladimir[email protected] 3Center for Biosystems, BioSense Institute, University of Novi Sad, Dr. Zorana Ðin ¯ di´ca 1, 21000 Novi Sad, Serbia; aleksandar[email protected] *Correspondence: [email protected] Abstract Accurate and timely wheat yield prediction is valuable globally for enhancing agricultural planning, optimizing resource use, and supporting trade strategies. Study addresses the need for precision in yield estimation by applying machine-learning (ML) regression models to high-resolution Unmanned Aerial Vehicle (UAV) multispectral (MS) and Red-Green-Blue (RGB) imagery. Research analyzes five European wheat cultivars across 400 experimental plots created by combining 20 nitrogen, phosphorus, and potassium (NPK) fertilizer treatments. Yield variations from 1.41 to 6.42 t/ha strengthen model robustness with diverse data. The ML approach is automated using PyCaret, which optimized and evaluated 25 regression models based on 65 vegetation indices and yield data, resulting in 66 feature variables across 400 observations. The dataset, split into training (70%) and testing sets (30%), was used to predict yields at three growth stages: 9 May, 20 May, and 6 June 2022. Key models achieved high accuracy, with the Support Vector Regression (SVR) model reaching R 2 = 0.95 on 9 May and R 2 = 0.91 on 6 June, and the Multi-Layer Perceptron (MLP) Regressor attaining R 2 = 0.94 on 20 May. The findings underscore the effectiveness of precisely measured MS indices and a rigorous experimental approach in achieving highaccuracy yield predictions. This study demonstrates how a precise experimental setup, large-scale field data, and AutoML can harness UAV and machine learning’s potential to enhance wheat yield predictions. The main limitations of this study lie in its focus on experimental fields under specific conditions; future research could explore adaptability to diverse environments and wheat varieties for broader applicability. Keywords: yield prediction; wheat; machine learning; multispectral indices 1. Introduction Early prediction of wheat yield is crucial not only for estimating final production quantities but also for enabling farmers to optimize crop management practices through the timely application of agronomic measures aimed at maximizing yield. Agriculture 2025,15, 1534 https://doi.org/10.3390/agriculture15141534
Agriculture 2025,15, 1534 2 of 32 In this paper, we have studied the application of machine-learning methodologies, particularly using the PyCaret library, in processing data acquired from UAVs equipped with both MS and RGB cameras. The goal was to develop highly accurate machine-learning models for the estimation of wheat yield based on drone images of the plant canopy, emphasizing achieving high precision and a significant coefficient of determination. While there are numerous studies on wheat yield prediction, many suffer from inadequate precision and lower determination coefficients, making them less reliable for practical applications. This work aims to address these issues by employing advanced machine-learning techniques and comprehensive UAV-based data collection on five European wheat cultivars, which, in terms of varietal characteristics, are similar to wheat varieties grown across Europe. This study’s novelty lies in setting a new benchmark for high predictive accuracy in European wheat yield estimation models by using machine-learning methodologies and UAV imagery. Traditional crop yield prediction methods—often based on field surveys, statistical modeling, or satellite imagery—frequently lack the spatial resolution, timeliness, and adaptability required for modern precision agriculture. These approaches are typically labor-intensive, delayed in response, or unable to capture subtle within-field variability. Hence, there is a growing interest in integrating UAV technology and machine learning to overcome these limitations and deliver more accurate, site-specific, and timely yield predictions. Machine learning (ML) is a field of artificial intelligence that focuses on developing algorithms enabling computers to recognize patterns and make data-driven predictions or decisions without being explicitly programmed for each task. In agriculture, ML models, particularly regression models, are widely used to predict outcomes such as crop yield by analyzing complex datasets [ 1 , 2 ]. These models identify relationships between variables (e.g., soil properties, environmental factors, or plant characteristics) and target outcomes, allowing for more accurate and timely predictions. A recent advancement in ML is Automated Machine Learning (AutoML), which automates the process of selecting, tuning, and optimizing ML models. AutoML enables the building of effective predictive models by simplifying the model selection and hyperparameter tuning process [ 3 ]. In our study, we used AutoML (PyCaret library) to identify the most suitable ML models for predicting wheat yield from UAV-collected MS and RGB imagery, enhancing the speed, accuracy, and reproducibility of the modeling process. The integration of AutoML in yield estimation offers a powerful approach to maximize the utility of agricultural data, optimize resource allocation, and improve decision-making in precision agriculture. Unmanned aerial vehicles (UAVs) have experienced a surge in use for capturing aerial images across numerous sectors. Their application has notably advanced in agriculture due to recent technological developments [ 4 ]. These advancements have enabled the integration of complex technology such as MS and RGB cameras into UAVs, thus solidifying their role as remote sensing systems in precision agriculture [5]. Their principal function is to capture images at multiple wavelength bands beyond the visible spectrum. Among the most employed indices are the Normalized Difference Vegetation Index (NDVI), Normalized Difference Red Edge (NDRE), Green Normalized Difference Vegetation Index (GNDVI), Leaf Chlorophyll Index (LCI), and Optimized Soil Adjusted Vegetation Index (OSAVI) [6–8]. Beyond these MS indices, the role of RGB indices such as Green–Red Ratio Index (GRRI), Green–Blue Ratio Index (GBRI), Red–Blue Ratio Index (RBRI), Excess Green (ExG), Colour Index of Vegetation (CIVE), and Vegetation Index Green (VIg) should not be underestimated [ 3 ]. While MS indices provide insights into specific aspects of crop health not
Agriculture 2025,15, 1534 3 of 32 visible to the naked eye, RGB indices can provide useful insights that are more visually intuitive. They can provide data about the general health and vigor of the crop based on the visual color, and they can help detect issues such as wilting or visible pests. Moreover, contemporary literature emphasizes the application of high-resolution hyperspectral and RGB cameras for estimating above-ground biomass (AGB), which is a significant phenotypic index for evaluating photosynthesis capacity, healthy growth, and estimating crop yield. Further insights into this topic can be found in the works of Yang Liu et al. [9,10]. Despite promising results in previous studies, many suffer from limited scale, use of low-resolution satellite imagery, or narrow crop diversity. For example, some studies focus on sugarcane or corn and use datasets restricted to 48–80 small plots, reducing generalizability. In contrast, this study leverages UAV-acquired high-resolution data across 400 plots and five wheat cultivars, reflecting the heterogeneity of European production systems. MS and RGB cameras, thanks to their ability to capture high-resolution images across multiple spectral bands, including the infrared, near-infrared, red, green, and blue range, provide a wealth of data. This data is invaluable for assessing crop health, identifying nutrient deficiencies, or tracking pest infestations. Therefore, UAVs equipped with these cameras have become pivotal in crop monitoring, analysis, and yield estimation, which are fundamental tasks for efficient and effective precision agriculture. A combination of remote sensing systems and machine-learning techniques is attracting particular attention today, especially in the estimation of crop yields [ 11 ]. This integrated approach holds significant promise for improving the accuracy and efficiency of crop yield prediction. Ref. [ 12 ] investigated early prediction of sugarcane crop yield using high-resolution MS imagery from UAVs and three advanced machine-learning techniques: Random Forest Regression (RFR), SVR, and Nonlinear Autoregressive Exogenous Artificial Neural Network (NARX ANN). The study focused on plot-level prediction and addressed challenges such as high ratooning capacity, limited high-resolution data, and yield complexity. Results showed that vegetation indices exhibited stronger correlations with crop yield during the middle growth stage. NARX ANN outperformed other models with an R 2 of 0.96 and the lowest Root Mean Square Error (RMSE) of 4.92 t/ha. SVR and RFR showed similar performance, with R 2 values of 0.52 and 0.48, and RMSE values of 14.85 t/ha and 11.20 t/ha, respectively. While this study demonstrates the potential of machine-learning approaches for plotlevel sugarcane yield prediction, a limitation lies in experimental scale, conducted over 48 plots, as well as the lower resolution of satellite-acquired imagery. Expanding the experimental field size and using higher-resolution imagery captured at optimal growth stages would enhance model accuracy and reliability. Another study to estimate corn grain yield using a neural network model based on MS and RGB vegetation indices, canopy cover, and plant density was conducted by [ 13 ]. The results showed high correlations between the estimated and observed corn grain yield. The variables with the highest relative importance for yield estimation varied depending on the stage of crop development. For the early stage (47 days after sowing), wide dynamic range vegetation index (WDRVI), plant density, and canopy cover exhibited the highest correlation coefficient and the smallest errors. At a later stage (79 days after sowing), a combination of NDVI, NDRE, WDRVI, ExG, triangular greenness index (TGI), plant density, and canopy cover provided the best estimation of corn grain yield. The study highlights the effectiveness of remote sensing data and machine-learning techniques for the accurate estimation of crop yield. While this study demonstrates the effectiveness of remote sensing data and machine-learning techniques for yield estimation, it is limited by its small plot sizes, with each plot containing only 15–20 plants and a total of 80 plots overall. Increasing
Agriculture 2025,15, 1534 4 of 32 the number and size of plots, along with more extensive plant sampling, could further enhance model robustness and generalizability, leading to more reliable yield predictions. However, a common limitation across these studies is the lack of diverse crop types and insufficient testing across varying fertilization or cultivar-specific responses. The current research addresses these gaps by applying a wide range of vegetation indices to a heterogeneous wheat dataset, enabling greater model adaptability and robustness across real-field scenarios. One of the most used machine-learning systems is the PyCaret software [ 3 , 14 ], which can also be used for analyzing and interpreting MS and RGB indices in the context of yield estimation. PyCaret, an open-source Python library, simplifies complex machinelearning tasks for data scientists and analysts. In the field of agriculture, PyCaret automates yield estimation from MS and RGB camera data through pattern recognition, streamlining end-to-end experiments. Similar research using the PyCaret machine-learning algorithm models to estimate crop health was conducted by [ 3 ]. In the study, UAVs equipped with MS and RGB cameras were used to capture images of spring wheat (Triticum aestivum L.). Their objective was to remotely estimate the maximum quantum efficiency of the photosystem (F v /F m ), a key indicator of photosynthetic health in plants. They constructed 51 vegetation indices and used 26 different machine-learning algorithms in PyCaret to analyze these indices. Their results showed strong correlations between most of the MS and half of the RGB vegetation indices with Fv/Fm. After comparing the performance of the algorithms, the Automatic Relevance Determination (ARD) model provided the highest accuracy in estimating F v /F m when using a combination of RGB and MS vegetation indices. The research underscores the potential of using UAVs and machine learning for rapid and precise monitoring of photosynthetic health in crops, providing critical technical support for agriculture. Building on this potential, the present study introduces an approach to achieving high prediction accuracy in wheat yield modeling by combining UAV-based MS and RGB data with AutoML. By applying an AutoML tool, 25 machine-learning models were optimized to assess the potential of 65 vegetation indices for yield estimation. This research stands out for its large-scale experimental design, which involved diverse treatments on multiple wheat cultivars across 400 plots, representing conditions typical of European wheat varieties. The methodology and comprehensive data collection offer a scalable framework for yield prediction, advancing the integration of UAV and AutoML technologies in precision agriculture. In summary, this study builds upon the existing body of research while addressing key limitations related to scale, data resolution, crop diversity, and experimental complexity. It introduces a robust and scalable AutoML-based framework that integrates MS and RGB imagery across diverse genotypes and treatment conditions, aiming to push the boundaries of predictive accuracy in crop yield estimation. 2. Materials and Methods The Workflow Diagram for Wheat Yield Prediction Using UAV-Based Data and Machine-Learning Models is presented in the Appendix Aas Figure A1. 2.1. Experimental Field Design and Setup The study was part of a long-term experiment conducted by the Institute of Field and Vegetable Crops, Novi Sad, Vojvodina, Serbia. Its precise geographic coordinates are 45 ◦ 19 ′ 58.0 ′′ N, 19 ◦ 49 ′ 53.6 ′′ E, and it stands at an altitude of 82 m (Figure 1). The site was partitioned into 400 sub-plots, each spanning 5 × 10 m, methodically organized into a
Agriculture 2025,15, 1534 5 of 32 structured grid of 20 ×20. The soil at the site is the dominating soil type in the Vojvodina region, Haplic Chernozem Aric [ 15 ], and is characterized as highly fertile (approximately 43% of the total arable land). Figure 1. Geographic location of the experimental field. Five distinct wheat cultivars were selected for the study: NS Igra, NS Rajna, NS Futura, NS Epoha, and NS Obala. The sowing density ranged between 200 and 230 kg/ha, contingent upon the specific cultivar. Critically, each cultivar was subjected to a diverse array of treatments, comprising twenty different NPK mineral fertilizer combinations. It is noteworthy that each of these treatments was replicated four times to ensure robustness in the findings. The experiment aimed to create a diverse set of conditions, similar to the natural changes seen in farming fields. This matrix incorporated four hundred sub-plots, each interspersed with one of the five wheat cultivars, and was further nuanced by twenty distinct NPK treatment variations. This design allowed for variation changes in spectral channels, closely mimicking real-world farming events. This design not only accentuates the study’s precision but also enables the replication of naturally occurring agronomic phenomena in a controlled setting. For a detailed breakdown of NPK treatment combinations and the corresponding yields across each wheat cultivar, readers are directed to Table 1. Table 1. Comprehensive overview of NPK treatment combinations and corresponding yields for each wheat cultivar. Mineral Fertilization Yield Intervals/Average Yield (t/ha) Variant N (kg/ha) P (kg/ha) K (kg/ha) NS Igra NS Rajna NS Futura NS Epoha NS Obala 10001.63–2.93 (2.3) 1.41–2.85 (2.26) 1.93–2.76 (2.29) 1.90–2.63 (2.33) 2.02–4.51 (2.97) 2 100 0 0 3.53–5.46 (4.73) 3.11–5.19 (4.53) 3.27–5.00 (4.32) 3.67–5.20 (4.59) 4.34–5.67 (5.15)
Agriculture 2025,15, 1534 6 of 32 Table 1. Cont. Mineral Fertilization Yield Intervals/Average Yield (t/ha) Variant N (kg/ha) P (kg/ha) K (kg/ha) NS Igra NS Rajna NS Futura NS Epoha NS Obala 3 0 100 0 1.92–3.25 (2.53) 1.82–2.88 (2.46) 2.25–2.97 (2.60) 2.71–3.15 (2.96) 3.04–4.42 (3.63) 4 0 0 100 1.62–2.74 (2.21) 1.96–2.57 (2.16) 1.72–5.34 (2.97) 1.93–2.38 (2.22) 1.94–3.91 (2.66) 5 100 100 0 4.85–5.86 (5.38) 5.06–5.53 (5.34) 5.07–5.42 (5.25) 5.37–5.62 (5.48) 5.07–5.44 (5.29) 6 100 0 100 4.24–5.54 (4.82) 4.15–5.18 (4.72) 4.27–4.81 (4.49) 4.29–5.11 (4.73) 4.83–5.68 (5.16) 7 0 100 100 2.19–3.19 (2.72) 1.99–3.34 (2.69) 1.43–3.26 (2.47) 2.20–3.25 (2.81) 2.16–5.37 (3.70) 8 50 50 50 4.36–5.18 (4.83) 4.31–5.07 (4.64) 4.39–4.96 (4.67) 4.48–4.93 (4.66) 4.59–6.32 (5.23) 9 50 100 50 4.63–5.26 (5.03) 4.44–5.25 (4.91) 4.77–5.38 (5.06) 4.75–5.01 (4.88) 5.22–6.28 (5.54) 10 50 100 100 4.54–5.34 (4.91) 4.56–5.26 (4.94) 4.88–5.30 (5.13) 4.50–5.48 (4.97) 5.06–5.95 (5.42) 11 100 50 50 5.28–5.72 (5.46) 5.35–5.90 (5.62) 4.88–5.64 (5.35) 5.34–5.62 (5.48) 5.30–6.16 (5.62) 12 100 100 50 5.46–5.88 (5.65) 5.21–5.80 (5.53) 5.27–5.77 (5.61) 5.64–5.68 (5.66) 5.18–6.19 (5.64) 13 100 100 100 5.50–6.12 (5.73) 5.45–6.00 (5.82) 5.41–5.82 (5.66) 5.38–5.75 (5.52) 5.24–6.22 (5.67) 14 100 150 50 5.29–6.05 (5.70) 5.36–6.19 (5.79) 5.39–6.20 (5.78) 5.40–6.09 (5.73) 5.16–6.03 (5.50) 15 100 150 150 5.65–6.14 (5.89) 5.43–6.14 (5.86) 5.55–6.04 (5.81) 5.46–5.82 (5.68) 5.18–6.17 (5.62) 16 150 50 50 5.20–5.78 (5.38) 5.08–6.05 (5.63) 5.18–5.80 (5.47) 4.65–5.65 (5.12) 5.03–6.33 (5.46) 17 150 100 50 5.06–6.14 (5.64) 4.73–6.20 (5.48) 4.97–5.96 (5.49) 5.49–6.03 (5.74) 5.19–5.95 (5.52) 18 150 100 100 5.17–5.75 (5.50) 4.30–5.81 (5.34) 4.64–5.62 (5.36) 5.26–6.01 (5.67) 5.39–5.76 (5.51) 19 150 150 100 5.38–5.83 (5.64) 4.57–5.66 (5.19) 5.27–5.70 (5.50) 5.49–5.72 (5.65) 5.12–6.24 (5.65) 20 150 150 150 5.22–5.97 (5.46) 4.28–5.68 (4.92) 5.18–5.53 (5.28) 5.46–5.83 (5.69) 5.22–6.42 (5.65) 2.2. Unmanned Aerial Vehicle (UAV) Data Acquisition For the purposes of data acquisition, high-resolution photographic images were captured on the following dates: 9 May 2022, during the Heading phase, 20 May 2022, during the Flowering phase, and 6 June 2022, during the Grainfilling phase. The DJI P4 Multispectral UAV, equipped with an MS camera, was utilized for this purpose. To ensure optimal lighting conditions and minimize potential sensor discrepancies, the aerial surveys were conducted under clear sky conditions, specifically between 12:00 and 13:00 h, corresponding to the local solar noon.
Agriculture 2025,15, 1534 7 of 32 To ensure data consistency and flight stability, all UAV missions were conducted under wind speeds below 3 m/s, and identical flight parameters were maintained across all sessions, including altitude (30 m AGL), speed (3 m/s), and image overlap (75% both longitudinally and laterally). An autonomous waypoint-based flight plan was used in each session, allowing for precise replication of flight paths. Prior to every mission, the UAV system underwent a pre-flight checklist to ensure proper sensor calibration, battery condition, and GPS connectivity. Additionally, the same pilot operated all missions to reduce variability in execution. The DJI P4 Multispectral, a quadrotor UAV (DJI, Shenzhen, China), was equipped with an integrated RGB camera and five distinct monochromatic sensors. These sensors encompass blue (B), green (G), red (R), red edge (RE), and near-infrared (NIR) bands, each tailored to specific spectral regions. Leveraging the data from the collected MS and RGB images, the study derived a set of 65 indices. This encompassed 40 MS and 25 RGB indices, offering a holistic perspective on the health and growth dynamics of the crop. A comprehensive specification of these monochromatic sensors can be found in Table 2. Table 2. Parameters of the monochromatic sensors of the multispectral camera. Band Center Wavelength/nm Bandwidth/nm Blue (B) 450 16 Green (G) 560 16 Red (R) 650 16 Red Edge (RE) 730 16 Near-infrared (NIR) 840 26 All flights were planned and executed using the DJI GS Pro software (Version: V2.0 2018.11, DJI, Shenzhen, China), ensuring consistent navigation and image capture parameters. The calibration of the sensors was verified prior to takeoff using the integrated sunlight sensor to normalize light intensity, ensuring reliable spectral readings across different times and dates. When deployed at an operational altitude of 30 m, the UAV achieved a spatial resolution of 1.6 cm per pixel. To ascertain the drone’s geospatial accuracy during its flight, we employed the D-RTK 2 high-precision mobile station (DJI, Shenzhen, China). The camera’s exposure setting was calibrated to 2 s with specified flight margins set at 5 m. The imagery was systematically captured, with each mission covering an area of 2.7 hectares. During our study, high-resolution imaging generated a substantial volume of data, culminating in a total digital footprint of 36.484 GB of raw data across three separate imaging sessions covering a combined area of 8.1 hectares. This yielded a digitization footprint of approximately 4.504 GB/ha per session. Each session entailed the acquisition of multispectral data across five spectral channels: blue, red, green, red-edge, and nearinfrared, each contributing an equal partition of the overall data volume. Consequently, the digitization footprint for each spectral channel amounted to roughly 0.901 GB/ha per session. Furthermore, orthomosaic images were created for each spectral channel in each of the three sessions, contributing an additional 7.898 GB in total to the digital footprint. When this is normalized across the total surveyed area, it results in an added digital footprint of approximately 0.975 GB/ha for the orthomosaic images. Upon summing up the total data volume for all imaging sessions (36.484 GB) and the orthomosaic images (7.898 GB), a cumulative digital footprint of 44.382 GB was calculated.
Agriculture 2025,15, 1534 8 of 32 When this total digital footprint is normalized over the surveyed area of 8.1 hectares, an overall digitization footprint of approximately 5.479 GB per hectare was obtained. In total, there were 9375 raw images captured across all sessions, with an even distribution across the five spectral channels, and the overall digitization footprint per session and per channel encompasses data from these individual images as well as from the compiled orthomosaic images. The UAV’s flight plan was algorithmically generated using the DJI GS Pro software suite (DJI, Shenzhen, China). To ensure the UAV’s optimal navigational trajectory, a solar radiation spectral sensor was used. This facilitated adaptive adjustments of photographic parameters such as ISO, white balance, and shutter speed based on real-time solar irradiance data, eliminating the nuances of manual calibrations. The drone could maintain flight for 25 to 30 min on a single battery cycle. 2.3. Data Processing Upon successful data retrieval using the DJI P4 Multispectral drone, Pix4D (Pix4Dmapper Enterprise 4.5.6, 2020, Pix4D, Prilly, Switzerland) software for a processing methodology was used, as delineated by [ 16 ]. This multifaceted workflow entailed the construction of an orthomosaic representation, aligning and amalgamating individual frames to architect a unified, geometrically congruent depiction of the experimental tract. Orthomosaic mappings for the entire RGB and MS bandwidths were subsequently synthesized, culminating in an ensemble of GeoTiff files. These georeferenced constructs not only summarized the raw image data but also provided rich metadata detailing the spatial orientation and geolocation of each constituent pixel. The preprocessing steps included radiometric correction based on real-time sunlight intensity measurements from the onboard sunlight sensor, ensuring reflectance normalization across all spectral bands. Geometric correction was achieved using the D-RTK 2 high-precision GNSS base station, enabling accurate alignment of image coordinates. Furthermore, noise filtering and shadow reduction were automatically performed within the Pix4D pipeline, based on point cloud quality and texture variation. In the subsequent phase, the synthesized GeoTiff datasets were integrated into the ArcGIS (ArcGIS-ArcMap 10.8.2, 2021, Esri Inc., 380 New York St, Redlands, CA, USA) platform for a granular analytical exercise, inspired by methodologies presented by [ 17 ]. A salient feature of this exploration was the derivation of numerical metrics for the RGB and MS spectra across each delineated plot. Additionally, the extraction of statistical descriptors for each demarcated polygonal segment within the experimental expanse was conducted. 2.4. Development of RGB and MS Indices Leveraging Extracted Spectral Data This study dissected and analyzed the numerical representations associated with the intensities observed within the red, blue, and green spectral sections, supplemented by data from the red-edge domain and the near-infrared region. After extracting these spectral data, equations for RGB and MS indices were derived, as specifically outlined in Tables A1 and A2, in the Appendix A , based on the methodologies presented in the scientific paper by [ 3 ]. The application of such indices provides a profound comprehension of the spectral characteristics of the analyzed terrain, elucidating details about their intrinsic physical and optical properties. At the specific stages of development and with the planting density employed, the canopy of the wheat plants is closed such that soil visibility through the canopy is almost nonexistent. Therefore, background noise such as soil and shadows was minimal and did not significantly impact the analysis. This dense canopy ensures that the measurements are primarily from the wheat plants themselves, providing accurate canopy information.
Agriculture 2025,15, 1534 9 of 32 2.5. Applying the PyCaret Library for Yield Prediction PyCaret is an open-source, high-level machine-learning library that offers efficient data preparation and modeling through a simple and user-friendly API [ 18 ]. The key step of this research is applying the PyCaret software on MS and RGB indices obtained by using UAVs for yield estimation. For that purpose, the correlation coefficient between the collected vegetation indices and the actual yield was examined. This step was necessary to statistically establish the relationship between the 65 calculated vegetation indices and the actual yield. After that, the normalization of 65 numerical features of the dataset was performed using the z-score method. This step is crucial for machine-learning algorithms sensitive to feature scales, as it transforms each feature to have a mean of zero and a standard deviation of one. Normalization of the features resulted in a quicker convergence of the algorithm, ultimately leading to more generalizable and robust predictions. These normalized indices were then matched with corresponding yield data collected manually from the field. This approach led to the establishment of a database containing 400 observations, each encompassing 66 feature variables, including 65 vegetation indices and wheat yield, which was employed in this research. Upon initialization, a unique session identifier (Session ID: 4797) was generated to facilitate the reproducibility of the experiment, in accordance with good scientific practice. Before training, the entire set of 65 vegetation indices was retained without applying automatic feature selection or dimensionality reduction techniques. This decision was made to preserve the full spectral information available in the dataset, allowing the machine-learning algorithms to independently evaluate the relevance of each index during model fitting. Since the PyCaret library internally handles algorithm-specific regularization and weight assignment, models such as Lasso, Ridge, and tree-based ensembles inherently perform implicit feature prioritization during training. The data was divided into a training set and a test set, with 280 and 120 samples (0.7/0.3 ratio), respectively. This partitioning was implemented to allow for robust training and evaluation of machine-learning models. The variable targeted for regression was denoted as ‘Yield’, which represents the wheat yield that the experiment aims to predict. Using machine-learning algorithms, the software learns how to recognize patterns and relationships between MS and RGB indices and yield data. Through an iterative process, the software is optimized to achieve as accurate a yield estimation as possible. A 10-fold cross-validation approach was implemented to evaluate the models’ performance, using the KFold algorithm exclusively on the training set. In this process, each fold was used once as a validation set while the k-1 remaining folds were used for the training. This validation technique provides a robust way to assess the model’s performance within the training set, minimizing the bias and variance associated with a single random partition of this subset. The built-in ‘tune model’ function of PyCaret used in our study enables automated hyperparameter tuning on the training folds, reducing manual effort and expertise required to optimize model performance within the training data. This process ensures that the model is not only optimized but also validated in a rigorous manner before being finally evaluated on a separate test set, which is critical in scientific studies where optimum model settings are essential and need to be validated effectively [19]. The tuning process utilizes grid search and random search strategies internally, depending on the algorithm, and selects hyperparameters based on performance metrics such as R 2 and RMSE across the cross-validation folds. All optimization steps were automated and reproducible, contributing to model stability. After completing the learning process, PyCaret software (Version: 3.0, Toronto, ON, Canada) can apply the learned models to new MS and RGB indices data to estimate crop yields on agricultural crop fields.
Agriculture 2025,15, 1534 16 of 32 6 June 2022 SVR Model Analysis: • Performance Metrics: For this date, the model produced a training R 2 of 0.937 and a test R 2 of 0.916. Although these numbers are somewhat lower than those from the previous dates, they still indicate strong predictive performance. With an explanation of over 93% of the data variance, the model remains valuable for agricultural planning and optimization based on its forecasts. • Residual Plot Analysis: The residuals, although mostly surrounding the zero line, showcase a more dispersed pattern as compared to the SVR model evaluated on 9 May 2022. This indicates a slightly lesser prediction accuracy for this specific dataset. • Distribution Analysis: The distribution of residuals on the right displays a near-normal pattern, albeit with a hint of right skewness, pointing towards minor prediction biases. While individual model diagnostics for each date provide granular insight, comparing residual plots and distribution shapes reveals broader trends. For example, the increase in residual spread from 9 May to 6 June reflects growing prediction uncertainty as the wheat matures. This could be linked to increasing canopy closure and spectral saturation in later stages, reducing the distinctiveness of vegetation indices. Additionally, the shift from nearly normal residual distributions to slightly skewed or bimodal forms suggests more complex error structures in later measurements, possibly due to field heterogeneity or weather variability. These insights imply that while regression models perform well throughout the season, their precision may be highest during the heading to flowering phases, as captured in the earlier dates. Considering the regression models’ performance on three distinct evaluation dates, the diagrams illustrating their capabilities are presented in Figure 3. 9 May 2022—SVR Model Analysis: • Performance Metrics: The SVR model on this date achieved an R 2 value of 0.947. In the context of wheat yield predictions, this high R 2 shows the model’s accuracy. It indicates that the model can effectively capture most of the data variance, making it reliable for wheat yield predictions. • Diagram Insight: The plot portrays predicted values against true values. A closer alignment of points with the identity line indicates better predictions. While most data points align with the “best fit” line, slight deviations highlight the model’s areas of potential improvement. 20 May 2022—MLP Regressor Model Analysis: • Performance Metrics: The MLP Regressor on this date recorded an R 2 of 0.938. Even though marginally lower than the SVR model on 9 May 2022, it remains high in accuracy. • Diagram Insight: The plotted data points largely follow the identity line, suggesting a good match between predicted and actual values. However, a few data points diverge, suggesting areas where the model might have faced challenges, possibly due to intricate patterns or outliers in the dataset. 6 June 2022—SVR Model Analysis: • Performance Metrics: On this date, the SVR model reported an R 2 of 0.916. While slightly lower than previous results, an R 2 above 0.9 in agriculture still highlights a strong predictive capability. • Diagram Insight: Most data points are near the identity line, showing the model’s consistency in predictions. The scatter, although slightly more pronounced compared to the evaluation of 9 May 2022, is still within acceptable bounds for wheat yield forecasting.
Agriculture 2025,15, 1534 17 of 32 (a) (b) (c) Figure 3. Comparative visualization of prediction errors across regression models: (a) data for 9 May 2022; (b) data for 20 May 2022; (c) data for 6 June 2022. In summary, both the residual analysis and predicted-vs-actual plots confirm the stability of model performance across all dates, while also revealing subtle variations that are relevant for practical implementation. The slightly declining accuracy and increasing residual dispersion over time underscore the importance of early-season data collection for maximizing prediction precision. These findings support the integration of temporal optimization in UAV-based monitoring workflows. Furthermore, the observed fluctuations in model rankings across dates can also be attributed to the intrinsic sensitivity of individual algorithms to the underlying data distribution and noise patterns. For instance, Support Vector Regression (SVR), which relies on optimal margin boundaries, may perform exceptionally well when the data exhibits linear or quasi-linear trends with minimal noise, as observed on 9 May and 6 June. However, on 20 May, where data characteristics may have been more nonlinear or affected by subtle outliers, the MLP Regressor, known for its capability to capture complex patterns through neural layers, outperformed SVR. These differences highlight the importance of aligning
Agriculture 2025,15, 1534 18 of 32 model selection with data-specific traits, ensuring the robustness of predictive analytics in agricultural scenarios. The results from the specified evaluation dates provide a nuanced understanding of the different regression models’ capabilities. While each model showed high predictive power, subtle differences in performance metrics and visual insights underline the importance of continuous evaluation. It is evident that, while all models offer high value in predicting wheat yield, their effectiveness can vary depending on the data’s characteristics and complexities. This rigorous assessment underscores the significance of selecting the right model for specific datasets and the need for ongoing validation to ensure optimal performance in the field of wheat yield prediction. 4. Discussion The findings from this study can offer novel insights into the application of machine learning for predicting wheat yields [ 19 ]. The following sections break down the discussions based on the derived results. 4.1. The Potential of Spectral Indices for Yield Prediction The results of the correlation coefficient unveil the relationship between various indices and wheat yield. The focus here is not only on the strength of the correlation but also on the type (positive or negative) and the consistency over time. The MS indices show a consistently high correlation with wheat yield across the three dates. NDVI, a widely recognized index for vegetation vigor, shows consistently high correlation, underscoring its reliability in predicting wheat yields [ 20 ]. Blue Normalized Difference Vegetation Index (BNDVI) and Structure Insensitive Pigment Index (SIPI), particularly for the 9 May 2022 measurement, display strikingly high positive correlations, suggesting their potential utility in predicting wheat yields. The almost identical correlation coefficient of these indices suggests they capture similar variability in the data, which might be indicative of chlorophyll content or general plant health [ 21 ]. However, it is crucial to note the changing nature of these correlations over different dates. The indices showed the highest average correlation on 9 May 2022 and 20 May 2022, indicating a possible relationship with a specific growth stage of wheat. These high correlations might correspond to periods when the wheat plants are at their heading vegetative stage, or when the grain-filling is taking place, emphasizing the significance of the timing in utilizing these indices for yield prediction [22,23]. On the other hand, the correlation coefficients for the RGB indices seem to be more varied and generally weaker compared to the MS indices. Modified Green–Red Vegetation Index (MGRVI) on 9 May 2022 and Color Intensity (INT) on 20 May 2022 stood out with their strong positive and negative correlations, respectively. Intriguingly, there was a marked decrease in the absolute average value of the correlation coefficient from 0.83 to 0.78 during the initial two dates, culminating in a mere 0.40 by 6 June 2022 [ 24 ]. This decline can potentially be ascribed to the physiological shifts in pigment composition as the wheat undergoes maturation processes [ 25 ]. The dynamism of these pigment alterations throughout the developmental stages is likely instrumental in the observed diminishment of RGB index values by 6 June 2022 [26]. Drawing from these results, the selection of features for machine-learning models dedicated to wheat yield prediction becomes paramount. Indices with consistently high correlation coefficients, such as BNDVI, SIPI, and Simple Ratio Red/NIR Ratio VegetationIndex (ISR) from the MS indices, should be prioritized in model training [ 27 ]. Their reliability suggests that they encapsulate vital information related to wheat growth and, by extension, yield [ 28 ]. Conversely, while RGB indices offer valuable insights, their inclusion
Agriculture 2025,15, 1534 19 of 32 should be approached with caution, especially considering their fluctuating correlations over time. Nonetheless, they should not be entirely dismissed as they might enhance the model’s predictive capability in conjunction with other indices. 4.2. Implications for Machine-Learning Models The analysis using PyCaret stands out for its capability to automatically fine-tune hyperparameters for 25 different machine-learning models, simplifying the process of identifying the most suitable model for a given dataset, as stated by [ 3 , 19 ]. PyCaret highlighted that the SVR model was especially proficient in predicting wheat yield for 9 May 2022 and 6 June 2022, with notable R 2 values of 0.95 and 0.91, and RMSEs of 0.26 and 0.33 t/ha, respectively. This high coefficient of determination suggests that the SVR model can explain a significant proportion of the variability in the wheat yield for these specific dates. Meanwhile, for 20 May 2022, the MLP Regressor (Neural Network) stood out with an R 2 of 0.94 and an RMSE of 0.28 t/ha, indicating its capability to model the dataset effectively for that date. The results of this study, with significantly higher precision, likely stem from the comprehensive nature of the experimental setup. A diverse dataset comprising 400 distinct experimental plots was utilized, providing broad variation and richness of data. This extensive dataset facilitated the effective application of machine-learning models, ensuring a high degree of accuracy in yield predictions. The use of a large number of vegetation indices and the application of 25 different machine-learning models helped in precisely classifying mathematical models by accuracy criteria. The study [ 28 ] also focused on predicting wheat yield using UAV imagery during the same vegetative period as this research. They reported that, like these findings, the SVR model and the Deep Neural Network (DNN) model were among the most effective for yield prediction. However, their approach utilized a smaller set of vegetative indices (10 RGB and 16 MS indices) compared to this study. This limitation in the range of indices may contribute to the generally lower precision of their models, with the SVR model achieving an R 2 ranging from 0.502 to 0.666, and the DNN model ranging from 0.489 to 0.670 in terms of R 2 . Importantly, ref. [ 2 ] emphasized that the variation in R 2 values within their models is directly linked to the incorporation of a more extensive array of vegetative indices, suggesting that a combination of both RGB and MS indices yields a more accurate model compared to using either RGB or MS indices in isolation. Ref. [ 29 ] noted a significant correlation between NDVI and early yield prediction for winter wheat, with R 2 ranging from 0.69 to 0.90 for different varieties, reinforcing the importance of vegetation indices in yield prediction. Ref. [ 2 ] combined the Agricultural Production Systems Simulator (APSIM)—simulated biomass with extreme climatic conditions and vegetation indices like NDVI and Standardized Precipitation Evapotranspiration Index (SPEI) in a hybrid approach using Random Forest (RF) and regression models. Their findings underline the critical role of adapting models to different environmental conditions, particularly highlighting drought as a significant disruptor of yield predictions. Ref. [ 30 ] demonstrated that deep-learning models like Convolutional Neural Network (CNN) and Long Short-Term Memory (LSTM) are highly effective in predicting crop yields, suggesting the potential for the methodology applied in this study to incorporate these models for more robust predictions. Comparatively, ref. [ 31 ] showed that machine-learning methods, particularly when integrating climatic and satellite data, outperform traditional regression techniques in wheat yield prediction. Their use of Enhanced Vegetation Index (EVI) provided better results compared to Sun-Induced Fluorescence (SIF) due to lower inherent noise, with optimal predictive capability achieved approximately two months before wheat maturity. This aligns with the approach used in this study of leveraging a wide range of indices,
Agriculture 2025,15, 1534 20 of 32 including BNDVI, SIPI, and ISR, which have demonstrated strong correlations with yield and contributed to the high accuracy of the models. Furthermore, ref. [ 1 ] indicated that Partial-Least-Squares Regression (PLSR) achieved the most precise results using both spectral indices and plant height data, highlighting the advantage of PLSR in interpreting hyperspectral data for yield estimation. This insight suggests that incorporating plant structural information alongside spectral data could enhance model accuracy. Ref. [ 32 ], who employed UAVs equipped with MS and thermal infrared cameras in a study on winter wheat, achieved the best results using a combination of LSTM neural networks and RF with an R 2 of 0.78 and RMSE of 0.68 t/ha. Ref. [ 33 ] demonstrated that machine-learning models like Support Vector Machine (SVM), RF, AdaBoost, and DNN outperformed linear regression techniques for wheat yield prediction in the United States. Particularly, AdaBoost achieved an R 2 of 0.86 and RMSE of 0.51 t/ha, showcasing the effectiveness of combining various data sources for yield forecasting up to 2.5 months before harvest. Finally, ref. [ 34 ] highlighted the efficiency of CNNs in predicting wheat yield in Germany with an R 2 of 0.78, underscoring the applicability of deep-learning models in diverse geographic contexts. While their findings reflect the potential of advanced machine-learning models in yield prediction, the lower R 2 values compared to the results of this study highlight the effectiveness of using an extensive dataset and a broad range of indices, as well as a comprehensive experimental design. Furthermore, when compared to previous studies in the domain of UAV-based yield prediction, the models developed in this research demonstrate notably superior predictive performance, particularly in terms of R 2 and RMSE values. Prior research typically reported R2values ranging from 0.50 to 0.90, depending on factors such as crop type, phenological stage, and the spectral data sources utilized. The exceptionally high accuracy achieved in this study—exceeding an R 2 of 0.94 and RMSE below 0.30 t/ha—can be primarily attributed to several key methodological strengths. First, the use of a large and diverse dataset encompassing 400 experimental plots enabled a broad representation of yield variability. Second, the integration of a comprehensive set of vegetation indices, including both RGB and multispectral (MS) indices, allowed for a more detailed and nuanced modeling of crop biophysical characteristics. Third, the consistent application of 25 different machine-learning algorithms, followed by performance-based model selection, ensured both robust optimization and fair comparative evaluation. Unlike many earlier studies that focused on a narrow temporal window or specific phenological phases, this research assessed model performance across multiple key growth stages. These methodological advantages collectively provided a strong framework for minimizing bias and overfitting, thereby resulting in more generalizable and reliable models. This study clearly demonstrates that precision and reproducibility in yield prediction can be significantly enhanced when spectral diversity, timely measurements, and automated modeling frameworks are systematically integrated. The broader temporal assessment further contributes to the robustness of the results, reinforcing this research’s contribution to the expanding field of UAV-based yield prediction. 4.3. Limitations and Future Scope This study accentuates the importance of understanding wheat yield variability and harnessing the power of spectral indices for its prediction. By leveraging the insights derived, machine-learning models can be optimized for predicting wheat yields with enhanced precision [ 27 ]. While certain indices prove more consistent and reliable, a holistic approach that considers both the strengths and limitations of each index is pivotal for advancing the agricultural domain through technology. Such insights not only propel the
Agriculture 2025,15, 1534 21 of 32 scientific understanding forward but also promise tangible benefits for the agricultural community, ensuring food security and sustainability in the long term [33]. While this study achieved high accuracy in wheat yield prediction using MS UAV data, there is still a need for further research to enhance the robustness and generalizability of these models. Refs. [ 9 , 10 ] demonstrated in their works the effectiveness of combining a successive projections algorithm (SPA) with an LSTM for improving the precision of aboveground biomass estimation. This approach not only helped in identifying the most relevant spectral features but also in capturing the temporal dynamics of crop growth across different phenological stages. By integrating these advanced methodologies, future studies could address current limitations and improve the adaptability of wheat yield prediction models to varying environmental conditions and crop management practices. Incorporating SPA and LSTM techniques with hyperspectral data, as exemplified by Liu et al., could potentially lead to even more accurate and reliable predictions, ensuring better resource allocation and management in agricultural systems [ 35 , 36 ]. Another limitation of this study is its reliance solely on spectral variables, which, while valuable, may not capture the full complexity of factors affecting wheat yield. Future research should consider integrating additional data types, such as environmental and agronomic variables, including precipitation levels, cumulative temperature sums, and soil composition, to improve the robustness of yield predictions. Additionally, this study focuses on European wheat varieties, limiting the generalizability of findings to other regions with different wheat cultivars and environmental conditions. To push the boundaries of current methodologies, future research could explore the integration of real-time UAV data streams with edge-computing systems for in-field yield forecasting. The incorporation of federated learning approaches may also enable collaborative model training across geographically distributed farms without compromising data privacy. Furthermore, the fusion of spectral data with emerging technologies such as soil health sensors, autonomous ground robots, and high-resolution satellite constellations could unlock new frontiers in precision agriculture. Adopting such forward-looking, multidisciplinary strategies would align future research with global trends in smart farming and agricultural digitalization. Incorporating diverse wheat varieties, along with climate and soil data, would help create a more adaptable model suited to a broader range of conditions and management practices, thereby enhancing the applicability and precision of yield predictions across varied agricultural settings. Furthermore, certain limitations stem from potential inconsistencies in UAV image quality due to varying light conditions, atmospheric interference, or flight path variations, which could introduce noise into the data and affect index accuracy. Although care was taken to conduct flights during consistent midday conditions, slight variabilities remain a possible source of error. In terms of model selection, despite the use of 25 machine-learning algorithms, the study did not incorporate deep-learning architectures such as CNNs or hybrid models, which have shown promise in related studies. The omission of these models might have limited the full exploration of performance boundaries. Additionally, due to the high dimensionality of the dataset, there is a risk of overfitting, especially when using complex algorithms on relatively small samples from specific measurement dates. Although cross-validation and tuning were performed to mitigate this, future studies should validate model performance with external datasets to confirm generalizability. 5. Conclusions This study evaluated 25 regression models for estimating wheat yield using multispectral (MS) and RGB vegetation indices derived from UAV imagery, tested across
Agriculture 2025,15, 1534 22 of 32 three different measurement dates. Among these, Support Vector Regression (SVR) and Multi-Layer Perceptron (MLP) Regressor consistently delivered the highest predictive performance, with SVR achieving R 2 values of 0.95 and 0.91 and RMSE values of 0.26 and 0.33 t/ha, while MLP Regressor achieved R 2 = 0.94 and RMSE = 0.28 t/ha. These models were capable of explaining up to 95% of the variation in wheat yield. The study also revealed important differences among five European wheat cultivars. NS Obala recorded the highest average yield (5.03 t/ha) and lowest variability ( CV = 21.63% ), while NS Rajna had the lowest average yield (4.69 t/ha) and the highest variability (CV = 27.55%). These results underscore the robustness of the models when applied to genetically diverse material. Key MS indices such as BNDVI and SIPI exhibited strong and consistent correlations with yield, while RGB indices were more variable. These insights suggest the need for caution when using RGB indices for yield prediction. Overall, the findings advance our understanding of integrating UAV-derived data with machine-learning models for crop yield prediction. This work contributes to filling the knowledge gap in temporal model performance, cultivar-specific variability, and the effectiveness of MS vs. RGB indices. The practical implications support the development of adaptable tools for precision agriculture and informed decision-making in wheat production. Author Contributions: Conceptualization, K.K., Z.S., M.K. and V.A.; methodology, K.K., Z.S., M.K. and V.A.; software, K.K., D.T. and T.N.; validation, Z.S., M.K. and N.M.; formal analysis, K.K., D.T., T.N. and A.I.; investigation, K.K. and Z.S.; resources, V.A.; data curation, K.K. and M.I.; writing— original draft preparation, K.K. and Z.S.; writing—review and editing, M.K., M.I. and N.M.; visualization, M.I. and A.I.; supervision, N.M. and M.I.; project administration, N.M.; funding acquisition, N.M. All authors have read and agreed to the published version of the manuscript. Funding: This publication is part of the TALLHEDA project that has received funding from the European Union’s Horizon Europe research and innovation programme under grant agreement No. 101136578. Funded by the European Union. Views and opinions expressed are however those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Executive Agency (REA). Neither the European Union nor the granting authority can be held responsible for them. Data Availability Statement: Data are contained within the article. Acknowledgments: The financial support mentioned in the Funding part is gratefully acknowledged. Conflicts of Interest: The authors declare no conflict of interest.
Agriculture 2025,15, 1534 23 of 32 Appendix A Figure A1. Workflow Diagram for Wheat Yield Prediction Using UAV-Based Data and MachineLearning Models. Table A1. Vegetation indices, accompanied by their computational equations, derived from the numerical intensities of RGB bands [3]. Index Type Formula Red (R), Green (G), Blue (B) Numerical values Normalized Red r=Red Red+Green+Blue Normalized Green g=Green Red+Green+Blue Normalized Blue b=Blue Red+Green+Blue Green–Red Ratio Index GRRI =Green Red
Agriculture 2025,15, 1534 24 of 32 Table A1. Cont. Index Type Formula Green–Blue Ratio Index GBRI =Green Blue Red–Blue Ratio Index RBRI =Red Blue Excess Red Vegetation Index ExR =1.4·r−g Excess Green Vegetation Index ExG =2·g−r−b Excess Blue Vegetation Index ExB =1.4·b−g Excess Green Minus Excess Red Index ExGR =ExG −ExR Woebbecke Index WI =Green−Blue Green+Red Normalized Difference Index NDI =r−g r+g+0.01 Color Intensity INT =Red+Blue+Green 3 Green Leaf Index 1 GLI1=2·Green−Red−Blue 2·Green+Red+Blue Green Leaf Index 2 GLI2=2·Green−Red+Blue 2·Green+Red+Blue Vegetative Index VEG =Red(2/3)·b(1/3) Red Color Index of Vegetation CIVE =0.441·r−0.811·g+0.3856·b+18.79 Combination COM =0.25·ExG +0.3·ExGR +0.33·CIVE +0.12·VEG Normalized Green–Red Vegetation Index NGRVI =Green−Red Green+Red Kawashima Index IKAW =Red−Blue Red+Blue Visible-band difference vegetation Index VDVI =2·g−r−b 2·g+r+b Visible Atmospherically Resistance Index VARI =g−r g+r−b Principal Component Analysis Index IPCA = 0.994·|Red −Blue|+0.961·|Green −Blue|+0.914·|Green −Red| Modified Green–Red Vegetation Index MGRVI =Green2−Red2 Green2+Red2 Red–Green–Blue Vegetation Index RGBVI =Green2−Blue·Red Green2+Blue·Red Table A2. Vegetation indices, along with their associated equations, found on the numerical intensities of MS bands. Index Type Formula Reference Normalized Difference Vegetation Index NDVI =NIR−Red NIR+Red [6] Renormalized Difference Vegetation Index RDVI =NIR−Red (NIR+Red)1/2 [7] Difference vegetation index DVI =NIR −Red [4] Blue normalized difference vegetation index BNDVI =NIR−Blue NIR+Blue [6] Green normalized difference vegetation index GNDVI =NIR−Green NIR+Green [37] Modified Soil Adjusted Vegetation Index MSAVI =2NIR+1−√(2NIR+1)2−8(NIR−Red) 2[38] Red-Edge Chlorophyll Vegetation Index ReCI =NIR Red−1[39] Normalized Difference Red Edge Index NDRE =NIR−Red Edge NIR+Red Edge [37] Normalized Difference Water Index NDWI =Green−NIR Green+NIR [40]
Agriculture 2025,15, 1534 25 of 32 Table A2. Cont. Index Type Formula Reference Optimized Soil-Adjusted Vegetation Index OSAVI =(1+0.16)·(NIR−Red) (NIR+Red+0.16)[37] The Simple Ratio SR =NIR Red [37] Modified Simple Radio MSR =NIR Red −1 (NIR Red )1/2+1 [41] Infrared percentage vegetation index IPVI =NIR NIR+Red [38] Enhanced Vegetation Index EVI =2.5·(NIR−Red) (NIR+6·Red−7.5·Blue+1)[42] Green atmospherically resistant vegetation index GARI =(NIR−(Green−(Blue−Red))) (NIR−(Green+(Blue−Red))) [43] Soil-Adjusted Vegetation Index SAVI =(NIR−Red) (NIR+Red+0.5)·(1+0.5)[38] Green Soil Adjusted Vegetation Index GSAVI =(NIR−Green) (NIR+Green+0.5)·(1+0.5)[44] Green Optimized Soil Adjusted Vegetation Index GOSAVI =NIR−Green NIR+Green+0.16 [44] Green Chlorophyll Vegetation Index GCI =NIR Green −1[45] Plant Senescence Reflectance Index PSRI =Red−Green NIR [46] Nonlinear vegetation index NLI =NIR2−Red NIR2+Red [47] Transformed difference vegetation index TDVI =1.5·(NIR−Red) q(NIR2+Red+0.5)[48] Visible Atmospherically Resistant Index VARI =Green−Red Green+Red−Blue [49] Wide Dynamic Range Vegetation Index WDRVI =0.1NIR−Red 0.1NIR+Red [43] Green-Red NDVI GRNDVI =NIR−(Green+Red) NIR+(Green+Red)[41] Green-Blue NDVI GBNDVI =NIR−(Green+Blue) NIR+(Green+Blue)[50] Red-Blue NDVI GRNDVI =NIR−(Red+Blue) NIR+(Red+Blue)[50] Pan NDVI PNDVI =NIR−(Green+Red+Blue) NIR+(Green+Red+Blue)[50] Simple Ratio Red/NIR Ratio Vegetation-Index ISR =Red NIR [51] Leaf Chlorophyll Index LCI =NIR−Red Edge NIR+Red [4] Ratio Between NIR and Green Bands VI(NIR Green )=NIR Green [52] Ratio Between NIR and Red Edge Bands VI(NIR Red Edge )=NIR Red Edge [53] Simplified Canopy Chlorophyll Content Index SCCCI =NDRE NDVI [3] Modified Chlorophyll Absorption Reflectance Index MCARI = (Red Edge −Red −0.2·(Red Edge −Green))·Red Edge Red [4] Transformed Chlorophyll Absorption Reflectance Index TCARI = 3·(Red Edge −Red)−0.2·(Red Edge −Green)·Red Edge Red [37] Structure Insensitive Pigment Index SIPI =NIR−Blue NIR−Red [41] TCARI/OSAVI TCARI OSAVI [41] MCARI/OSAVI MCARI OSAVI [41] Red-Edge Chlorophyll Index 1 CI1=NIR Red Edge −1[54] Red-Edge Chlorophyll Index 2 CI2=Red Edge Green −1[55] Table A3. List of Regression Models with Corresponding Abbreviations. Model Type Model Abbreviation Linear Regression lr Lasso Regression lasso
Agriculture 2025,15, 1534 32 of 32 54. Zhang, H.; Li, J.; Liu, Q.; Lin, S.; Huete, A.; Liu, L.; Croft, H.; Clevers, J.G.P.W.; Zeng, Y.; Wang, X.; et al. A novel red-edge spectral index for retrieving the leaf chlorophyll content. Methods Ecol. Evol. 2022,13, 2771–2787. [CrossRef] 55. Clevers, J.G.P.W.; Gitelson, A.A. Remote estimation of crop and grass chlorophyll and nitrogen content using red-edge bands on Sentinel-2 and -3. Int. J. Appl. Earth Obs. Geoinf. 2013,23, 344–351. [CrossRef] Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.