Document Type : Research Article
Author
Department of Agrotechnology, Faculty of Agriculture, Ferdowsi University of Mashhad, Mashhad, Iran
Abstract
Introduction
The escalating degradation of water and soil resources, coupled with global warming, population growth, and climate change, necessitates a reevaluation of food production systems. Corn (Zea mays L.), as the third most strategic crop worldwide, faces increasing demand due to rising food needs. Optimizing seed yield under varying environmental conditions is critical, with arbuscular mycorrhizal fungi (AMF) and plant growth-promoting rhizobacteria (PGPR) playing pivotal roles by enhancing nutrient uptake (e.g., phosphorus and nitrogen) and improving plant resilience to stresses and pathogens. However, the complex eco-physiological interactions influencing seed yield remain poorly understood, underscoring the need for advanced modeling tools. Traditional regression models often fail to capture non-linear and interactive relationships among factors such as photosynthesis rates, nutrient levels, and plant anatomical traits. Consequently, machine learning models like Deep Neural Networks (DNN) and Transformers have emerged as powerful alternatives for precision agriculture. This study aims to identify key features affecting corn seed yield using stepwise regression and compare eight machine learning algorithms, Enhanced DNN, Transformer, XGBoost, SVM, ANN, ANFIS, LightGBM, and SVR, to determine the most accurate model for predicting yield and unveiling hidden eco-physiological relationships.
Materials and Methods
The experiments were conducted over two consecutive years at the research farm of Ferdowsi University of Mashhad, Iran (latitude 36°15'N, longitude 59°28'E, altitude 985 m), located in the Kashafroud River basin. The region features a semi-arid climate with an average annual rainfall of 252 mm and a mean temperature of 15°C, with loamy soil of moderate organic carbon content. The dataset comprised 96 samples and 73 features, collected over two growing seasons, including plant traits (e.g., leaf chlorophyll content via SPAD index, nitrogen concentration) and soil characteristics. Of these, 32 were primary features (e.g., photosynthesis rate, leaf area index, canopy temperature), and 41 were engineered interaction features (e.g., Pmax, Mean_Leaf Area Index, Canopy Temp, 1_Root Colonization), designed based on domain knowledge of agro-ecophysiology, soil ecology, AMF, and PGPR. Data were randomly split (random_state = 42) into 70% training (67 samples), 15% validation (14 samples), and 15% test (15 samples) sets, with standardization applied using StandardScaler. Feature selection employed stepwise backward regression, reducing 73 features to 13 key variables (e.g., Canopy Temp_3, % P plant) based on adjusted R², multicollinearity, and variance inflation factor (VIF). Among the 15 machine learning algorithms evaluated, eight were configured with model-specific architectures, including an Adaptive Neuro-Fuzzy Inference System (ANFIS) implemented using TensorFlow-based layers, a Transformer model with two attention heads, an Artificial Neural Network (ANN) with three hidden layers, and a Support Vector Regression (SVR) model with a radial basis function (RBF) kernel. Model performance was evaluated using the coefficient of determination (R²), root mean square error (RMSE), mean absolute error (MAE), and Willmott's index of agreement (d). In addition, the Shapiro–Wilk test was used to assess the normality of model residuals.
Results and Discussion
Stepwise regression yielded a model with an adjusted R² of 58.53% and a predictive R² of 51.06%, identifying 13 features (e.g., Canopy Temp_3, coefficient = 0.2451, p = 0.002). Machine learning results identified ANFIS (R² = 0.555), the Transformer model (R² = 0.545), and ANN (R² = 0.518) as the top-performing models, whereas SVM (R² = 0.325) exhibited the lowest performance. The Nemenyi diagram ranked ANN first (mean rank = 1.50, Willmott's d = 0.848), followed by the Transformer model, with a critical distance of 1.11 indicating significant differences (α = 0.05) between the highest- and lowest-ranked models. Taylor diagrams highlighted the superior performance of the Transformer model (RMSE = 2.834, R² = 0.545), with ANFIS and ANN performing similarly. SHAP plots revealed that interactions such as Leaf Area Index–Dry Matter Yield were among the most influential predictors, while SVR shared nine features with the regression model. Correlation analysis grouped the features into physiological (e.g., SPAD_Mean) and yield-related (e.g., Dry Matter Yield) categories, with strong positive correlations (e.g., SPAD_2 and cob diameter, r = 0.82) and strong negative correlations (e.g., Canopy Temp_Mean and Root Colonization, r = −0.82). Skewed distributions in KDE plots underscored the need for non-linear models. Force plots (e.g., LightGBM sample) showed features like Leaf Area Index_Dry Matter Yield (+280.047) driving predictions, while loss curves for Enhanced DNN indicated effective convergence.
Conclusion
The study suggests that neural network-based models excel in capturing complex eco-physiological interactions, with canopy temperature and root-nitrogen interactions as key predictors, offering insights for sustainable corn production despite data size limitations.
Keywords
Subjects