Full text
Supplementary Materials for ”Adjusting Expected Goals (xG) and Shots on Target for Game Context in Soccer” Andrey Skripnikov1∗ , Ahmet Cemek1, and David Gillman1 1New College of Florida, Sarasota, FL, USA October 11, 2025 1 Goodness-of-Fit Tests 1.1 Expected Goals (xG) as response The DHARMa goodness-of-fit tests for the models we used for expected goals (xG) clearly showcase the inadequacy of Gaussian approaches (top row of Figure 1). The quantile–quantile plots show large deviations from assumptions, and the residuals-vs-fitted plots reveal systematic skew. For Tweedie models, the quantile–quantile plots demonstrate much better alignment with the theoretical quantiles, while the residuals-vs-fitted plots show a more adequate fit, especially for the log-link Tweedie model. ∗Corresponding author: askripniko[email protected] 1
0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.2 0.4 0.6 0.8 1.0 QQ plot residuals Expected Observed KS test: p= 0 Deviation significant Outlier test: p= 0 Deviation significant Dispersion test: p= 0.984 Deviation n.s. Model predictions (rank transformed) DHARMa residual 0.0 0.2 0.4 0.6 0.8 1.0 0.00 0.25 0.50 0.75 1.00 DHARMa residual vs. predicted Quantile deviations detected (red curves) Combined adjusted quantile test significant Bundesliga, Gaussian GLM 0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.2 0.4 0.6 0.8 1.0 QQ plot residuals Expected Observed KS test: p= 0 Deviation significant Outlier test: p= 0 Deviation significant Dispersion test: p= 0.88 Deviation n.s. Model predictions (rank transformed) DHARMa residual 0.0 0.2 0.4 0.6 0.8 1.0 0.00 0.25 0.50 0.75 1.00 DHARMa residual vs. predicted Quantile deviations detected (red curves) Combined adjusted quantile test significant Bundesliga, Log−Link Gaussian GLM 0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.2 0.4 0.6 0.8 1.0 QQ plot residuals Expected Observed KS test: p= 0 Deviation significant Outlier test: p= 0 Deviation significant Dispersion test: p= 0 Deviation significant Model predictions (rank transformed) DHARMa residual 0.0 0.2 0.4 0.6 0.8 1.0 0.00 0.25 0.50 0.75 1.00 DHARMa residual vs. predicted Quantile deviations detected (red curves) Combined adjusted quantile test significant Bundesliga, Log−Link Tweedie 0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.2 0.4 0.6 0.8 1.0 QQ plot residuals Expected Observed KS test: p= 0 Deviation significant Outlier test: p= 0 Deviation significant Dispersion test: p= 0 Deviation significant Model predictions (rank transformed) DHARMa residual 0.0 0.2 0.4 0.6 0.8 1.0 0.00 0.25 0.50 0.75 1.00 DHARMa residual vs. predicted Quantile deviations detected (red curves) Combined adjusted quantile test significant Bundesliga, Inverse−Link Tweedie Figure 1: Visual representation of DHARMa model diagnostics with expected goals (xG) as the response variable. Diagnostics include quantile–quantile plots (left) and residuals-vs-fitted plots (right) for four models: Gaussian (top left), log-link Gaussian (top right), log-link Tweedie (bottom left), and inverse-link Tweedie (bottom right). 2
1.2 Shots on Target as Response After running models for each league–season combination with shots on target as the response variable, Figures 2 and 3 display the DHARMa goodness-of-fit results. Figure 2 shows Holm-adjusted p-values for various tests (uniformity, outlier, overdispersion, and zero-inflation), while Figure 3 shows the quantile–quantile and residuals-vs-fitted plots. Both figures point to Poisson-family models (Poisson and Negative Binomial) as superior to Gaussian models, with a slight edge to Negative Binomial, particularly for outlier tests. Figure 2: P-values from DHARMa model diagnostics for each model type predicting shots on target. Tests include uniformity (Kolmogorov–Smirnov), outlier, overdispersion, and zero-inflation tests. The null hypothesis in each test is that model assumptions are satisfied. 3
0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.2 0.4 0.6 0.8 1.0 QQ plot residuals Expected Observed KS test: p= 0 Deviation significant Outlier test: p= 0 Deviation significant Dispersion test: p= 0.984 Deviation n.s. Model predictions (rank transformed) DHARMa residual 0.0 0.2 0.4 0.6 0.8 1.0 0.00 0.25 0.50 0.75 1.00 DHARMa residual vs. predicted Bundesliga, Gaussian GLM 0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.2 0.4 0.6 0.8 1.0 QQ plot residuals Expected Observed KS test: p= 0 Deviation significant Outlier test: p= 0 Deviation significant Dispersion test: p= 0.984 Deviation n.s. Model predictions (rank transformed) DHARMa residual 0.0 0.2 0.4 0.6 0.8 1.0 0.00 0.25 0.50 0.75 1.00 DHARMa residual vs. predicted Bundesliga, Log−Link Gaussian GLM 0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.2 0.4 0.6 0.8 1.0 QQ plot residuals Expected Observed KS test: p= 0.65993 Deviation n.s. Outlier test: p= 0.22118 Deviation n.s. Dispersion test: p= 0.032 Deviation significant Model predictions (rank transformed) DHARMa residual 0.0 0.2 0.4 0.6 0.8 1.0 0.00 0.25 0.50 0.75 1.00 DHARMa residual vs. predicted Bundesliga, Poisson 0.0 0.2 0.4 0.6 0.8 1.0 0.0 0.2 0.4 0.6 0.8 1.0 QQ plot residuals Expected Observed KS test: p= 0.49806 Deviation n.s. Outlier test: p= 0.57962 Deviation n.s. Dispersion test: p= 0.04 Deviation significant Model predictions (rank transformed) DHARMa residual 0.0 0.2 0.4 0.6 0.8 1.0 0.00 0.25 0.50 0.75 1.00 DHARMa residual vs. predicted Bundesliga, Negative Binomial Figure 3: Visual representation of DHARMa diagnostics with shots on target as the response variable. Diagnostics include quantile–quantile plots (left) and residuals-vs-fitted plots (right) for four models: Gaussian (top left), log-link Gaussian (top right), Poisson (bottom left), and Negative Binomial (bottom right). 4