R. Fonzetti , D. Bailo and L. Trani 1Istituto Nazionale di Geofisica e Vulcanologia, Rome (RM), Italy. 2Royal Netherlands Meteorological Institute (KNMI), De Bilt, Netherlands. *
[email protected] *,1 2 1 Unbalanced Learning: Deep Training Challenges in Seismic Phase Detection PROBLEM AND AIM TRAINING MODEL TESTING ON REAL DATA INVESTING IN THE ROOT CAUSE OF THE PROBLEM TAKE HOME MESSAGES -- 63,704 earthquakes (2009-04-06 up to 2009-12-20); -- Multiple P and S pick arrivals that allow improvements in current Machine Learning algorithms development and analysis (recognise seiemic sequence patterns). 50,000 earthquakes (time period: 2000-2021) with magnitude scale of events ranges from 0 - 6.5. Seismograms with the pickings of the validation test. For the selected tests, the Loss functions of this model are provided. The Low Recall of S-wave (max about 60%) depends on the different amount of Pand S-waves, with a ratio of ~ 4 (P: 1559462 vs. S: 388996). Bagagli, M., Valoroso, L., Michelini, A., Cianetti, S., Gaviano, S., Giunchi, C., Jozinović, D., & Lauciani, V. (2023). AQ2009 – The 2009 Aquila Mw 6.1 earthquake aftershocks seismic dataset for machine learning application. Istituto Nazionale di Geofisica e Vulcanologia (INGV). https://doi.org/10.13127/AI/AQUILA2009. Michelini, A., Cianetti, S., Gaviano, S., Giunchi, C., Jozinović, D., & Lauciani, V. (2021). INSTANCE - The Italian Seismic Dataset For Machine Learning. Istituto Nazionale di Geofisica e Vulcanologia (INGV). https://doi.org/10.13127/INSTANCE. Valoroso, L., Chiaraluce, L., Piccinini, D., Di Stefano, R., Schaff, D., & Waldhauser, F. (2013). Radiography of a normal fault system by 64,000 high‐precision earthquake locations: The 2009 L'Aquila (central Italy) case study. Journal of Geophysical Research: Solid Earth, 118(3), 1156-1176. Zhu, W., & Beroza, G. C. (2019). PhaseNet: a deep-neural-network-based seismic arrival-time picking method. Geophysical Journal International, 216(1), 261-273. Zhu, W., McBrearty, I. W., Mousavi, S. M., Ellsworth, W. L., & Beroza, G. C. (2022). Earthquake phase association using a Bayesian Gaussian mixture model. Journal of Geophysical Research: Solid Earth, 127(5), e2021JB023249. Phase Picking Phase Association Seismic catalog Seismological analysis Model Training AQ2009 (Bagagli et al., 2023; Valoroso et al., 2013) PhaseNet (Zhu and Beroza, 2019) GaMMA (Zhu et al., 2022) Application on 2016-2017 Amatrice-Visso-Norcia seismic sequence (July 2016-Febraury 2017) 4D Non Linear Seismic Tomography Creation of a robust workflow that, from the picking of continuous waveforms (specifically downloaded and/or locally stored) to the association of seismic phases, produces a seismic catalog for use in seismological purposes. New deep learning algorithms for picking and association seismic waves have been implemented. The neural network in question has also undergone specific training in the study area. AQ2009 DATASET (Bagagli et al., 2023; Valoroso et al., 2013) INSTANCE DATASET (Michelini et al., 2021) E BS Loss Function LR T P TS PR-P PR-S R-P R-S F1 P F1 S 60 8192 Focal Loss: gamma 3, alpha= (p=1, s=7, noise=1) 5,00E-04 0,55 0,55 0,93 0,64 0,67 0,81 0,78 0,72 70 8192 Focal Loss: gamma 3, alpha= (p=3, s=7, noise=1) 5,00E-04 0,55 0,55 0,88 0,64 0,74 0,81 0,80 0,71 100 8192 Focal Loss: gamma 3, alpha= (p=3, s=7, noise=1) 5,00E-04 0,55 0,55 0,89 0,65 0,74 0,81 0,81 0,72 50 4096 Focal Loss: gamma 1.0 (1,1.25,3.0) 1,00E-04 0,55 0,55 0,95 0,87 0,62 0,36 0,75 0,51 50 4096 Cross Entropy 1,00E-04 0,55 0,55 0,92 0,80 0,70 0,55 0,80 0,65 50 8192 Focal Loss: gamma 1.0 (1,1.25,3.0) 1,00E-05 0,55 0,55 0,91 0,78 0,69 0,46 0,78 0,58 100 8192 Focal Loss: gamma 3, alpha= (p=3, s=7/4, noise=1) 5,00E-04 0,55 0,55 0,88 0,70 0,74 0,75 0,81 0,72 130 8192 Focal Loss: gamma 3, alpha= (p=4, s=4, noise=1) 5,00E-04 0,55 0,55 0,87 0,71 0,75 0,74 0,81 0,72 130 8192 Focal Loss: gamma 3, alpha= (p=3, s=4, noise=1) 5,00E-04 0,55 0,55 0,89 0,72 0,74 0,73 0,81 0,81 90 4096 Focal Loss: gamma 3, alpha= (p=1, s=4, noise=2) 1,00E-04 0,55 0,55 0,95 0,78 0,63 0,63 0,75 0,69 50 4096 adaptive 1,00E-04 0,55 0,60 0,92 0,77 0,69 0,63 0,78 0,69 50 4096 adaptive 1,00E-04 0,55 0,60 0,88 0,70 0,70 0,63 0,78 0,66 35 1024 Cross Entropy 5,00E-04 0,55 0,55 0,92 0,81 0,72 0,62 0,81 0,70 70 4096 Focal Loss: gamma 3, alpha= (p=1, s=4, noise=2) 1,00E-04 0,55 0,55 0,95 0,77 0,62 0,62 0,75 0,69 100 8192 Focal Loss: gamma 3, alpha= (p=1, s=4, noise=2) 1,00E-04 0,55 0,55 0,95 0,77 0,62 0,62 0,75 0,72 70 8192 Focal Loss: gamma 3, alpha= (p=1, s=4, noise=2) 1,00E-04 0,55 0,55 0,95 0,76 0,72 0,61 0,75 0,68 50 8192 Focal Loss: gamma 3, alpha= (p=1, s=4, noise=2) 1,00E-04 0,55 0,55 0,94 0,76 0,62 0,60 0,75 0,67 50 2048 adaptive 1,00E-04 0,55 0,55 0,92 0,79 0,69 0,60 0,79 0,68 35 1024 Cross Entropy 1,00E-04 0,55 0,55 0,92 0,81 0,71 0,57 0,80 0,67 55 8192 adaptive 1,00E-04 0,55 0,55 0,91 0,79 0,69 0,52 0,78 0,63 57 2048 adaptive 1,00E-04 0,55 0,55 0,92 0,81 0,69 0,54 0,78 0,65 70 8192 Focal Loss: gamma 3, alpha= (p=1, s=4, noise=2) 5,00E-04 0,55 0,55 0,92 0,73 0,64 0,57 0,75 0,64 50 2048 adaptive 1,00E-03 0,55 0,60 0,93 0,85 0,70 0,49 0,80 0,62 35 2048 Cross Entropy 1,00E-04 0,55 0,55 0,92 0,81 0,71 0,56 0,80 0,66 90 2048 adaptive 1,00E-04 0,55 0,55 0,93 0,82 0,69 0,56 0,79 0,66 60 1024 Cross Entropy 1,00E-04 0,55 0,55 0,92 0,81 0,72 0,60 0,80 0,69 60 4096 adaptive 1,00E-04 0,55 0,55 0,92 0,71 0,69 0,74 0,78 0,72 TRANSFER LEARNING (PRETRAINED INSTANCE + AQ2009) E BS Loss Function LR T P T S PR-P PR-S R-P R-S F1 P F1 S 35 1.024 Cross Entropy 5,00E-04 0,55 0,55 0,92 0,74 0,70 0,71 0,80 0,72 35 1.024 Mse 5,00E-04 0,55 0,55 0,93 0,79 0,72 0,64 0,81 0,70 40 1.024 Cross Entropy 5,00E-04 0,55 0,55 0,93 0,80 0,72 0,61 0,81 0,69 45 2.048 Cross Entropy 1,00E-04 0,50 0,50 0,91 0,76 0,71 0,61 0,80 0,68 35 1.024 Cross Entropy 5,00E-04 0,55 0,55 0,93 0,81 0,72 0,60 0,81 0,69 60 2.048 Cross Entropy 1,00E-04 0,50 0,50 0,91 0,79 0,72 0,60 0,81 0,68 12 2.048 Cross Entropy 1,00E-03 0,50 0,50 0,92 0,78 0,72 0,59 0,81 0,67 40 2.048 Cross Entropy 1,00E-04 0,50 0,50 0,92 0,77 0,71 0,59 0,80 0,67 25 1.024 Cross Entropy 1,00E-04 0,55 0,55 0,92 0,78 0,70 0,58 0,80 0,66 35 2.048 Cross Entropy 1,00E-04 0,50 0,50 0,91 0,76 0,71 0,58 0,80 0,66 35 1.024 Cross Entropy 5,00E-04 0,55 0,55 0,92 0,82 0,73 0,57 0,82 0,67 10 2.048 Cross Entropy 1,00E-03 0,50 0,50 0,92 0,79 0,72 0,57 0,81 0,66 40 2.048 Cross Entropy 1,00E-04 0,55 0,55 0,93 0,77 0,69 0,57 0,79 0,66 40 2.048 Cross Entropy 5,00E-04 0,55 0,55 0,93 0,82 0,72 0,56 0,81 0,66 25 512 Cross Entropy 1,00E-04 0,55 0,55 0,93 0,80 0,70 0,55 0,80 0,65 25 2048 Cross Entropy 5,00E-04 0,55 0,55 0,93 0,80 0,71 0,55 0,80 0,65 25 1024 Cross Entropy 5,00E-04 0,55 0,55 0,93 0,82 0,72 0,54 0,81 0,65 40 2048 MSe 5,00E-04 0,55 0,55 0,93 0,83 0,72 0,54 0,81 0,66 40 1024 Cross Entropy 1,00E-04 0,55 0,55 0,92 0,81 0,71 0,54 0,80 0,65 40 2048 Cross Entropy 1,00E-03 0,60 0,60 0,93 0,84 0,71 0,51 0,81 0,64 25 4096 Cross Entropy 5,00E-04 0,55 0,55 0,93 0,81 0,69 0,50 0,79 0,61 20 2048 Cross Entropy 1,00E-04 0,50 0,50 0,90 0,77 0,70 0,48 0,79 0,59 40 2048 Cross Entropy 5,00E-04 0,60 0,60 0,94 0,84 0,70 0,47 0,80 0,61 40 2048 Cross Entropy 1,00E-04 0,60 0,60 0,93 0,82 0,68 0,47 0,79 0,60 50 512 Cross Entropy 1,00E-05 0,55 0,55 0,91 0,79 0,67 0,44 0,57 0,77 35 1024 Focal Loss: gamma 1.0 (1,1.25,3.0) 5,00E-04 0,55 0,55 0,96 0,88 0,63 0,43 0,76 0,58 35 1024 Cross Entropy 5,00E-04 0,70 0,70 0,95 0,87 0,67 0,41 0,79 0,41 100 4096 Focal Loss: gamma 1, alpha= (p=1, s=1.25, noise=3) 1,00E-04 0,55 0,55 0,95 0,88 0,63 0,37 0,76 0,52 40 2048 Cross Entropy 1,00E-04 0,70 0,70 0,95 0,86 0,64 0,36 0,76 0,51 40 2048 Cross Entropy 1,00E-03 0,70 0,70 0,96 0,88 0,65 0,35 0,77 0,52 35 512 Focal Loss: gamma 1.0 (1,1.25,5.0) 5,00E-04 0,55 0,55 0,96 0,92 0,59 0,28 0,73 0,43 35 512 Focal Loss: gamma 2.0 (1,1.25,5.0) 5,00E-04 0,55 0,55 0,97 0,92 0,57 0,26 0,72 0,41 40 2048 Cross Entropy 1,00E-05 0,50 0,50 0,92 0,63 0,40 0,25 0,56 0,35 35 512 Cross Entropy 1,00E-05 0,55 0,55 0,91 0,81 0,63 0,17 0,75 0,29 60 1024 Cross Entropy 5,00E-04 0,55 0,55 0,92 0,8 0,72 0,65 0,80 0,71 60 1014 adaptive focal loss (num: 3) 1,00E-04 0,55 0,55 0,92 0,67 0,65 0,75 0,76 0,71 80 2048 adaptive focal loss (num: 3) 1,00E-04 0,55 0,55 0,88 0,61 0,71 0,81 0,78 0,7 60 1024 adaptive focal loss (num: 2) 5,00E-04 0,55 0,55 0,92 0,77 0,69 0,68 0,79 0,72 ONLY AQ2009 DATASET Legend: E: Epoch; BS: Batch Size; LR: Learning Rate; T: Threshold; PR: Precision; R: Recall 60 Epochs;1024 Batch Size; 0.0005 Learning Rate Cross Entrophy 020 40 60 Epoch 10-1 Loss Train Loss Test Loss 0 01000 2000 3000 Counts -4 4 0 -4 4 0.2 1.0 0.6 0.2 1.0 0.6 Prob. Ampl. Prob. Ampl. 0 -10 10 0.2 1.0 0.6 Prob. 0.2 1.0 0.6 Prob. Ampl. 01000 2000 3000 Counts 0.2 1.0 0.6 Prob. 0.2 1.0 0.6 Prob. 0 -4 4 Ampl. 01000 2000 3000 Counts 80 Epochs; 2048 Batch Size; 0.0001 Learning Rate Adaptive Focal Loss 60 Epochs; 4096 Batch Size; 0.0001 Learning Rate Adaptive Focal Loss 020 40 60 Epoch 10-2 Loss Train Loss Test Loss 6x10-3 -2 2x10 -2 3x10 01000 2000 3000 Counts 01000 2000 3000 Counts 0.2 1.0 0.6 Prob. 0.2 1.0 0.6 Prob. 0 Ampl. 6 -6 0 -10 10 0.2 1.0 0.6 Prob. Ampl. 0.2 1.0 0.6 Prob. The best PhaseNet performance (with the best Loss Function curve and Hyperparmeters) were selected to produce a seismic catalog for the 2016-2017 Amatrice-Visso-Norcia seismic sequence. We tested the models producted on the 30th October 2016 and visualized some pickings after the seismic catalog production. 12 3 4 5 AI in Physical Sciences: Toward a Unified Approach for Physics and Geophysics 19 September 2025 Sala Baldini Origin Time: 2016-10-30 16:26:06 Station T1245 Training dataset: AQ2009. Epoch 41; 1024 Batch Size; 0.0005 Learning Rate; Cross Entrophy Time (UTC) 16:26:10 16:26:15 16:26:20 16:26:25 10000 -10000 0 Z (Counts) 20000 -20000 40000 -40000 0 N (Counts) 20000 -20000 40000 -40000 0 -60000 E (Counts) S P Precision P: 0.92 Precsion S: 0.80 Recall P: 0.72 Recall S: 0.65 F1 P: 0.80 F1 S: 0.71 200000 0 Z (Counts) N (Counts) E (Counts) Time (UTC) 14:17:25 S P Origin Time: 2016-10-30 14:17:23 Training Model: AQ2009. Epoch 80; 2048 Batch Size; 0.0001 Learning Rate; Focal Loss Station MZ104 14:17:30 14:17:35 14:17:40 200000 0 -200000 250000 0 -250000 Precision P: 0.88 Precision S: 0.61 Recall P: 0.81 Recall S: 0.78 F1 P: 0.78 F1 S: 0.72 E (Counts) Time (UTC) 06:48:50 Training Model: INSTANCE + AQ2009. Epoch 60; 1024 Batch Size; 0.0005 Learning Rate; Cross Entropy Station MZ24 500000 Precision P: 0.92 Precision S: 0.71 Recall P: 0.69 Recall S: 0.74 F1 P: 0.78 F1 S: 0.72 Origin Time: 2016-10-30 06:48:50 S P 06:48:55 06:49:00 06:49:05 06:49:10 2 0 1e6 1e6 1 0 -1 N (Counts) -500000 0 Z (Counts) IT’S OBVIOUS MEN!!! 20 Z (Counts) Time (UTC) 05:23:50 Origin Time: 2016-10-30 05:23:47 PhaseNet trained on INSTANCE dataset Station T1211 S P 05:23:55 05:24:00 05:24:05 -20 0 100 -100 0 N (Counts) 100 -100 0 E (Counts) 10000 Z (Counts) Time (UTC) 10:10:40 Origin Time: 2016-10-30 10:10:37 PhaseNet trained on Northern Califorinia Seismic Network Station T1215 10:10:45 10:10:50 10:10:55 -10000 0 10000 0 -10000 20000 N (Counts) 10000 E (Counts) -10000 0 42°00'N 42°15'N 42°30'N 42°45'N 43°00'N 43°15'N 43°30'N 010 km 12°45'E 13°15'E 13°45'E "Depth (km)" 050 100 150 42°00'N 42°15'N 42°30'N 42°45'N 43°00'N 43°15'N 43°30'N "Depth (km)" 010 km 12°45'E 13°15'E 13°45'E 050 100 150 42°00'N 42°15'N 42°30'N 42°45'N 43°00'N 43°15'N 43°30'N 12°45'E 13°15'E 13°45'E 010 km "Depth (km)" 050 100 150 12°45'E 13°15'E 13°45'E 42°00'N 42°15'N 42°30'N 42°45'N 43°00'N 43°15'N 43°30'N "Depth (km)" 050 100 150 010 km 42°00'N 42°15'N 42°30'N 42°45'N 43°00'N 43°15'N 43°30'N 010 km "Depth (km)" 050 100 150 The use of the Californian seismicity as training dataset also generates seismic events that are well localized in the seismicity zone involved in the sequence. However, seismicity is less clustered than when using INSTANCE. The geometry of the seismicity mapped is similar to that obtained using INSTANCE, but there is a “delay” in phase picking. By using this training model, PhaseNet confused the phases. Although seismicity is similar to that recorded using standard methods, a delay in picking is observed. The P wave is almost defined near the S wave. The reasons for a unreliable pickings predictions by the model may be caused by: - unappropriate hyperparameters configuration; - low-resolution dataset; We investigated the dataset choosen. EXAMPLE OF AQ2009 DATASET Different sampling rate: the tracks are sampled at 125Hz and PhaseNet samples at 100Hz (this explains the “delay” in picking). Some noisy and/or malfunctioning stations have false triggers which, when entered into the neural network, caused it to “learn badly” but well! Station: LG03; Magnitude: 0.49 Origin Time: 2009-09-04 23:32:56.330 Lat: 42.42°; Lon: 13.41°; Depht: 8.9 km GAP: 86.0°; RMS: 0.13 s 020 40 6010 30 50 70 Time (sec) 0 2000 6000 -2000 -4000 4000 8000 Amplitude Z N E Pick P 0 200 600 -200 -400 400 800 Amplitude 020 40 6010 30 50 Time (sec) 70 Station: AQU; Magnitude: 0.39 Origin Time: 2009-05-05 19:49:56.660 Lat: 42.42°; Lon: 13.36°; Depht: 3.3 km GAP: 69.0°; RMS: 0.10 s -600 -800 Z N E Pick P 020 40 6010 30 50 Time (sec) 70 Station: CERT; Magnitude: 0.59 Origin Time: 2009-11-13 13:06:44.560 Lat: 42.58°; Lon: 13.34°; Depht: 6.9 km GAP: 233.0°; RMS: 0.03 s 0 200 -200 -400 400 Amplitude Z N E Pick P 020 40 6010 30 50 Time (sec) 70 Station: CAMP; Magnitude: 0.67 Origin Time: 2009-05-29 12:18:27.620 Lat: 42.29°; Lon: 13.53°; Depht: 3.0 km GAP: 169.0°; RMS: 0.03 s Z N E Pick P 0 200 -200 Amplitude -100 100 300 020 40 6010 30 50 Time (sec) 70 0 200 600 -200 -400 400 Amplitude Station: CERT; Magnitude: 0.95 Origin Time: 2009-07-03 13:23:23.550 Lat: 42.32°; Lon: 13.39°; Depht: 7.3 km GAP: 62.0°; RMS: 0.08 s Z N E Pick P Z N E Pick P Station: LNSS; Magnitude: 0.82 Origin Time: 2009-06-03 08:56:27.390 Lat: 42.26°; Lon: 13.53°; Depht: 5.9 km GAP: 150.0°; RMS: 0.16 s 0 1000 3000 -1000 2000 4000 Amplitude 020 40 6010 30 50 Time (sec) 70 020 40 6010 30 50 Time (sec) 70 0 200 600 -200 -400 400 Amplitude Station: CESI; Magnitude: 0.49 Origin Time: 2009-05-02 23:22:52.110 Lat: 42.27°; Lon: 13.47°; Depht: 5.7 km GAP: 94.0°; RMS: 0.07 s Z N E Pick P 020 40 6010 30 50 Time (sec) 70 0 2500 -2500 Amplitude 5000 7500 -5000 -7500 -10000 Station: CESI; Magnitude: 0.57 Origin Time: 2009-07-12 08:37:30.900 Lat: 42.39°; Lon: 13.39°; Depht: 0.4 km GAP: 77.0°; RMS: 0.09 s Z N E Pick P 020 40 6010 30 50 Time (sec) 70 Z N E Pick P Pick S 0 200 -200 Amplitude 100 -100 -300 Station: LG06; Magnitude: 0.48 Origin Time: 2009-04-24 23:34:59.420 Lat: 42.47°; Lon: 13.38°; Depht: 7.9 km GAP: 33.0°; RMS: 0.10 s 020 40 6010 30 50 Time (sec) 70 Z N E Pick P Station: FIAM; Magnitude: 0.39 Origin Time: 2009-09-20 23:26:17.810 Lat: 42.57°; Lon: 13.19°; Depht: 4.0 km GAP: 118.0°; RMS: 0.02 s 0 250 -250 Amplitude 500 750 1000 -500 -750 -1000 020 40 6010 30 50 Time (sec) 70 Z N E Pick P Station: FIAM; Magnitude: 0.64 Origin Time: 2009-11-25 00:31:13:660 Lat: 42.59°; Lon: 13.17°; Depht: 4.7 km GAP: 138.0°; RMS: 0.04 s 0 200 -200 -400 400 Amplitude -600 -800 020 40 6010 30 50 Time (sec) 70 Z N E Pick P Station: FIAM; Magnitude: 0.69 Origin Time: 2009-04-27 23:08:55.660 Lat: 42.39°; Lon: 13.43°; Depht: 8.2 km GAP: 100.0°; RMS: 0.06 s 0 1000 3000 -1000 -2000 2000 Amplitude FUTURE WORK 6 Train PhaseNet with the “cleaned” dataset. To do so, we’ll: -select the best seismic events (according to locations quality criteria); -optimise data and code in order to have a coherent sampling frequency both in the data and in the metadata. Z N E Pick P (from metadata) Pick P (after resample) Pick S (from metadata) Pick S (after resample) 020 40 60 10 30 50 70 Time (sec) 0 250 750 -250 -500 500 1000 Amplitude -700 020 40 60 10 30 50 70 Time (sec) 0 1000 2000 3000 -1000 -2000 -3000 -4000 Amplitude 020 40 60 10 30 50 70 Time (sec) 0 1000 -1000 -500 500 Amplitude 020 40 60 10 30 50 70 Time (sec) 50 0 100 150 200 -50 -100 Amplitude RESAMPLE EXAMPLES - A grid search over hyperparameters guides towards an optimal solution, avoiding GPU-time consumption; - Selection of most appropriate loss function improves the training quality; - The quality of the training dataset is crucial to ensure the scientific reliability of the resulting seismic catalogues; -When training datasets rely on catalogues produced by automatic procedures of questionable reliability, there is a concrete risk of error propagation — from one generation of catalogues to the next. REFERENCES Excellent performance is observed when applying PhaseNet, trained on Italian earthquakes from 2000 to 2021, to waveforms from 2016. This is expected, as these data were included in the training set — the model had already “seen” them and likely memorized their patterns. Such evaluation does not reflect the model’s generalization capability. The use of the Focal Loss function improved the S-Recall value (about 20%). This is due to the introduction of the weighting of Pand Swaves compensating the different samples amout in the training dataset. Adaptive Focal Loss focusing parameter balancing factor probability Precision Recall S P