نوع مقاله : مقاله پژوهشی
عنوان مقاله English
نویسنده English
Extended Abstract
Introduction
Accurate spatial aftershock forecasting remains a major challenge in seismology. While empirical laws describe temporal decay and magnitude distribution, explaining spatial patterns of aftershocks has proven far more difficult. The Coulomb failure stress change (ΔCFS) has been the dominant physical criterion; however, independent tests report area under the ROC curve (AUC) values as low as 0.583, indicating limited predictive power (Zhao et al., 2022). ΔCFS suffers from high sensitivity to unknown parameters (fault orientation, friction coefficient, pore pressure), contradictory empirical evidence from major earthquakes, and substantial uncertainties in slip inversion and secondary stress transfer (Hardebeck et al., 1998; Mallman & Zoback, 2007; Hainzl et al., 2009). Furthermore, Meade et al. (2017) demonstrated that no single scalar stress metric consistently outperforms others in explaining aftershock patterns. These limitations necessitate data-driven approaches.
Machine learning has opened new possibilities. DeVries et al. (2018) proposed a six-layer deep neural network (DeepMLP) achieving AUC of 0.849, substantially outperforming ΔCFS. However, DeepMLP requires extensive hyperparameter tuning, long training times, and its "black box" nature complicates geophysical interpretation. Subsequent research has challenged the necessity of complex architectures: Mignan and Broccardo (2020) showed that logistic regression with only two parameters achieves performance comparable to the 13,451-parameter DeVries network, while Liu et al. (2024) achieved 93.44% accuracy with hybrid XGBoost-LightGBM models. These findings suggest simpler, more interpretable models, when properly optimized, can be competitive.
Support Vector Machines (SVMs) with RBF kernels offer a compelling alternative. Grounded in statistical learning theory (Cortes & Vapnik, 1995; Vapnik, 1995), SVMs perform well on structured tabular data—the format of stress tensor inputs. The RBF kernel enables modeling complex non-linear relationships via implicit mapping to infinite-dimensional space, while convex optimization guarantees global optimality. However, SVM application to large-scale problems is limited by computational complexity—O(n²) to O(n³)—and sensitivity to the regularization parameter C and kernel width γ. This study conducts a systematic comparison between DeepMLP and an optimized RBF-SVM for spatial aftershock prediction, evaluating performance, probability calibration, uncertainty quantification, and computational trade-offs to provide practical model selection guidance.
Materials & Methods
This study uses the DeVries et al. (2018) dataset: 199 finite-fault rupture models (SRCMOD) with 5×5×5 km grid cells. Six static stress-change tensor components were calculated using Okada's (1992) half-space model, and aftershocks (1s–1yr) were extracted from the ISC catalogue. The balanced dataset (202,644 samples) has 12 features (absolute values of the stress components and their negative absolute values, normalized by 10⁻⁶). Data were split by distinct mainshocks (142 training, 57 testing), with 10% stratified validation to prevent data leakage.
DeepMLP follows the DeVries18 architecture: six hidden layers (50 neurons each, tanh, dropout 0.5), Adam optimizer, binary cross-entropy, 200 epochs, batch size 512, and class_weight='balanced'. The SVM was rescued with five key modifications: (1) class_weight='balanced' to address class imbalance; (2) scoring='roc_auc' instead of F1; (3) gamma search space extended to [1e-4, 1]; (4) probability calibration using CalibratedClassifierCV (isotonic, cv=5, ensemble=True); and (5) subsampling to 10,000 samples to reduce training time while preserving support vectors. Hyperparameters were optimized via Grid Search, Random Search, and Bayesian Optimization (Optuna TPE). Decision thresholds were optimized on validation by maximizing F1. Evaluation included AUC, AP, Brier score, NLL, DeLong test for correlated AUCs, McNemar test, AUSE for uncertainty quantification, and learning curves.
Results & Discussion
The rescued SVM achieved Test AUC = 0.8443, statistically significantly better than DeepMLP's 0.8432 (DeLong ΔAUC = +0.0011, z = −3.27, p = 0.0011). SVM also outperformed DeepMLP on Brier score (0.1659 vs. 0.1720) and NLL (0.5100 vs. 0.5253), with the two Brier 95% bootstrap CIs completely disjoint ([0.1641, 0.1675] vs. [0.1704, 0.1735]) — strong evidence of superior calibration. However, McNemar's test on hard classifications at optimal thresholds returned p = 0.574, indicating the two models are statistically equivalent in binary decision. This apparent paradox between DeLong and McNemar reveals that SVM's advantage lies in ranking ability and calibration rather than in the final decision boundary.
Uncertainty quantification via AUSE showed both models have moderate UQ quality (MLP: 0.1104, SVM: 0.1092), with negligible difference. Selective-prediction curves demonstrated that filtering the 50% most-uncertain samples improves accuracy from 0.77 to 0.87 (+10%) for both models — a practically useful gain. Learning curves saturated at ~5,000–10,000 training samples for both models; increasing data 30-fold did not improve performance by more than 0.3%, indicating that the limiting factor is feature representation, not sample size.
Computationally, the rescued SVM was ~20× faster to train (19 s vs. 376 s), while DeepMLP was ~56× faster at inference (2.1 s vs. 118 s). These differences make SVM attractive for rapid training scenarios, whereas DeepMLP is preferable for high-throughput real-time inference.
The DeepMLP's deep architecture with multiple hidden layers enables hierarchical feature extraction, capturing complex nonlinear relationships essential for aftershock pattern identification. The RBF SVM, despite its infinite-dimensional mapping, cannot achieve the same level of complexity due to its local kernel nature. SVM's advantages in ranking and calibration stem from proper optimization and isotonic calibration, which correct its raw decision scores into well-calibrated probabilities. The learning curves suggest that both models have effectively reached the ceiling of the current feature set; further improvement requires feature engineering rather than more data.
Conclusion
This study compared DeepMLP with a rescued RBF-SVM for spatial aftershock forecasting. The rescued SVM significantly outperformed DeepMLP in ranking ability (DeLong p = 0.0011) and probability calibration (Brier CI disjoint), while the two models were statistically equivalent in hard classification (McNemar p = 0.57). Both models saturated at ~5,000–10,000 samples, indicating that the current 12 stress features capture most of the learnable signal. Uncertainty quantification was moderate for both models (AUSE ≈ 0.11), but filtering high-uncertainty samples improved accuracy by 10%. Operationally, SVM is preferred when training speed or probability ranking matters, while DeepMLP is preferable for fast inference on large datasets. Future work should focus on feature engineering, hybrid ensembles, and integration of dynamic stress features.
کلیدواژهها English