فصلنامه علمی- پژوهشی اطلاعات جغرافیایی « سپهر»

فصلنامه علمی- پژوهشی اطلاعات جغرافیایی « سپهر»

تحلیل مقایسه‌ای شبکه عصبی عمیق و ماشین بردار پشتیبان بهینه‌شده برای برآورد مکانی پسلرزه با تأکید بر کالیبراسیون احتمال و کمّی‌سازی عدم‌قطعیت

نوع مقاله : مقاله پژوهشی

نویسنده
دانشیار گروه مهندسی نقشه برداری، دانشکده عمران، دانشگاه تبریز، تبریز، ایران
چکیده
برآورد دقیق مکانی پسلرزه‌ها یکی از چالش‌های اساسی در زلزله‌شناسی آماری است. در این مطالعه، دو الگوریتم یادگیری ماشین شامل شبکه عصبی عمیق (DeepMLP) و ماشین بردار پشتیبان با هسته تابع پایه شعاعی (SVM-RBF)برای طبقه‌بندی دودویی سلول‌های شبکه ۵ کیلومتری به دو کلاس «دارای پسلرزه» و «بدون پسلرزه» مقایسه شدند. مجموعه داده شامل ۲۰۲٬۶۴۴ نمونه از ۱۹۹ رخداد زمین‌لرزه بود که با تقسیم در سطح رخداد (۱۴۲ رخداد آموزش، ۵۷ رخداد آزمون) از نشتی داده جلوگیری شد. SVM با پنج اصلاح کلیدی بهبود یافت: وزن‌دهی متعادل کلاس‌ها، جایگزینی معیار F1 با AUC، گسترش فضای جستجوی γ به بازه (1، 0.0001)، کالیبراسیون احتمال، و کاهش زیرنمونه به ۱۰٬۰۰۰. SVM بهبودیافته نه‌تنها توانایی رتبه‌بندی بهتری نشان داد (DeLong ΔAUC = +0.0011، p = 0.0011) بلکه کالیبراسیون احتمال آن نیز بهتر بود (Brier: 0.166 در برابر 0.172، بازه اطمینان ۹۵٪ بدون همپوشانی). با این حال، تصمیم‌های دودویی دو مدل در آستانه بهینه تفاوت معناداری نداشتند (McNemar p = 0.57). ارزیابی عدم‌قطعیت با AUSE سیگنال متوسطی (0.11~) برای هر دو مدل نشان داد، اما فیلتر کردن 50% نمونه‌های با عدم‌قطعیت بالا، دقت را از 0.77 به 0.87 (10%+) رساند. منحنی‌های یادگیری حاکی از اشباع هر دو مدل در حدود ۵٬۰۰۰ تا ۱۰٬۰۰۰ نمونه بود؛ یعنی عامل محدودکننده، نمایش ویژگی است نه حجم داده. از نظر کارایی،SVM بهبودیافته حدود ۲۰ برابر سریع‌تر از DeepMLP آموزش می‌بیند (۱۹ در برابر ۳۷۶ ثانیه)، در حالی که DeepMLP در استنتاج ۵۶ برابر سریع‌تر عمل می‌کند. این یافته‌ها به مدیران بحران در انتخاب مدل مناسب بر اساس اولویت‌های عملیاتی (سرعت در برابر دقت مثبت) کمک می‌کند.
کلیدواژه‌ها
موضوعات

عنوان مقاله English

A comparative analysis of deep neural networks and optimized support vector machines for spatial aftershock estimation with emphasis on probability calibration and uncertainty quantification

نویسنده English

Asghar Rastbood
Associate professor, Surveying engineering department, Civil engineering faculty, University of Tabriz, Tabriz, Iran
چکیده English

Extended Abstract
Introduction
Accurate spatial aftershock forecasting remains a major challenge in seismology. While empirical laws describe temporal decay and magnitude distribution, explaining spatial patterns of aftershocks has proven far more difficult. The Coulomb failure stress change (ΔCFS) has been the dominant physical criterion; however, independent tests report area under the ROC curve (AUC) values as low as 0.583, indicating limited predictive power (Zhao et al., 2022). ΔCFS suffers from high sensitivity to unknown parameters (fault orientation, friction coefficient, pore pressure), contradictory empirical evidence from major earthquakes, and substantial uncertainties in slip inversion and secondary stress transfer (Hardebeck et al., 1998; Mallman & Zoback, 2007; Hainzl et al., 2009). Furthermore, Meade et al. (2017) demonstrated that no single scalar stress metric consistently outperforms others in explaining aftershock patterns. These limitations necessitate data-driven approaches.
Machine learning has opened new possibilities. DeVries et al. (2018) proposed a six-layer deep neural network (DeepMLP) achieving AUC of 0.849, substantially outperforming ΔCFS. However, DeepMLP requires extensive hyperparameter tuning, long training times, and its "black box" nature complicates geophysical interpretation. Subsequent research has challenged the necessity of complex architectures: Mignan and Broccardo (2020) showed that logistic regression with only two parameters achieves performance comparable to the 13,451-parameter DeVries network, while Liu et al. (2024) achieved 93.44% accuracy with hybrid XGBoost-LightGBM models. These findings suggest simpler, more interpretable models, when properly optimized, can be competitive.
Support Vector Machines (SVMs) with RBF kernels offer a compelling alternative. Grounded in statistical learning theory (Cortes & Vapnik, 1995; Vapnik, 1995), SVMs perform well on structured tabular data—the format of stress tensor inputs. The RBF kernel enables modeling complex non-linear relationships via implicit mapping to infinite-dimensional space, while convex optimization guarantees global optimality. However, SVM application to large-scale problems is limited by computational complexity—O(n²) to O(n³)—and sensitivity to the regularization parameter C and kernel width γ. This study conducts a systematic comparison between DeepMLP and an optimized RBF-SVM for spatial aftershock prediction, evaluating performance, probability calibration, uncertainty quantification, and computational trade-offs to provide practical model selection guidance.
Materials & Methods
This study uses the DeVries et al. (2018) dataset: 199 finite-fault rupture models (SRCMOD) with 5×5×5 km grid cells. Six static stress-change tensor components were calculated using Okada's (1992) half-space model, and aftershocks (1s–1yr) were extracted from the ISC catalogue. The balanced dataset (202,644 samples) has 12 features (absolute values ​​of the stress components and their negative absolute values, normalized by 10⁻⁶). Data were split by distinct mainshocks (142 training, 57 testing), with 10% stratified validation to prevent data leakage.
DeepMLP follows the DeVries18 architecture: six hidden layers (50 neurons each, tanh, dropout 0.5), Adam optimizer, binary cross-entropy, 200 epochs, batch size 512, and class_weight='balanced'. The SVM was rescued with five key modifications: (1) class_weight='balanced' to address class imbalance; (2) scoring='roc_auc' instead of F1; (3) gamma search space extended to [1e-4, 1]; (4) probability calibration using CalibratedClassifierCV (isotonic, cv=5, ensemble=True); and (5) subsampling to 10,000 samples to reduce training time while preserving support vectors. Hyperparameters were optimized via Grid Search, Random Search, and Bayesian Optimization (Optuna TPE). Decision thresholds were optimized on validation by maximizing F1. Evaluation included AUC, AP, Brier score, NLL, DeLong test for correlated AUCs, McNemar test, AUSE for uncertainty quantification, and learning curves.
Results & Discussion
The rescued SVM achieved Test AUC = 0.8443, statistically significantly better than DeepMLP's 0.8432 (DeLong ΔAUC = +0.0011, z = −3.27, p = 0.0011). SVM also outperformed DeepMLP on Brier score (0.1659 vs. 0.1720) and NLL (0.5100 vs. 0.5253), with the two Brier 95% bootstrap CIs completely disjoint ([0.1641, 0.1675] vs. [0.1704, 0.1735]) — strong evidence of superior calibration. However, McNemar's test on hard classifications at optimal thresholds returned p = 0.574, indicating the two models are statistically equivalent in binary decision. This apparent paradox between DeLong and McNemar reveals that SVM's advantage lies in ranking ability and calibration rather than in the final decision boundary.
Uncertainty quantification via AUSE showed both models have moderate UQ quality (MLP: 0.1104, SVM: 0.1092), with negligible difference. Selective-prediction curves demonstrated that filtering the 50% most-uncertain samples improves accuracy from 0.77 to 0.87 (+10%) for both models — a practically useful gain. Learning curves saturated at ~5,000–10,000 training samples for both models; increasing data 30-fold did not improve performance by more than 0.3%, indicating that the limiting factor is feature representation, not sample size.
Computationally, the rescued SVM was ~20× faster to train (19 s vs. 376 s), while DeepMLP was ~56× faster at inference (2.1 s vs. 118 s). These differences make SVM attractive for rapid training scenarios, whereas DeepMLP is preferable for high-throughput real-time inference.
The DeepMLP's deep architecture with multiple hidden layers enables hierarchical feature extraction, capturing complex nonlinear relationships essential for aftershock pattern identification. The RBF SVM, despite its infinite-dimensional mapping, cannot achieve the same level of complexity due to its local kernel nature. SVM's advantages in ranking and calibration stem from proper optimization and isotonic calibration, which correct its raw decision scores into well-calibrated probabilities. The learning curves suggest that both models have effectively reached the ceiling of the current feature set; further improvement requires feature engineering rather than more data.
Conclusion
This study compared DeepMLP with a rescued RBF-SVM for spatial aftershock forecasting. The rescued SVM significantly outperformed DeepMLP in ranking ability (DeLong p = 0.0011) and probability calibration (Brier CI disjoint), while the two models were statistically equivalent in hard classification (McNemar p = 0.57). Both models saturated at ~5,000–10,000 samples, indicating that the current 12 stress features capture most of the learnable signal. Uncertainty quantification was moderate for both models (AUSE ≈ 0.11), but filtering high-uncertainty samples improved accuracy by 10%. Operationally, SVM is preferred when training speed or probability ranking matters, while DeepMLP is preferable for fast inference on large datasets. Future work should focus on feature engineering, hybrid ensembles, and integration of dynamic stress features.

کلیدواژه‌ها English

Aftershock forecasting
Support vector machine
Deep neural network
Probability calibration
Uncertainty quantification
DeLong test
AUSE

مقالات آماده انتشار، پذیرفته شده
انتشار آنلاین از 18 مهر 1405