Run-Level Fault Detection and SHAP-Based Diagnosis of Persistent Classification Difficulty in the Tennessee Eastman Process
PROCESSES, cilt.14, sa.16, 2026 (SCI-Expanded, Scopus)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 14 Sayı: 16
- Basım Tarihi: 2026
- Doi Numarası: 10.3390/pr14162569
- Dergi Adı: PROCESSES
- Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Compendex, Natural Science Collection (ProQuest), Biological Science Database (ProQuest), Materials Science & Engineering Collection (ProQuest), Technology Collection (ProQuest)
- Gazi Üniversitesi Adresli: Evet
Özet
Reliable fault detection in nonlinear process systems requires both accurate classification and interpretable analysis of faults with weak or near-normal signatures. This study develops an explainable, data-driven, and temporally informed multiclass framework for the Tennessee Eastman Process (TEP), retaining all 21 operating conditions (20 fault types and the normal operating condition). Six sliding-window statistics were extracted from 52 process variables and classified using Extreme Gradient Boosting (XGBoost). Performance was evaluated at both sample and run levels through fault-specific window analysis, an a priori validation-driven hierarchical decomposition, SHapley Additive exPlanations (SHAP), a Relative Sensitivity Index (RSI) based on detection-delay sensitivity, and computational benchmarking. Aggregating sample-level predictions (macro F1 = 0.8823) into run-level decisions via majority voting improved performance substantially (macro F1 = 0.9515). Across five seeded repetitions, the mean run-level macro F1-score was 0.9522 +/- 0.0035 (95% CI: +/- 0.0044 ). Under the primary evaluation, 18 of 21 classes achieved F1 >= 0.97 . Fault 3 benefited strongly from extended temporal context, whereas Normal operation, Fault 9, and Fault 15 retained a structured but asymmetric confusion pattern dominated by Normal-Fault 9 errors. SHAP identified model-specific attribution patterns associated mainly with cooling-water-related variability features, while RSI indicated greater prediction-stream sensitivity for the historically difficult faults. Feature extraction and inference required approximately 6.2 ms on CPU and 35.3 ms on GPU, negligible relative to the 180 s sampling interval. These findings indicate that temporal context, run-level aggregation, and explainability can jointly support accurate data-driven fault diagnosis while revealing persistent fault-specific ambiguity.