Clearance CLR-4159 · SAF156
SAFXGB
Safety & OperationsClearance sheet
XGBoost Model Predicts Airport Construction Risk at 92.7% Accuracy
A study of 412 construction events at a 45-million-passenger Chinese hub shows XGBoost predicts live-operations risk at 92.7% accuracy, with NOTAM delays over two hours the strongest single factor.
Read-back
- XGBoost achieved 92.7% accuracy (AUC 0.875) predicting risk events across 412 construction events at a three-runway, 45-million-passenger hub in eastern China, 2019–2024.
- NOTAM release delays exceeding two hours carried the highest odds ratio for risk events (4.28), pushing incidence to 47.1% — 2.3 times the rate for on-time releases.
- Risk incidence rose from 25.0% to 41.8% when flight density exceeded 35 movements per hour; crews with over eight years of experience cut risk by 58.9% relative to sub-three-year crews.

A machine learning model trained on 412 construction events at a 45-million-passenger hub airport in eastern China predicted safety risk events during live operations with 92.7% accuracy, according to a study published in Scientific Reports on 4 April 2026.
The single-author study, by Xian Yang of Guangxi Airport Management Group, drew on data spanning January 2019 through December 2024 at a three-runway airport operating 24/7 with an average of 1,200 daily flights. Of the 412 construction events analyzed — 216 runway maintenance jobs, 128 taxiway renovations and 68 apron repairs — 103 qualified as risk events, a 25.0% incidence rate. Risk events were defined as documented safety violations, near-misses recorded in the airport's safety management system, operational disruptions affecting flight schedules, or incidents formally reported to the Civil Aviation Administration of China.
The framework rests on 42 variables across six domains: personnel, environment, equipment, management, facilities and operations. Data came from five channels — aviation safety reporting systems, the airport's operational database, internal construction records, meteorological observations and NOTAM issuance logs — with a 98.5% completion rate.
Among five algorithms tested, XGBoost delivered the best performance: 92.7% accuracy, 85.7% precision, 85.7% recall, an F1 score of 85.7% and a cross-validated AUC of 0.875 ± 0.021. It outperformed Random Forest (86.6% accuracy) and Stacking (87.8%), while deep neural networks reached only 82.9% accuracy despite requiring the longest training time. The class imbalance — 103 risk events against 309 normal operations — was handled through SMOTE oversampling and a 3:1 class weight ratio in the loss function.
The analysis produced quantified risk thresholds with direct operational consequences. When flight density exceeded 35 movements per hour, risk event occurrence rose from the 25.0% average to 41.8%, an adjusted odds ratio of 3.76. Visibility below 3 km carried an odds ratio of 3.18. The most severe single factor was NOTAM release delay: postponements beyond two hours pushed risk incidence to 47.1%, an odds ratio of 4.28 — evidence, the study argues, that information communication failures outweigh weather in the risk hierarchy. Peak-hour construction (OR 2.64) and nighttime work between 22:00 and 06:00 (OR 1.87, with incidence of 31.2% versus 20.8% daytime) completed the top-tier factors.
Crew composition matters in measurable terms. Risk incidence was 38.2% when construction workers had fewer than three years of experience but dropped to 15.7% — a 58.9% relative reduction — for crews with more than eight years. Air traffic controller workload showed a similar pattern: continuous, non-rotated duty beyond four hours produced a 33.6% incidence rate against 20.3% under regular rotations.
SHAP interpretability analysis exposed nonlinear threshold behavior and interaction effects. Below 25 flights per hour, flight density contributed almost nothing to predicted risk; above 35, its average SHAP value jumped to +0.52. NOTAM delays beyond two hours averaged +0.61, and the delay effect amplified during peak hours — the combined SHAP contribution of high flight density plus NOTAM delay exceeded the sum of their individual effects, a clear positive interaction between operational stress variables.
Validation was extensive. On an independent test set of 82 events, the model correctly identified 18 of 21 risk events (recall 85.7%) with only three false positives among 61 normal operations (specificity 95.1%). Bootstrap resampling of 1,000 datasets produced a 95% confidence interval for accuracy of [88.6%, 94.7%]. Performance held across scenarios: 93.5% accuracy for daytime construction, 89.2% at night, 95.2% in good weather, 83.3% in severe conditions, and 90–92% across all three construction types. In backtracking against real-world cases, the model flagged seven of eight major risk events, missing one caused by a sudden hydraulic failure of construction equipment — a rare-event category with only four samples in training data.
The regulatory context frames the practical stakes. FAA Circular 150/5370-2G mandates systematic risk assessment and automated protection during airport construction, and ICAO Annex 19 identifies construction without airport closure as a key risk source requiring data-driven dynamic assessment. Hub airports cannot shut down without severe economic loss, so the efficiency-safety trade-off is structural: conservative constraints cut throughput, aggressive schedules invite incidents. The study positions its adjustable-threshold model as a way to manage that trade-off — a lower threshold of 0.3 raises recall to 92% for risk-averse hubs, while a higher threshold of 0.7 cuts the false positive rate to 2% for airports with limited management resources. The author notes inference runs in milliseconds versus days for conventional assessment methods.
Limitations are acknowledged. The model is trained on a single airport in a subtropical monsoon climate experiencing roughly 45 low-visibility days and 15–20 typhoon-affected days annually, plus about 70 concurrent construction projects per year. Deployment elsewhere would require local calibration for operational scale, climate and regulatory framework, and rare equipment-failure events remain weakly captured. The source code is publicly available on GitHub and archived in Zenodo.
Future work identified in the study includes multi-airport transfer learning, time-series features capturing construction progress, online learning to track regulatory changes, and integration of physiological monitoring of worker fatigue — a path toward a full human-machine collaborative decision platform embedded in airport safety management systems.
via nature.com (Original)
More from Priya Raman
Same bay
- FDR674ATSB Opens Preliminary Reviews Into Sydney Airport Incidents · September 30, 2026
- FDR965Qatar's aviation safety system rests on training and coordination · September 27, 2026
- FDR858EASA counts 116 aviation fatalities in 2025 annual safety review · September 30, 2026
- FDR884Flight Safety Foundation to Honor Two Safety Leaders at IASS 2026 · September 28, 2026
- FDR245CASA Finds No Immediate Safety Risk in Sydney Airport Incidents · September 30, 2026