01 Background
The safe, reliable, high-quality, and efficient operation of a major airport depends on substantial network infrastructure and the support of advanced information technology. To date, the airport has deployed 9 major platforms and over 100 business systems across more than a thousand servers, leveraging a range of emerging technologies including big data, IoT, cloud computing, Docker, and mobile applications.
As air traffic volume continues to grow year over year, the airport's IT resources have expanded significantly. Daily IT resource monitoring and management have encountered multiple challenges, including a lack of management measures, declining alert effectiveness, insufficient Monitoring and Alerting coverage, and a lack of sustained operations. Traditional monitoring tools can no longer meet management standards or business requirements.
Against this backdrop, the airport set out to build an IT O&M management platform encompassing automated O&M tools, a Configuration Management Center, an IT Service Management Center, and a Monitoring and Alerting Center. The goal was to integrate all O&M tools, further enhance business continuity and fault-handling efficiency, and meet the evolving demands of future IT O&M operations.
02 Objectives
To address alert management challenges, the airport introduced the CanWay BlueWhale Alert Management Platform to enhance alert lifecycle management capabilities, with the following objectives:
Deepen and expand Monitoring and Alerting coverage;
Improve alert effectiveness and alert processing efficiency;
Optimize management measures to achieve closed-loop alert management;
Use alert governance to drive optimization of monitoring strategies;
Ensure business continuity and reduce business incidents caused by growing operations.
03 Solution
Unified Alert Integration for Linked O&M Scenarios
The IT Operations Management Platform, built on a PaaS foundation, ingests alert data from various monitoring systems to establish unified alert standards and management. Leveraging the platform's CMDB, operations dashboards, ITSM, and standard O&M workflows, it achieves unified alert lifecycle management. Throughout this process, each component interacts closely with various O&M tools, not only significantly improving O&M efficiency, but also providing valuable data and in-depth analytical insights for system optimization and improvement.
Multiple Alert Sources in Parallel to Improve Alert Coverage
Monitoring coverage and completeness directly impact alert effectiveness and reliability. Building on existing monitoring tools such as Zabbix, out-of-band monitoring, and VCenter, the airport integrated the capabilities of the BlueKing Monitoring Platform to add service uptime testing, log keyword monitoring, customized business monitoring, component monitoring, and NTP monitoring — comprehensively enhancing alert coverage.
Multi-layer, multi-object, multi-metric, and multi-dimensional monitoring, combined with alert convergence and alert correlation, provides richer alert data and more comprehensive alert information. This assists in investigating and pinpointing the root causes of faults, enabling 24/7 information system operational assurance.
Dashboard Display of Business Health for Rapid Alert Response
To ensure normal business operations and timely resolution of O&M alerts, the airport's standard is to maintain an "alert-free screen."
The ECC (Emergency Command Center) duty team includes over ten service providers with a comprehensive duty system, along with well-established management protocols for alert response responsibilities. For rapid response, a large dashboard screen in the ECC duty room displays the health status of each business system, enabling duty personnel to quickly respond to and handle alerts based on health status.
When an alert is generated, the CMDB enriches it with business ownership information. Alerts are then aggregated by business dimension, and the dashboard displays the status of all business systems. A green status indicates no active alerts. When alerts occur, the system displays the corresponding health status based on alert severity, accompanied by an audible notification. The designated duty personnel then respond and resolve the issue. The goal of ECC duty O&M personnel is to resolve all alerts and achieve a fully green, healthy screen.

Alert Self-Healing for Rapid Alert Recovery
For alerts with known and recurring resolution procedures, waiting for manual response and handling extends alert resolution time. Through alert self-healing, corresponding remediation actions are automatically triggered to restore normal operations or mitigate potential risks.
During earlier O&M duty operations, the airport had already accumulated a set of routine and standardized alert handling procedures. For example, alerts caused by process errors in certain non-critical business systems are matched with resolution strategies based on alert information, and automated remediation is executed to restart the process — rapidly restoring the system to normal operation.
Alert Governance to Improve Alert Effectiveness and Efficiency
As monitoring coverage expands, the volume of alerts increases, making effective alert convergence through noise reduction critically important. The airport uses a tagging approach to categorize alert dispositions and conducts periodic alert reviews to summarize alert handling methods, false alerts, and unreasonable alert strategies. Based on these reviews, monitoring strategies, alert convergence strategies, and alert handling strategies are optimized and fine-tuned, progressively improving alert effectiveness. Currently, the alert hit rate has reached 75%.
Additionally, through alert report analysis, the airport evaluates alert handling efficiency by vendor and business system, and leverages performance metrics to further improve alert processing efficiency.
04 Results
Timely Alert Detection
Alert notification timeliness improved by 150%, with alerts accurately delivered within 1 minute of occurrence. This has reduced business impact duration, improved business stability, and enhanced user satisfaction.
Effective Assistance in Root Cause Analysis (RCA)
Through increased alert coverage and effective alert convergence, combined with alert topology views and correlated alert information, the platform enables correlation analysis upon alert occurrence — allowing faster identification of root causes and key issues. This accelerates alert resolution and reduces the scope and duration of alert impact.
Ensuring Business Continuity
Integrated with the monitoring platform, the airport enhanced its existing monitoring systems with service uptime testing, log keyword monitoring, customized business monitoring, component monitoring, and NTP monitoring, achieving broader and more granular monitoring capabilities. Alert coverage reached 90%. Localized alerts are detected before business disruptions occur, and through alert correlation analysis, related business systems and associated alerts are identified and resolved promptly — preventing cascading failures caused by accumulated faults. This maximizes business continuity operations.
05 Product Applicability
The CanWay BlueWhale Alert Management Platform is designed for alert lifecycle management. By aligning with enterprise organizational structures and business needs, it delivers solutions tailored to improve alert coverage and business continuity. It is suitable for enterprises with the following requirements:
Multiple monitoring systems without a centralized alert management system, requiring unified alert coverage for streamlined management;
Large duty teams or multiple outsourcing vendors experiencing untimely or inadequate notifications;
Numerous systems and O&M personnel, requiring alert correlation analysis for rapid issue and personnel identification;
Alert governance needs, optimizing monitoring metrics and O&M frameworks through systematic alert governance;
High requirements for business continuity and user satisfaction.

















