01 The Necessity of Alerts in the IT O&M System
1. Challenges in Enterprise Monitoring and Alerting Management
Alert management is a critical component of enterprise IT O&M management. It enables organizations to monitor and diagnose the status of business systems in real time and promptly identify potential failures or anomalies. However, in practice, enterprise alert management faces several challenges, primarily in the following areas:
Scattered and Non-Standardized Alerts
Across multiple isolated monitoring systems, various alerts exist without a unified format or content standard. There is a lack of centralized management tools, and alert information is often incomplete with poor readability.
Delayed Alert Notifications
For urgent failures or anomalies, if alert information cannot be delivered to relevant personnel in a timely manner, it may lead to irreversible consequences. Enterprise alert management systems must ensure the timeliness and reliability of alert notifications.
Excessive Alert Noise
Different monitoring systems use manually configured fixed thresholds with inconsistent standards, and a single failure may trigger alerts across different systems, resulting in a large volume of false positives, missed alerts, and duplicate alerts.
No Global View
It is impossible to intuitively understand the overall alert status and the scope of correlated impact across application systems and object models.
Lack of Tool Integration
Alert handling involves excessive manual intervention, minimal automated processing, low alert workflow efficiency, and a lack of tracking and closed-loop processes.
Lack of Accumulated IT O&M Experience
For similar or highly complex alerts, less experienced IT O&M personnel must spend significant time troubleshooting, resulting in low alert handling efficiency.
2. Alert Management Pain Points by Role
Within an enterprise, different roles have different priorities regarding alerts and consequently experience different pain points in alert management. The following roles may encounter alert management pain points:
3. The Value of Alert Management
Alert management is an indispensable part of ensuring system stability. Its value lies in helping organizations discover and resolve issues promptly, safeguarding system stability and user experience. By significantly improving response speed, effectively reducing human errors, and optimizing system maintenance workflows, alerts play a vital role in daily IT O&M and management.
Real-Time Monitoring and Timely Detection
By configuring alert rules and metrics, the operational status of various systems, networks, and applications is monitored. Once an alert rule is triggered, the responsible personnel are notified immediately, enabling them to intervene and address the issue in a timely manner.
Rapid Problem Localization, Reduced Troubleshooting Time
Through clear metrics, detailed data, and intelligent knowledge base recommendations provided in alert information, problems can be quickly localized and effective countermeasures can be taken, shortening fault resolution time.
Automated Handling, Improved Efficiency
Through automated alerting and handling, the time and cost previously required for manual maintenance can be significantly reduced. Alerts can easily and automatically trigger emergency response workflows, reducing manual intervention and errors.
Global Data Analysis for Alert Governance
Alerts provide real-time data and statistical information, serving as the basis for business decisions or performance optimization. Through systematic organization and in-depth analysis of alert information, organizations can not only more effectively support management in making precise business decisions, but also discover potential growth opportunities and development prospects.
02 Alert System Implementation Path
1. Alert System Maturity Model
Alert system maturity refers to the assessment of how mature an enterprise or organization is in implementing an effective alert system. An alert system is one that monitors critical business operations and systems — including systems, applications, and devices — and issues alerts, helping users discover problems and handle them swiftly. Below is the industry-standard classification of alert system maturity:
Currently, most enterprises operate at the L2–L4 levels of alert management, having completed basic alert lifecycle management. Higher levels enable more efficient alert closed-loop operations. Alert system maturity must be built progressively from low to high — only after lower-maturity alert management has been completed can higher-level optimization be built upon that foundation.
2. Implementation Approach
Achieving automated alert management or alert governance optimization requires forming a closed loop from standardized alert onboarding → alert handling process → post-incident review and knowledge accumulation. Implementing this closed-loop management scenario involves people, tools, and management standards. Combining these elements, the final implementation produces the following closed-loop path:
Standardized Alert Onboarding: Based on the requirements of alert information standardization and scenario-driven consumption, alerts from various monitoring systems are unified through plugin development, alert enrichment, and other methods, standardizing alert data and formats. This includes: alert severity definitions, alert metrics, alert objects, etc.
Alert Handling Process: Based on alert information, problem investigation and resolution are carried out, including identifying root causes and taking appropriate measures to resolve issues. This encompasses alert enrichment, alert convergence, alert analysis, alert handling, and other stages.
Post-Incident Review and Knowledge Accumulation: A review of alert activity over a given period is conducted. Based on the results of alert post-mortems, alert handling processes and rules are optimized and improved to enhance alert accuracy and handling efficiency.
3. Implementation Path
Based on the implementation approach, the alert implementation is divided into the following key steps: Standardized Alert Onboarding, Alert Convergence Standards, Alert Handling Standards, and Alert Post-Mortem Governance.
① Standardized Alert Onboarding
Based on the requirements of alert information standardization and scenario-driven consumption, alerts from various monitoring systems are unified through plugin development, alert enrichment, and other methods, standardizing alert data and formats.
By consolidating all monitoring tool alert events through a unified alert center and standardizing all alert fields, alerts must conform to the following onboarding specification template:
② Alert Convergence Standards
Alert convergence, as a critical task during the alert handling phase, filters, merges, and simplifies repetitive alert information generated multiple times to reduce alert volume and improve alert handling efficiency and accuracy. Establishing alert convergence standards helps alleviate the burden on IT O&M personnel and prevents the chaos and delays caused by alert flooding. Below are key points for defining alert convergence standards:
Alert Suppression
For monitoring system alert sources lacking built-in convergence capabilities, on-duty personnel configure alert suppression strategies to effectively prevent alert storms.
Common Alert Suppression Scenario — Anti-Jitter Suppression Strategy:
Alert events sporadically generated by fluctuating metrics
Fluctuating metrics include: CPU utilization, memory utilization, disk I/O, network traffic, etc.
Ineffective alerts caused by metric fluctuations can be suppressed using "N occurrences within X minutes"; configured based on the probability of metric fluctuations.
Alert Shielding
For IT O&M change windows, on-duty personnel configure alert shielding strategies to prevent false alerts. Alert shielding is generally divided into two methods — time-based shielding and dependency-based shielding — with the following common use cases:
Time-Based Shielding Strategy: For alerts generated by known events that require no attention.
Common scenarios: During system maintenance periods or change windows; a strategy can be configured to shield all alerts for a given system during a specified time window.
Dependency-Based Shielding Strategy: For correlated alert events caused by dependency relationships.
Common scenarios include:
Components installed and running on a host;
Host disks mounted on storage volumes provided by a storage system;
Virtual machines running on a host or host cluster;
Hosts and devices connected to the network via switches;
Internal service call dependencies within an application, e.g., front-end applications calling back-end services or databases;
External service call dependencies, e.g., a Taobao application calling Alipay's payment service. If Object A depends on Object B, a strategy can be configured so that when Object B generates a specific alert, the corresponding alert for Object A is automatically shielded.
③ Alert Handling Standards
The alert handling phase primarily involves event acknowledgment and documentation, ensuring that problems can be quickly and accurately identified, analyzed, and resolved. Key activities in the alert handling phase include:
Alert Dispatching
For valid alert events, on-duty personnel configure alert dispatching strategies to route alerts matching specified time and rule criteria to designated personnel and groups for handling.
Alert Self-Healing
For common alerts with well-defined handling procedures, alert self-healing strategies can be configured:
Log files too large — automatic log cleanup;
Disk space full — automatic cleanup of files in specified directories;
Service anomalies — automatic process restart;
Load balancer node anomalies — automatic removal of the anomalous node from the load balancer pool;
Website access anomalies — automatic DNS record modification to redirect to a backup address;
Cluster synchronization anomalies — automatic synchronization triggered;
Time synchronization anomalies — automatic time synchronization triggered;
Host failure — standby machine automatically brought online.
Automatic Ticket Creation
For complex alert handling that requires manual intervention, alerts can be routed to the appropriate team or experts via a ticketing system for processing, with a complete handling record maintained. Common scenarios:
Issues requiring escalation to second-line, third-line support, or external vendors;
Alerts escalated to incidents requiring ticket-based incident management;
Alerts transferred to other departments, requiring workflow routing via tickets.
Alert Post-Mortem Governance
Through alert operations analysis, metrics such as alert distribution, MTTA and MTTR for alert handling, and alert closure rates are tracked to continuously optimize alert strategies and management processes. Additionally, a knowledge base is built from historical alert resolution records to provide handling guidance for similar future issues.
4. Success Factors
Alert management requires coordination across people, systems, and management standards — all of which are complex. These factors influence whether alert management implementation can succeed. Several key success factors are as follows:
Alert Data Standardization
For alert source data with multiple formats and definitions, a unified standard for onboarding into the alert center must be established to ensure data consistency. Format standards include alert structure definitions (alert information, object information, and extended information), alert severity definitions, alert event ID definitions, and other critical field definitions.
Standardized CMDB Implementation
Based on CMDB data, accurate alert enrichment information can be matched to facilitate efficient dispatching and problem localization. Correct correlation relationships to alert objects can be established, improving the accuracy of alert correlation analysis.
Standardized Alert Management Practices
Formalized alert management policies must be established, with alert handling strategies configured according to these policies to enable efficient processing. Rules for different alert types and severity levels must be standardized. Developing alert management standards is a complex process encompassing personnel, role responsibilities, alert severity definitions, alert handling SLAs, alert handling strategies, and other intricate specifications. Once comprehensive management standards are defined, they are implemented through the alert center's feature set.
Continuous Process Improvement Through Operations Analysis
Through alert handling and workflow records, alert response time and resolution time are measured and assessed to promote rapid alert response and resolution. Optimization adjustments are made based on assessment results. Alert governance is a continuous process of improvement and optimization — in response to architectural changes or issues exposed by operational incidents, monitoring items and alert strategies must be continuously refined to enhance timeliness and effectiveness.
03 Driving Alert System Construction with Products
Building an enterprise IT O&M fault closed-loop alert system hinges on the equal emphasis on standardized processes and quality products. Processes ensure the steady construction of the alert system to effectively handle various alerts and safeguard system stability. Robust product support serves as an accelerator — not only strengthening system capabilities but also driving the overall evolution of the IT O&M system, significantly improving IT O&M response speed and efficiency, and enhancing system reliability.
CanWay BlueWhale Alert Management Platform is the ideal platform for achieving this goal. By combining the alert implementation path with this platform, an efficient and reliable alert management system can be built. The system's automated processes, tightly integrated with manual intervention, not only improve the speed and accuracy of alert handling but also provide robust support for enterprise IT O&M management, ensuring business continuity and stability.
1. Product Introduction
CanWay BlueWhale Alert Management Platform is a full alert lifecycle management tool that effortlessly consolidates alert information from various monitoring systems. It enables alert enrichment, suppression, shielding, handling, dispatching, and analysis, helping IT O&M teams manage alert events in a unified closed-loop manner. By freeing up human resources while dramatically improving fault handling efficiency, it better safeguards business stability.
Through CanWay BlueWhale Alert Management Platform, the lifecycle workflow of alert source onboarding, alert enrichment, alert convergence, and alert handling can be fully realized.
2. Product Features
Effortless Alert Consolidation
Easily integrates with various monitoring systems to comprehensively collect alerts and achieve centralized alert event management. Provides a low-cost, low-barrier alert source adapter development framework — online script debugging, flexible extension, and self-service development.
Alert Storm Prevention
Convenient and flexible alert noise reduction configuration, supporting automatic deduplication, anti-jitter alert suppression, correlation-based aggregation, dependency-based shielding, and maintenance window shielding. Standard noise reduction effectiveness exceeds 70%, protecting IT O&M personnel from the barrage of ineffective alerts.
Precise Alert Event Management
Integrates with CMDB, ticketing system, and CanWay BlueWhale Automation Operation Center to enable alert enrichment, dynamic dispatching, ticket creation, self-healing, and automatic closure. Supports integration with WeChat and enterprise WeChat for mobile alert management, shortening fault repair time and reducing business risks caused by missed alerts.
Alert Impact and Alert Correlation Analysis
Presents the full picture of alerts from multiple dimensions — including basic alert event information fields, correlated information (metric trend charts, associated logs, etc.), statistical reports, and business topology — helping IT O&M personnel quickly locate faults.
Algorithm-Assisted Analysis and Handling
Leveraging large language model algorithm capabilities to further strengthen alert handling, lower the IT O&M barrier, and accelerate fault resolution speed and efficiency.
3. Product Advantages
Integrated Linkage
Based on the BlueKing PaaS Platform;
Full-scenario observability chain integration:
① IT monitoring – Alert linkage: Confirm alert severity and type. Alert severity and type reflect the seriousness and urgency of an alert, helping classify alert information, prioritize responses, and formulate corresponding emergency response plans.
② APM – Alert linkage: Greatly enhances organizational capabilities in monitoring, management, and system maintenance. APM tools provide the system call relationships needed by the alert management system, and alerts can be matched with APM call data. This integration helps organizations better discover alert generation chains, shortening fault repair time and reducing congestion.
③ Log management – Alert linkage: Through this mechanism, log context information associated with alerts can be quickly retrieved and displayed. This information typically covers key elements such as timestamps, event locations, and environment configurations, helping users better understand alert information and more accurately identify the root cause.
Combined with CanWay BlueWhale Automation Operation Center / ITSM to achieve a complete anomaly-to-resolution closed loop;
Combined with CMDB to enable alert data enrichment.
Powerful Alert Noise Reduction
Effectively avoids being overwhelmed by alert storms, reducing business risks caused by missed faults.
With powerful alert filtering rules + flexible notification configuration, notification effectiveness is ensured.
Rapid Problem Localization and Resolution
Integrates with CMDB to consume instance information;
Based on business topology and correlation relationships, quickly view the alert impact scope;
Integrates with multiple IT O&M tools, supports mobile alert handling, simplifies the alert handling workflow, and shortens fault response and recovery time.
Easy Rollout and Deployment
Supports rapid onboarding and extension of multiple alert sources;
Convenient configuration of alert enrichment, shielding, aggregation, and handling rules;
Supports multi-channel notification expansion.
Intelligent Analysis
Leverages large language model algorithms for supervised learning and training on knowledge base content, enabling knowledge recommendations when alerts are generated;
Provides an intelligent assistant that offers fault localization analysis and recommended resolution plans through a conversational interface.
4. Implementation Outcomes
Designed for alert lifecycle management, the CanWay BlueWhale Alert Management Platform supports a structured approach to alert implementation, helping enterprises accelerate closed-loop IT O&M fault resolution. Representative implementation outcomes are presented below:
① Major Airport
Timely Alert Detection
The timeliness of alert notifications improved by 150%, with alerts accurately delivered within one minute of being generated. This reduced the duration of business impact, improved business stability, and increased user satisfaction.
Effective Support for Root Cause Analysis (RCA)
With broader alert coverage and effective alert convergence, together with alert topology views and correlated alert information, the platform enabled correlation analysis as soon as an alert occurred. This helped teams identify root causes and key issues more quickly, improved alert handling efficiency, and reduced both the scope and duration of alert impact.
Ensuring Business Continuity
By integrating with the monitoring platform, the airport enhanced its existing monitoring systems with Network Service Synthetic Monitoring, log keyword monitoring, customized business monitoring, component monitoring, and NTP monitoring, enabling broader and more granular monitoring. Alert coverage reached 90%.
Component-level alerts were detected before they escalated into business disruptions. Alert correlation analysis then helped identify affected business systems and related alerts so that issues could be resolved promptly, preventing multiple accumulated faults from cascading into wider business failures. This safeguarded business continuity to the greatest extent possible.
② Large Insurance Group
Achieved unified alert lifecycle management across multiple monitoring systems, bringing 100% of alerts under management.
Implemented dynamic alert dispatch and accurate notification, with less than one minute from alert generation to notification.
Achieved alert convergence tailored to the characteristics of the financial industry, reducing resource waste caused by invalid alerts and reaching a 70% alert convergence rate.
Combined multiple automation scenarios to reduce routine manual maintenance and management costs.
Provided data support for alert governance, enabling optimization at every stage through data-driven retrospective analysis.
③ Major Securities Firm
Achieved unified alert lifecycle management across multiple monitoring systems, bringing 100% of alerts under management.
Supported high-concurrency alert scenarios with daily alert volumes of 1–2 million, peak rates of more than 1,000 alerts per minute, an alert-to-notification time of under one minute, and alert query/detail page load times of approximately three seconds.
Reduced daily alert volume from more than 1 million to approximately 1,000 effective alerts through convergence, achieving a convergence rate of over 99%.
Improved alert processing efficiency by 16% through mobile alert handling.
Alert self-monitoring ensured the normal operation of all Alert Center components, reduced information gaps at critical stages, and supported continuous operation.

















