01 Three Major Stages of Alert Event Management
The full lifecycle management of alert events can be divided into three major stages: pre-incident, during-incident, and post-incident. The pre-incident stage focuses on prevention and early detection. The during-incident stage focuses on rapid discovery and resolution of issues to ensure business continuity and minimize losses. The post-incident stage focuses on incident review, knowledge accumulation, and continuous optimization of business systems to ensure healthy operations.
Key Metrics for Alert Event Management
The most commonly used industry metrics for defining the full lifecycle of alert events include MTBF (Mean Time Between Failures), MTTR (Mean Time to Repair/Recover/Respond/Resolve), MTTF (Mean Time to Failure), and MTTA (Mean Time to Acknowledge). These metrics help technical teams understand how frequently failures occur and how quickly incidents are recovered.
The CanWay BlueWhale Alert Management Platform (hereinafter referred to as the "Alert Center"), built on the CMDB model and instances, centers on alert events and provides unified management of enterprise business system alerts — encompassing alert ingestion, alert enrichment, alert convergence, alert handling & notification, and alert analysis. Below is the full lifecycle workflow of an alert within the Alert Center.

02 Alert Center Product Features
Feature 1: Alert Ingestion
The Alert Center supports standardized plugins for over 20 common monitoring systems, including Zabbix, Prometheus, VMware, Huawei Cloud, and Alibaba Cloud — enabling out-of-the-box, rapid integration with various types of monitoring systems. It also supports ingesting alerts pushed from third-party systems via REST API.

Feature 2: Alert Enrichment
Plugin-Based Cleansing: When ingesting alerts from different systems, key alert field contents are extracted and output based on the data cleansing logic defined in each plugin.
Standard Enrichment Rules: If the output alert content does not meet the standard format requirements, standard enrichment rules can be applied to replace, extract, or adjust alert field contents.
CMDB Enrichment Rules: The Alert Center integrates with the Configuration Management Database, automatically enriching alert details with configuration information from CMDB based on object model instance relationships.

Alert Detail Information

Standard Enrichment Rules

CMDB Enrichment Rules
Feature 3: Alert Convergence
For enterprise scenarios involving alert storms and various types of false positives or missed alerts, the Alert Center provides a mature alert noise reduction solution. This includes automatic deduplication algorithms, alert suppression, alert muting, and alert merging. These convergence strategies can be flexibly configured for different business scenarios, achieving alert compression ratios of over 90%.
Automatic Deduplication Algorithm
The built-in automatic deduplication generates an alert event ID using a hash algorithm based on four fields: alert source ID, alert object, alert metric, and alert severity. Alerts with the same ID are automatically deduplicated by the system.
Alert Debounce Suppression
Debounce alert suppression is primarily designed for high-fluctuation metrics such as CPU utilization and NIC traffic. It can be configured so that a valid alert is generated only after a specified number of occurrences within a defined time window.

Debounce Suppression Rules
Correlation-Based Aggregation Suppression
Alerts can be suppressed based on custom field matching. For example, alerts with the same business name, alert object, alert metric, and alert severity can be considered identical. By applying combined conditions on these fields, duplicate alerts are suppressed.

Correlation-Based Aggregation Rules
Time-Based Muting
Time-based muting is typically used during enterprise system maintenance windows or when business systems require it, to centrally mute alerts and avoid generating a large volume of alerts and notifications.

Time-Based Muting Rules
Dependency-Based Muting
Dependency-based muting configures alert muting policies based on custom dependency relationships or the association relationships between models in the CMDB.
For example, when a server's NIC triggers an alert, the switch on that server will inevitably generate an alert as well. For such scenarios, dependency-based muting policies can be configured based on the association relationships between these objects, thereby reducing the generation of noise alerts.

Dependency-Based Muting Rules
Alert Merging
The alert merging feature handles enterprise scenarios where a single fault triggers a large volume of related alerts by consolidating them.
For example, when the transaction rate in a certain business domain drops, this may often be attributed to multiple factors — such as persistently high CPU utilization of the services the business depends on, significantly increased service response times, and so on. When alert signals for these factors are triggered simultaneously, they can be consolidated into a single comprehensive valid alert to improve handling efficiency.

Alert Merging Rules
Feature 4: Alert Handling
After applying the series of alert convergence strategies, IT O&M personnel only need to focus on and handle the valid alerts. The Alert Center provides both manual and automated handling options to accelerate alert event response and resolution. Additionally, the Alert Center offers rich notification channels covering both PC and mobile platforms, ensuring that relevant personnel receive notifications immediately and are promptly aware of system issues.
Auto-Close
For alerts that may not affect core system functionality or are not urgent — such as performance alerts from test machines or alerts that do not require handling on non-working days — auto-close policies can reduce the workload of alert management.

Auto-Close Policy
Auto-Dispatch
Alerts can be automatically dispatched and notified to the corresponding person, team, or on-duty personnel based on IT O&M management requirements.
For example, when a server goes down or experiences performance anomalies, the Alert Center automatically dispatches the alert to the server maintenance team. When switch, router, or network device fault alerts are triggered, the system automatically dispatches them to the network operations team.

Auto-Dispatch Policy
Auto-Remediation
The Alert Center supports automated remediation capabilities for alerts. Common auto-remediation scenarios include server restart, log cleanup, and disk cleanup. Corresponding scripts can be assigned to execute remediation workflows for each scenario. The platform also supports remediation workflow parameter input, enabling rapid execution of remediation scripts to address faults.

Auto-Remediation Policy
Auto-Ticket Creation
Supports built-in integration with ITSM and third-party ticketing system platforms, enabling automated workflows from alert generation to ticket creation. It also supports ticket template creation for quick template-based parameter filling, making it convenient for IT O&M personnel to promptly create incident management tickets, change management tickets, and more — accelerating the alert-to-resolution workflow.

Auto-Ticket Creation Policy
Feature 5: Alert Notification
The Alert Center provides robust alert notification capabilities, including flexible notification frequency configuration, diverse notification channels, and customizable notification templates.
Notification Frequency
For critical and urgent alerts — such as host CPU utilization, disk utilization, and network unreachable conditions — the system should be configured to send immediate emergency notifications upon triggering. When no one responds, the system will send cyclic notifications at defined intervals; after acknowledgment, cyclic notifications continue if the alert remains unresolved.
For relatively less urgent but still noteworthy warnings — such as network bandwidth utilization reaching approximately 70% — notifications can be delayed.

Alert Notification Frequency
Alert Notification Channels
Supports diverse notification channel configurations, including common channels such as email, SMS, ESB WeChat, voice calls, DingTalk, WeCom/DingTalk mobile apps, WeCom/Feishu/DingTalk group bots, and page-based voice broadcast functionality for duty monitoring screens.

Alert Notification Channels
Alert Notification Templates
Custom notification templates can be configured for different notification scenarios according to enterprise alert notification requirements, ensuring that alerts reach the responsible personnel faster and with more detailed information.

Alert Notification Template Configuration
Feature 6: Alert Analysis
Association Topology
By integrating with the CMDB, the platform automatically retrieves topology relationship diagrams based on object models and instances. Nodes with active alerts are highlighted in red, providing an intuitive view of upstream and downstream fault associations for rapid identification of the fault impact scope.

Alert Association Topology
Alert Reports
Built-in multi-type, multi-style statistical report modules enable intuitive viewing of alert statistics and individual MTTA/MTTR metric performance.

Alert Reports
Assisted Analysis
The Alert Center supports integration with knowledge bases and ticket systems. After an alert is generated, it can quickly match associated solutions and related historical change management tickets, assisting IT O&M personnel in fault localization and handling.

Alert Assisted Analysis Module
Feature 7: Intelligent Processing
The Alert Center leverages large language model (LLM) capabilities to further enhance alert handling, lower the IT O&M barrier, and accelerate fault resolution speed and efficiency.
Knowledge Base Association
A built-in IT O&M knowledge base is included out of the box. Knowledge base files can be imported in bulk, and the LLM algorithm performs supervised learning on the knowledge base content, enabling matching between alert content and the knowledge base, with results displayed in order of relevance.

Automatic Knowledge Base Association
AI Operations Assistant
Leveraging generative AI capabilities of large language models — supporting models such as ChatGPT and LLaMA 2 — the AI Operations Assistant enables conversational fault analysis and provides recommended remediation solutions through dialogue-based interactions.

AI Operations Assistant

















