Against the backdrop of localization, digitalization, and cloud adoption, O&M targets are becoming increasingly complex. Financial IT architectures—running a multitude of intricate systems—inevitably encounter failures or anomalies. Alerts generated through appropriate monitoring mechanisms enable personnel to monitor system status in real time, address errors promptly, and avoid or mitigate potentially severe consequences.
However, as business scale expands and system complexity increases, alert volumes have surged dramatically. The original alert handling processes and staffing can no longer meet actual demands. During peak periods, the alert center must process hundreds of alerts per minute, leading to delayed handling. In severe cases, delayed alerts can escalate into real system issues that impact business operations.
To address this challenge, the alert center requires a rich set of alert convergence strategies and a sound technical architecture to cope with the rapid growth in alert volume—ensuring that every alert is handled promptly and accurately to safeguard stable system operation.
01 Business Scenario
After years of development, a major securities firm had built monitoring capabilities spanning operating systems, hardware, databases, network devices, business monitoring, and other log monitoring. However, due to the wide scope and sheer volume of alerts, the company's legacy alert system was unable to handle peak alert loads, causing the original unified alert center to crash and failing to meet the requirements for continuous alerting and IT O&M.
To continuously fulfill the monitoring and O&M needs of the technology department—and effectively address scenarios including multi-source alert integration, reasonable alert convergence, high peak alert volumes, real-time alert handling, and alert center self-monitoring—a new alert management platform needed to be built.
02 Pain Point Analysis
The business model and system architecture of this major securities firm led to the following pain points in real-world scenarios:
Fragmented Monitoring Tools
Infrastructure monitoring, log monitoring, network device monitoring, database monitoring, network probing, business SQL queries, and other monitoring scenarios were scattered across different tools. Each tool handled notifications and alert closure independently, requiring different roles to monitor multiple tools—making unified monitoring extremely difficult to achieve.
Performance Issues from Alert Peaks
Network probing and business SQL query monitoring requirements resulted in high alert volumes with frequent concurrency spikes. The company's daily alert volume reached the millions, with severe concurrency issues.
Lack of Proper Alert Convergence
High alert volumes without effective convergence led to massive numbers of invalid alert notifications, dispersing personnel attention and reducing processing efficiency.
Inability to Handle Alerts Anytime
O&M personnel could not be on-site 24/7, making it impossible to handle alerts remotely in a timely manner—especially during alert storms, when batch suppression was needed.
No Self-Monitoring for the Alert Center
After alerts are generated, the entire alert lifecycle is managed within the alert center. If the alert center itself experiences anomalies, alert information cannot be obtained in time, preventing timely handling.
03 Solution
Based on the customer's business model, organizational structure, and management requirements, an alert center was built to enable unified alert integration. Mobile capabilities were also developed to allow alert handling anytime, anywhere.
Multiple Alert Integration Methods for Multi-Source Access
With numerous and fragmented monitoring systems, different integration methods were required for different systems to unify alert data. These include Kafka, API integration, and plugin-based integration.
Kafka Push: For monitoring systems with push capabilities, alerts can be pushed to a Kafka queue upon generation. The alert center periodically retrieves alert information from the Kafka queue, enabling centralized alert ingestion.
API Integration: For monitoring platforms that provide APIs, alerts can be integrated into the alert center via both push and pull approaches.
Plugin Integration: For alerts generated natively by the CanWay BlueWhale Full-Stack Observability Center, built-in plugins enable direct integration with the alert center, minimizing delays in alert ingestion.
Architecture Supporting High-Concurrency Alerts
The alert center's technical architecture incorporates high-concurrency design to handle large-scale alert traffic.
First, a distributed system architecture is employed, distributing alert processing loads across multiple server nodes to reduce pressure on individual servers and increase overall processing capacity.
Second, message queues are leveraged for asynchronous processing. When the alert center receives a large volume of alert requests, these requests are first placed into a message queue, then read and processed asynchronously by dedicated processing tasks—ensuring smooth system operation.
Finally, different database types are adopted for different data categories, with optimized database design. For high-volume alert events, Elasticsearch is used for storage, combined with efficient indexing, table/database sharding, and read-write separation to significantly reduce database I/O pressure and enhance concurrent processing capacity.
Through this design, the alert center's architecture effectively supports high-concurrency alert processing, ensuring stable operation and timely response even when faced with massive alert volumes.
Rich Alert Convergence Strategies to Minimize Invalid Alerts
To accommodate convergence needs across various business scenarios, a rich and flexible set of alert convergence rules is provided. These rules consolidate multiple alerts originating from the same event or failure, reducing redundant alert information and significantly improving alert processing efficiency.
Matching conditions can be written against any field in the alert information and combined flexibly to suit various business scenarios. Through alert noise reduction, the processing pressure and decision-making confusion caused by large volumes of duplicate alerts are effectively eliminated.
Mobile Alert Handling Anytime, Anywhere
After alerts are generated, handling is not limited to the PC—alerts can also be notified, viewed, processed, and suppressed via mobile devices.
Alert Notification: Alert information is delivered through enterprise WeChat message push or group chatbot, enabling awareness at any time.
Alert Viewing: The mobile interface allows viewing of alert details, handling records, and associated information for continuous alert awareness.
Alert Handling: Synchronized with the PC interface, alerts can be routed and closed on mobile, with batch processing support to increase efficiency.
Alert Suppression: For alert storm scenarios, suppression policies can be created on mobile to reduce excessive invalid notifications.
Multi-Client Login: Catering to different user preferences, the enterprise WeChat PC client also supports mobile-style viewing, reducing PC login verification steps and improving work efficiency.
Alert Self-Monitoring to Ensure Normal Alert Routing
Core components of the alert center are monitored using Prometheus, with monitoring metrics displayed through Grafana dashboards, providing an at-a-glance view of the alert center's health status.
The monitored components and metrics are as follows:
04 Results Showcase
Alert Source Integration

Alert Convergence Policy Rules


Mobile Alert Handling
● Mobile Notification

● Viewing and Handling Alerts

● Suppression Policy Creation

● PC Client Viewing Mobile-Adapted Interface


Alert Self-Monitoring
● Self-Monitoring Dashboard

● Alert Stress Test Report
The following shows stress test data from 6 runs with different parameters. "Worker" refers to the number of processes updating alert status.
05 Implementation Outcomes
Achieved unified alert lifecycle management across multiple monitoring systems, with 100% of alerts under management.
High-concurrency alert scenario support: daily alert volume of 1–2 million, peak alert rate exceeding 1,000+ per minute, alert-to-notification time under 1 minute, and alert query/detail page load time approximately 3 seconds.
Daily alerts reduced from 1 million+ to approximately 1,000+ effective alerts through convergence, achieving a convergence rate exceeding 99%.
Mobile alert handling improved alert processing efficiency by 16%.
Alert self-monitoring ensures normal operation of all alert center components, reduces information breakpoints at critical stages, and enables continuous operation.
06 Scenario Applicability
The CanWay BlueWhale Alert Management Platform is designed for alert lifecycle management scenarios. Based on enterprise business models and management requirements, the following applicable customer types and demand scenarios have been identified:
Enterprises with numerous monitoring systems but no centralized alert management system.
Enterprises with high alert data volumes, high concurrency, and strict timeliness requirements.
Enterprises that need to handle alert information anytime via mobile devices.
Enterprises with high business continuity requirements that need self-monitoring of their Monitoring and Alerting systems.

















