In recent years, as enterprise business scales have grown and IT architectures have become increasingly complex, technologies such as cloud computing and microservices have begun to be progressively explored and implemented across organizations. These technological advancements not only pose significant challenges to internal IT O&M management, but also place higher demands on monitoring systems.
A telecom operator boldly adopted forward-thinking concepts when building its IT architecture and organizing its departments, implementing a fully distributed design for all business systems and establishing dedicated SRE (Site Reliability Engineering) operations divisions. The CanWay BlueWhale Full-Stack Observability Center provided critical tooling support for the R&D testing and rapid iteration of business systems, and offered a unified capability platform for SRE operations teams to observe business system operations and promptly locate, analyze, and handle alerts.
01 Business Scenario
Advanced application architectures such as distributed, microservices, and Cloud-Native — while enabling agile development, rapid iteration, and elastic scaling — break down monolithic applications into multiple independently deployed, interconnected composite applications. The number of applications grows exponentially, dependencies between business modules become highly intricate, and establishing real-time, effective mapping relationships across different business layers and dimensions becomes difficult. Meanwhile, as containers are frequently started and stopped, changes in monitoring targets and their metrics become the norm, making it difficult to preserve fault scenes and effectively locate the root causes of issues.
02 Pain Point Analysis
The observability challenges of Cloud-Native architectures described above pose severe challenges to fault analysis, root cause localization, and business continuity and stability in application monitoring. The key observability difficulties can be summarized as follows:
Complex information dimensions make it difficult to establish multi-dimensional data correlation mappings. Monitoring metrics for Cloud-Native applications span multiple layers of resource attributes and performance indicators, including application processes, middleware, container orchestration platforms, container processes, and resource infrastructure. Furthermore, application troubleshooting and performance profiling involve complex interactions across multiple services and components, requiring Root Cause Analysis (RCA) based on request chain dependency relationships.
Dynamic architectural changes make it difficult to preserve fault scenes and locate problems. Container deployment architectures are based on a declarative, desired-state design philosophy, where deployed resource instances change frequently and service node drifting becomes routine. An observational analysis matrix built on the correlation mapping of multi-dimensional detailed data and metric data can effectively replay historical fault scenes.
03 Solution
Vertical and Horizontal Fault Localization
Vertical: Establish runtime software architecture cascading object drill-down analysis logic. Build global dependency topologies for different services based on actual business traffic, enabling panoramic analysis of individual business domains within selectable time ranges. Through topology node size and color differentiation, effectively analyze service traffic loads and service health status. Support drill-down analysis on service nodes — including service requests, load, errors, and latency golden metrics within specified time ranges. Within a service, drill down further into individual interfaces or individual service instances for deeper fault localization analysis. By associating service instances with CMDB-managed resources (hosts, containers), drill down to the IaaS layer to analyze the impact of infrastructure monitoring metric anomalies on service traffic.
Horizontal: Build per-request trace tracking based on Trace correlation. Each business request generates a unique request identifier at the entry service. As traffic passes through multiple downstream services, the unique request identifier, current node request identifier, and upstream service information are propagated as context, thereby constructing the complete business call chain. Additionally, users can capture business-characteristic data from HTTP request headers, request parameters, cookies, and other sources based on actual business scenarios for data instrumentation. During trace analysis, request dependency relationships filtered by specified business characteristics assist in business anomaly analysis.
Correlating Call Chains with Log Details for Root Cause Localization
When KAPM and KLC are deployed together, traces can be correlated with log details to support efficient Root Cause Analysis (RCA). KAPM reveals request dependencies and narrows the troubleshooting scope, while linked log details provide the evidence needed to identify the underlying cause—bridging the "last mile" of log troubleshooting.
04 Results Showcase
Full Coverage of Core Application Systems


Application Overview Dashboard Based on Runtime Status

Automatic Discovery of Application-Associated Resources

Interface-Level Operational Health Monitoring

Real-Time Retrieval of System Request Traces

05 Outcomes
Achieved full-stack monitoring and observability coverage across all application systems with one-stop topology and metric monitoring;
Achieved automatic application topology generation, solving architecture analysis challenges in the new paradigm;
Achieved automatic association of peripheral assets, improving system fault analysis efficiency;
Achieved application call-chain sampling and analysis, enabling system operations to be traced down to each individual request.
06 Applicability
The CanWay BlueWhale Full-Stack Observability Center is designed for enterprises with distributed architecture designs and microservice-based system units. It is suitable for the following types of enterprises:
Enterprises that are currently undergoing or have completed distributed architecture transformation;
Enterprises whose application development or IT O&M teams recognize the need for application performance observability and can use it effectively;
Enterprises that already have a Monitoring and Alerting and log framework in place, and wish to deepen observability capabilities beyond basic monitoring;
Enterprises that have already deployed other CanWay BlueWhale module products and wish to achieve an integrated observability center;
Enterprises experiencing difficult application development troubleshooting and low iteration efficiency, seeking to leverage observability products to enable rapid R&D.

















