In many IT organizations, the focus is often on resolving incidents. However, focusing on this aspect over the long term can lead to a reactive cycle, such as high incident volumes and overwhelmed engineers. Over time, it can result in a continuous increase in problems, with unresolved root causes triggering even more incidents. If an IT organization aims to effectively implement problem management while balancing incident management, it is essential to strike a balance between the two.
01 How Problem Management Works
The sole objective of problem management is to identify and eliminate the root causes of recurring incidents. When incidents cannot be prevented, problem management strives to minimize the impact of these incidents on business operations.
If the focus is solely on "how to quickly identify and restore services," it is not problem management but incident management. The core goal of incident management is rapid service restoration. Problem management, on the other hand, is a completely different process, primarily divided into two types: reactive and proactive.
Reactive Problem Management
Reactive problem management is triggered by incidents. Many IT organizations conduct post-incident reviews for major incidents and initiate reactive problem management when potential underlying issues are identified.
Proactive Problem Management
Proactive problem management identifies potential problems through data trends and historical information. It can involve ongoing service improvement activities, data analysis, or even reliance on accumulated experience and intuition.
Regardless of the approach, problem management must prioritize issues based on their business value. For example, methods like "business impact analysis" can help identify which problems will deliver the greatest value to the business.
02 How Organizations of Different Scales Can Build Problem Management
The design of problem management in an IT organization should vary based on its scale. When determining the model for a problem management process, the following factors should be considered:
● Number of IT O&M personnel
● Scale of infrastructure
● Stability of the infrastructure
● Volume of recurring incidents
If these factors are difficult to measure effectively, the following guidelines can serve as a reference:
Problem Management for Small-Scale Organizations
In small organizations, problem management typically does not have a dedicated process manager. Instead, it is primarily discussed in regular periodic meetings. Before these meetings, it is recommended that responsible personnel from various domains summarize the most critical issues within their respective areas based on records from the previous cycle. These issues are then discussed and confirmed during the meeting, with investigation and resolution efforts planned for the next cycle.
Problem Management for Medium and Large-Scale Organizations
In medium and large organizations with multiple business domains, a unified problem-management model is usually adopted, with a focus on identifying and implementing solutions. Proactive problem management typically defines multiple sources of problems, such as Monitoring and Alerting events that recur during specific periods, incidents repeatedly reported by users, major incidents, potential issues discovered during routine inspections, and critical flaws occasionally identified in business processes or services. A problem manager usually collects and consolidates these issues regularly, coordinates resolution, and tracks progress.
In daily IT O&M work, potential problems should be identified proactively rather than only after incidents recur.
Proactive health checks: Conduct periodic health checks to analyze the operational status of application systems, proactively identify issues, prevent major incidents, and eliminate system vulnerabilities.
Continuous tracking and handling of identified issues: regularly report the progress of problem resolution to relevant personnel.
Continuous optimization: The problem manager or responsible personnel of the system should continuously refine health check methods, as well as the identification and resolution progress of problems.
03 How to Implement Effective Problem Management
Distinguishing Between Incidents and Problems, and Defining Management Responsibilities
As noted above, incident management and problem management have different objectives. Incident management focuses on resolving incidents promptly and restoring services, whereas problem management emphasizes preventive measures to identify and eliminate underlying issues that may cause incidents or other adverse impacts. Clearly distinguishing the two enables IT teams to shift from reactive emergency response to proactively identifying and mitigating potential risks, thereby improving overall service quality and stability.
Similarly, for an incident manager, the priority is the rapid resolution of incidents, while a problem manager's goal is prevention. By combining the efforts of these two roles, the continuity and availability of application systems can be fundamentally improved.
Conducting Thorough Problem Analysis
There are numerous methods for analyzing problems. Organizations can employ different approaches in various scenarios to achieve precise and effective problem analysis. Below are analytical tools applicable in different contexts:
Example of the 5 Whys:
Example of a Fishbone Diagram:
Adopting a Result-Oriented Approach
Many IT organizations tend to overemphasize the number of problems and resolution times during problem management. However, these are not the core criteria for measuring the effectiveness of problem management. Truly effective problem management should be evaluated through two key dimensions: the KPI of problem management and its actual impact on business operations. The following examples can be used for reference:
Leveraging the Known Error Database
This also reflects a knowledge-management best practice: grant different teams access to the Known Error Database and related solutions. This approach facilitates mutual learning, saves time in handling incidents and problems, and helps the organization operate more efficiently.
04 Conclusion
By implementing effective problem management, IT organizations can address recurring incidents at their root, while significantly improving service stability and customer satisfaction. Clearly distinguishing the responsibilities of incident management and problem management, and using appropriate analytical tools such as brainstorming, the 5 Whys, and fishbone diagrams, helps teams identify root causes more quickly and implement effective preventive measures. Regularly reviewing and using the Known Error Database further strengthens problem management. Ultimately, problem management aims to achieve efficient, reliable, and sustainable IT services through continuous improvement. An advanced ITSM platform integrates problem management with knowledge management to prevent recurring issues. An advanced ITSM platform integrates problem management with knowledge management to prevent recurring issues.


























