Tool sprawl, AI complicate enterprise network operations
Network World · · g5503766

Enterprise organizations typically use six to eight observability tools, creating integration, staffing, and visibility challenges as network operations teams adopt AI to help manage complex IT environments, according to research from Enterprise Management Associates (EMA). EMA surveyed 356 enterprise IT professionals and found that tool sprawl is consuming engineering time, increasing staffing demands, and slowing down incident resolution. Most organizations surveyed are already using AI in at least some areas of observability, including tasks such as alert correlation and anomaly detection, according to EMA’s report, The Reality of Observability Unification in Modern IT Operations Still, efforts to use AI to simplify observability across sophisticated environments won’t reduce tools down to a single platform. “One of the first things I can say is no one gets a single pane of glass,” said Shamus McGillicuddy , vice president of research covering network infrastructure and operations at EMA, during a web event sharing the research findings. Tool sprawl strains IT operations Many enterprises are still working to unify their observability environments. EMA found that 51% of respondents had tool unification efforts under way, while 32% were still planning or evaluating that approach and 17% had completed unification. Larger organizations reported that they had more observability tools in use across their environments. According to EMA, IT organizations cited several ways tool sprawl affects IT teams: 45% cited skills and staffing burdens from operating multiple tools. 44% said integration and API complexity consume engineering time. 41% said context switching among tools slows investigations and response. 40% reported increased manual effort. 34% said alert noise creates toil and cognitive overload. Tool sprawl also contributes to higher total cost of ownership, slow incident detection and resolution, inconsistent data quality and coverage gaps, and poor end-to-end visibility across technology domains. For instance, a network operations team can have tools showing connectivity is healthy, while cloud and application teams have their own tools indicating those environments are operating normally, but end users are still experiencing a problem. “All of our tools say everything’s fine,” McGillicuddy said, but the problem is that teams are not equipped to correlate data among the tools and recognize an anomaly associated with the incident. Enterprises can use AI to help correlate information and automate operations, but it doesn’t eliminate the need to address underlying operational problems. “AI can overcome data fragmentation across tools, but it’s not going to overcome bad data and bad operations in general,” McGillicuddy said. AI takes on more of the workload Enterprises are now applying AI to that operational data. Thirty-two percent of respondents said they use AI extensively across observability tools, 51% use it in selected areas, and 13% are still piloting it. Use cases shared with EMA included incident summaries and explanations, alert correlation and noise reduction, capacity forecasting, anomaly detection, predictive analytics, and root-cause analytics. Organizations that reported extensively using AI also said they had greater confidence in their observability unification strategies. Fragmented or inconsistent data can limit AI effectiveness, and organizations also cited concerns about inaccurate outputs and security and compliance risks. Parker Hathcock , research director covering IT service, operations, and ServiceOps at EMA, said during the webinar that the problems associated with tool sprawl can compound one another. He said strong tool governance is needed before organizations try to simplify their observability environments. “All of these issues can compound each other, so that’s why strong tool governance is essential,” Hathcock said. “It’s an essential step to get to a better place before you even start trying to simplify what you have.” About one-quarter (24%) of organizations reported fully centralized observability tool governance, and 48% said ownership was centralized for the most part with some domain-specific exceptions. Those with centralized ownership reported greater confidence in their unification strategies, deeper integration between observability and IT service management, more effective incident-impact analysis, and more mature use of AI-powered observability. McGillicuddy warned that centralization means more than putting one group in charge. And EMA found that understaffed tool governance and engineering teams had less success with their unification strategies. “It’s critical to not just centralize authority but give that centralized group the resources it needs to implement and deliver,” he said. When AI starts acting on infrastructure AI governance becomes more important when AI systems begin taking actions across network infrastructure, according to Christopher Steffen , vice president of research at EMA. Steffen explained that as enterprises give AI more autonomy, they also need to govern those systems appropriately. Agentic AI systems are probabilistic, meaning the responses and actions aren’t always predictable, while many traditional cybersecurity controls are designed to operate deterministically, which means following predefined rules. For instance, Steffen explained that a firewall rule could appear unnecessary to an AI system. The rule allows a particular IP address to access a database process for one day each quarter. An AI system might identify the rule as unnecessary and remove it without knowing the rule supports a legitimate business process. “It looks like a rule that was a mistake. It looks like a rule that should be immediately eliminated,” Steffen said in a separate webinar . This is just an example of a scenario in which network teams would have to determine where a human review remains necessary as enterprises allow AI to take on more operational functions. Steffen said that AI governance requires human oversight for unusual exceptions, business-critical rules, and potentially harmful or inaccurate outputs, with controls tailored specifically to an organization’s risk profile. For network operations teams, the increasing use of AI and automation requires integration of observability data and workflows, stronger tool governance, and controls over what AI systems can and can’t do. “AI is yet another tool in that tool chest that we’re going to continue to use. It’s not going anywhere. We’re going to continue using it. It’s going to continue to evolve. We’re going to have proper governance around it,” Steffen said.
Enterprise organizations typically use six to eight observability tools, creating integration, staffing, and visibility challenges as network operations teams adopt AI to help manage complex IT environments, according to research from Enterprise Management Associates (EMA). EMA surveyed 356 enterprise IT professionals and found that tool sprawl is consuming engineering time, increasing staffing demands, and slowing down incident resolution. Most organizations surveyed are already using AI in at least some areas of observability, including tasks such as alert correlation and anomaly detection, according to EMA’s report, The Reality of Observability Unification in Modern IT Operations Still, efforts to use AI to simplify observability across sophisticated environments won’t reduce tools down to a single platform. “One of the first things I can say is no one gets a single pane of glass,” said Shamus McGillicuddy , vice president of research covering network infrastructure and operations at EMA, during a web event sharing the research findings. Tool sprawl strains IT operations Many enterprises are still working to unify their observability environments. EMA found that 51% of respondents had tool unification efforts under way, while 32% were still planning or evaluating that approach and 17% had completed unification. Larger organizations reported that they had more observability tools in use across their environments. According to EMA, IT organizations cited several ways tool sprawl affects IT teams: 45% cited skills and staffing burdens from operating multiple tools. 44% said integration and API complexity consume engineering time. 41% said context switching among tools slows investigations and response. 40% reported increased manual effort. 34% said alert noise creates toil and cognitive overload. Tool sprawl also contributes to higher total cost of ownership, slow incident detection and resolution, inconsistent data quality and coverage gaps, and poor end-to-end visibility across technology domains. For instance, a network operations team can have tools showing connectivity is healthy, while cloud and application teams have their own tools indicating those environments are operating normally, but end users are still experiencing a problem. “All of our tools say everything’s fine,” McGillicuddy said, but the problem is that teams are not equipped to correlate data among the tools and recognize an anomaly associated with the incident. Enterprises can use AI to help correlate information and automate operations, but it doesn’t eliminate the need to address underlying operational problems. “AI can overcome data fragmentation across tools, but it’s not going to overcome bad data and bad operations in general,” McGillicuddy said. AI takes on more of the workload Enterprises are now applying AI to that operational data. Thirty-two percent of respondents said they use AI extensively across observability tools, 51% use it in selected areas, and 13% are still piloting it. Use cases shared with EMA included incident summaries and explanations, alert correlation and noise reduction, capacity forecasting, anomaly detection, predictive analytics, and root-cause analytics. Organizations that reported extensively using AI also said they had greater confidence in their observability unification strategies. Fragmented or inconsistent data can limit AI effectiveness, and organizations also cited concerns about inaccurate outputs and security and compliance risks. Parker Hathcock , research director covering IT service, operations, and ServiceOps at EMA, said during the webinar that the problems associated with tool sprawl can compound one another. He said strong tool governance is needed before organizations try to simplify their observability environments. “All of these issues can compound each other, so that’s why strong tool governance is essential,” Hathcock said. “It’s an essential step to get to a better place before you even start trying to simplify what you have.” About one-quarter (24%) of organizations reported fully centralized observability tool governance, and 48% said ownership was centralized for the most part with some domain-specific exceptions. Those with centralized ownership reported greater confidence in their unification strategies, deeper integration between observability and IT service management, more effective incident-impact analysis, and more mature use of AI-powered observability. McGillicuddy warned that centralization means more than putting one group in charge. And EMA found that understaffed tool governance and engineering teams had less success with their unification strategies. “It’s critical to not just centralize authority but give that centralized group the resources it needs to implement and deliver,” he said. When AI starts acting on infrastructure AI governance becomes more important when AI systems begin taking actions across network infrastructure, according to Christopher Steffen , vice president of research at EMA. Steffen explained that as enterprises give AI more autonomy, they also need to govern those systems appropriately. Agentic AI systems are probabilistic, meaning the responses and actions aren’t always predictable, while many traditional cybersecurity controls are designed to operate deterministically, which means following predefined rules. For instance, Steffen explained that a firewall rule could appear unnecessary to an AI system. The rule allows a particular IP address to access a database process for one day each quarter. An AI system might identify the rule as unnecessary and remove it without knowing the rule supports a legitimate business process. “It looks like a rule that was a mistake. It looks like a rule that should be immediately eliminated,” Steffen said in a separate webinar . This is just an example of a scenario in which network teams would have to determine where a human review remains necessary as enterprises allow AI to take on more operational functions. Steffen said that AI governance requires human oversight for unusual exceptions, business-critical rules, and potentially harmful or inaccurate outputs, with controls tailored specifically to an organization’s risk profile. For network operations teams, the increasing use of AI and automation requires integration of observability data and workflows, stronger tool governance, and controls over what AI systems can and can’t do. “AI is yet another tool in that tool chest that we’re going to continue to use. It’s not going anywhere. We’re going to continue using it. It’s going to continue to evolve. We’re going to have proper governance around it,” Steffen said.