Council Post: Why Autonomous Operations Require More Than AI Models
Dr. Dale Skeen, Chief Technology and AI Officer, Vitria Technology.

getty
The biggest limitation in today’s AI infrastructure is not model intelligence. It is the absence of operational understanding.
Over the last decade, enterprises and telecom operators have invested heavily in observability, telemetry pipelines and AI-driven automation. Modern systems can ingest enormous amounts of operational data, detect anomalies in real time and correlate events across highly distributed environments. Yet despite these advances, fully autonomous operations remain elusive.
The reason is becoming increasingly clear. Most AI systems are designed to recognize patterns, not understand operational reality. That distinction matters as organizations move from AI-assisted operations toward agentic AI systems capable of making autonomous decisions inside live production environments.
Autonomous operations require more than anomaly detection. They require contextual intelligence. A machine learning model may identify a degraded network function or predict an impending infrastructure issue. But autonomous resolution requires understanding service dependencies, infrastructure topology, remediation sequencing, operational policies, historical outcomes and business impact.
In other words, autonomous systems must understand operational causality, not simply statistical relationships. This is where many current AI architectures break down.
Why Operational Context Matters
Most operational environments contain vast amounts of institutional knowledge, but very little of it exists in a continuously usable form. Critical intelligence is fragmented across configuration management databases (CMDBs), tickets, logs, engineering notes, runbooks and years of accumulated operational experience. Human operators naturally synthesize this context when making decisions. Traditional AI systems generally do not. That gap creates a growing challenge for agentic AI.
As organizations give AI systems increasing authority to diagnose issues, initiate workflows and execute remediation actions, trust becomes the defining requirement. Operators need to understand not only what an AI system recommends, but why it reached that conclusion and whether the action is operationally safe.
Consider a cloud-native telecom core environment where rising storage latency begins affecting a virtualized user plane function (UPF). A conventional AI system may detect degraded performance and initiate isolated remediation actions, but the operational reality is far more complex.
The issue may cascade across dependent network functions, degrade subscriber sessions and impact revenue-generating services across multiple domains. Resolving the issue autonomously requires understanding the relationships between infrastructure, services, operational policies and downstream business impact before action is taken.
This is why probabilistic outputs alone are not enough. As AI systems become more autonomous, operational reasoning becomes more important than automation itself.
Three Ways A Knowledge Plane Transforms Operational Excellence
This is driving the emergence of the knowledge plane. A knowledge plane is not simply a static knowledge graph layered beside an AI platform. It is a continuously evolving operational intelligence layer embedded directly into detection, analysis and resolution workflows. Its role is to transform operational experience into machine-usable reasoning, which creates several important shifts:
1. AI systems become more explainable.
Recommendations can be traced through infrastructure relationships, operational context and learned remediation paths rather than functioning as opaque outputs. Explainability becomes essential if organizations expect operators to trust AI systems inside mission-critical environments.
2. Systems become adaptive.
Modern infrastructure changes continuously. Static operational models decay rapidly in cloud-native environments. AI systems must continuously acquire and refine operational knowledge as infrastructure, services and dependencies evolve.
3. Automation becomes safer.
The challenge with autonomous operations is no longer whether an action can be automated. The challenge is whether the system understands the broader operational consequences of taking that action. That requires contextual awareness, not just statistical confidence.
The Next Phase Of AIOps
Over the next several years, I believe the industry will begin separating AI platforms into two categories: systems that generate outputs and systems that can reason operationally. That distinction will become increasingly important as enterprises move toward autonomous infrastructure where AI systems are expected to make real-time operational decisions with measurable business consequences.
This shift is particularly important in telecom environments shaped by 5G, cloud-native architectures, edge computing and increasingly dynamic service orchestration. Human operators alone cannot scale to manage this level of complexity in real time.
At the same time, organizations cannot afford autonomous systems that act without operational understanding. That tension is what I believe will define the next era of AI operations (AIOps).
The industry has largely focused on building AI systems that can process more data, generate faster outputs and automate isolated workflows. But the next era of autonomous infrastructure will not be defined by which AI systems generate the fastest response. It will be defined by which systems can understand operational reality well enough to make the right decision.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?