Enterprise spending on AI-optimised infrastructure is projected to reach $758 billion annually by 2029, accelerating the shift toward autonomous digital ecosystems. For operations teams, this transition means AI must evolve from a creative tool into a deterministic, fact-based system to prevent hallucinations from causing costly system outages.
Balancing creative and deterministic AI
Generative AI functions as a creative artist, effective for brainstorming and generating ideas. However, critical IT operations require a scientist. Dynatrace's chief technology officer explained, "The generative AI everyone is talking about is a probabilistic system. It's a creative artist which is brilliant for brainstorming, generating novel content, and accelerating ideation. Probabilistic approaches shine where complexity exceeds deterministic reach. But for critical IT operations, you also need a scientist: a deterministic, fact-based system that is reliable and comprehensible."
While 50% of organisations currently run AI agents in production for limited use cases, only 23% have scaled these projects to mature, enterprise-wide integration. Hallucinations in agentic chains are not minor errors. They can accumulate, trigger incorrect actions, and expose the business to severe financial and security risks.
The three stages of autonomous operations
Organisations progress toward autonomy through three distinct maturity stages. The first stage is automated, where systems move beyond brittle scripts to reliably execute responses to known problems using real-time, contextual data.
The second stage is supervised autonomous. In this phase, AI analyses novel situations, understands the business impact, and generates a ready-to-implement action plan. Human experts must approve this plan before execution, keeping people in the loop for critical decisions while offloading initial cognitive burdens.
The final stage features fully autonomous systems. These systems operate independently to manage environments, optimise costs, and remediate issues before they affect users. Humans transition into an architect role, reviewing outcomes, adjusting strategies, and defining what the system should deliver.
Why observability is mandatory for reliable AI
AI agents lack real-world awareness without observability feeding them precise context. Organisations must build systems on a unified AI data lakehouse combined with a real-time dependency graph that maps every service and infrastructure component.
Large language models cannot directly process petabytes of heterogeneous observability data due to limited context windows. Curating and structuring only the most relevant information yields better results than flooding the model with raw data. This is where AI for IT & Development strategies must prioritize contextual analytics to deliver crisp, accurate data to agents at speed.
Why this matters for operations professionals
Operations teams must shift from managing manual workflows to overseeing AI-driven execution. Building a reliable AI ecosystem requires establishing strong guardrails and high-quality data foundations before scaling agentic systems. Implementing these AI for Operations frameworks ensures that automated systems remain transparent, accountable, and support business goals.
Your membership also unlocks: