Get in Touch

Course Outline

Introduction to Agentic AI in Operations

  • From static runbooks to reasoning agents: The evolution of IT automation
  • Agent anatomy: Reasoning loops, tool utilization, memory, and planning
  • Determining when to automate versus when to retain human oversight

Agent Frameworks and Architectures

  • Single-agent patterns: ReAct, Plan-and-Execute, and tool-calling loops
  • Multi-agent architectures: Supervisor, hierarchical, and swarm patterns
  • Framework comparison: LangGraph, CrewAI, AutoGen, and custom agents
  • Building your first operational agent: Querying monitoring, diagnosing issues, and proposing solutions

Tool Integration for IT Operations

  • Connecting agents to Prometheus, Grafana, Datadog, and PagerDuty APIs
  • Log querying via agents: Integrating Elasticsearch, Loki, and Splunk
  • Utilizing infrastructure tools: kubectl, Terraform, and Ansible through agent actions
  • Designing secure tool interfaces with parameter validation and idempotency

Incident Response Automation

  • Automated incident triage: Severity classification and routing
  • Generating root cause hypotheses and gathering evidence
  • Automated remediation: Executing restart, scale, rollback, and failover actions
  • Constructing an incident runbook agent with progressive levels of autonomy

Safety, Guardrails, and Human-in-the-Loop

  • Action classification: Distinguishing read-only, low-risk, high-risk, and destructive actions
  • Approval gates and escalation policies for critical operations
  • Guardrail patterns: Action allowlists, blast radius limitations, and rollback guarantees
  • Audit trails and decision provenance for compliance purposes

Multi-Agent Orchestration for Complex Incidents

  • Coordinating specialized agents: Triage, diagnosis, and remediation agents
  • Inter-agent communication and management of shared context
  • Resolving conflicts when agents propose contradictory actions
  • End-to-end simulation of major incidents with multi-agent response

Observability and Evaluation

  • Tracing agent reasoning chains for debugging and auditing
  • Evaluating agent decision quality: Precision, recall, and time-to-resolution
  • Feedback loops: Learning from operator overrides and outcomes
  • Cost tracking and token economics for operational agents

Production Deployment and Operations

  • Deploying agents as services: Utilizing APIs, webhooks, and scheduled jobs
  • Gradual autonomy rollout: Transitioning from shadow mode to full auto-remediation
  • Runbooks for agent failures: Managing scenarios where the agent itself breaks
  • Building the business case and measuring ROI for autonomous operations

Requirements

  • Practical experience with IT operations, DevOps, or SRE practices.
  • Familiarity with Python scripting and REST APIs.
  • A foundational understanding of LLM capabilities and prompt engineering.

Target Audience

  • SRE and DevOps engineers exploring AI-driven automation.
  • Platform engineers developing self-healing infrastructure.
  • IT operations leaders evaluating agentic AI for incident management.
 14 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories