Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Introduction to Agentic AI in Operations
- From static runbooks to reasoning agents: The evolution of IT automation
- Agent anatomy: Reasoning loops, tool utilization, memory, and planning
- Determining when to automate versus when to retain human oversight
Agent Frameworks and Architectures
- Single-agent patterns: ReAct, Plan-and-Execute, and tool-calling loops
- Multi-agent architectures: Supervisor, hierarchical, and swarm patterns
- Framework comparison: LangGraph, CrewAI, AutoGen, and custom agents
- Building your first operational agent: Querying monitoring, diagnosing issues, and proposing solutions
Tool Integration for IT Operations
- Connecting agents to Prometheus, Grafana, Datadog, and PagerDuty APIs
- Log querying via agents: Integrating Elasticsearch, Loki, and Splunk
- Utilizing infrastructure tools: kubectl, Terraform, and Ansible through agent actions
- Designing secure tool interfaces with parameter validation and idempotency
Incident Response Automation
- Automated incident triage: Severity classification and routing
- Generating root cause hypotheses and gathering evidence
- Automated remediation: Executing restart, scale, rollback, and failover actions
- Constructing an incident runbook agent with progressive levels of autonomy
Safety, Guardrails, and Human-in-the-Loop
- Action classification: Distinguishing read-only, low-risk, high-risk, and destructive actions
- Approval gates and escalation policies for critical operations
- Guardrail patterns: Action allowlists, blast radius limitations, and rollback guarantees
- Audit trails and decision provenance for compliance purposes
Multi-Agent Orchestration for Complex Incidents
- Coordinating specialized agents: Triage, diagnosis, and remediation agents
- Inter-agent communication and management of shared context
- Resolving conflicts when agents propose contradictory actions
- End-to-end simulation of major incidents with multi-agent response
Observability and Evaluation
- Tracing agent reasoning chains for debugging and auditing
- Evaluating agent decision quality: Precision, recall, and time-to-resolution
- Feedback loops: Learning from operator overrides and outcomes
- Cost tracking and token economics for operational agents
Production Deployment and Operations
- Deploying agents as services: Utilizing APIs, webhooks, and scheduled jobs
- Gradual autonomy rollout: Transitioning from shadow mode to full auto-remediation
- Runbooks for agent failures: Managing scenarios where the agent itself breaks
- Building the business case and measuring ROI for autonomous operations
Requirements
- Practical experience with IT operations, DevOps, or SRE practices.
- Familiarity with Python scripting and REST APIs.
- A foundational understanding of LLM capabilities and prompt engineering.
Target Audience
- SRE and DevOps engineers exploring AI-driven automation.
- Platform engineers developing self-healing infrastructure.
- IT operations leaders evaluating agentic AI for incident management.
14 Hours