Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps
- Defining AIOps and its significance
- Conventional oversight versus AIOps-led observability
- AIOps structure and essential elements
Gathering and Standardizing Operational Information
- Kinds of observability data: indicators, logs, and traces
- Importing data from diverse sources (servers, containers, cloud)
- Utilizing agents and exporters (Prometheus, Beats, Fluentd)
Information Linking and Anomaly Identification
- Temporal series linking and statistical approaches
- Applying ML models for spotting anomalies
- Identifying incidents across decentralized systems
Notification Strategies and Clutter Minimization
- Creating smart notification rules and limits
- Muting, removing duplicates, and grouping alerts
- Connecting with Alertmanager, Slack, PagerDuty, or Opsgenie
Root Cause Investigation and Visual Analysis
- Using control panels to display indicators and track patterns
- Analyzing events and timelines for RCA
- Tracking problems across layers using distributed tracing utilities
Automation and Resolution
- Activating automated scripts or workflows from incidents
- Integrating with ITSM platforms (ServiceNow, Jira)
- Application examples: self-healing, scaling, traffic redirection
Open Source and Commercial AIOps Solutions
- Overview of tools: Prometheus, Grafana, ELK, Moogsoft, Dynatrace
- Assessment criteria for choosing an AIOps platform
- Demonstration and practical work with a chosen stack
Recap and Future Directions
Requirements
- A solid grasp of IT operations and system oversight principles
- Practical experience with oversight instruments or control panels
- Basic knowledge of standard log and indicator structures
Target Audience
- Operational teams managing infrastructure and applications
- Site Reliability Engineers (SREs)
- IT oversight and observability specialists