Get in Touch
 Duration 14 hours

Course Outline

Introduction to AIOps

  • Defining AIOps and its significance
  • Conventional oversight versus AIOps-led observability
  • AIOps structure and essential elements

Gathering and Standardizing Operational Information

  • Kinds of observability data: indicators, logs, and traces
  • Importing data from diverse sources (servers, containers, cloud)
  • Utilizing agents and exporters (Prometheus, Beats, Fluentd)

Information Linking and Anomaly Identification

  • Temporal series linking and statistical approaches
  • Applying ML models for spotting anomalies
  • Identifying incidents across decentralized systems

Notification Strategies and Clutter Minimization

  • Creating smart notification rules and limits
  • Muting, removing duplicates, and grouping alerts
  • Connecting with Alertmanager, Slack, PagerDuty, or Opsgenie

Root Cause Investigation and Visual Analysis

  • Using control panels to display indicators and track patterns
  • Analyzing events and timelines for RCA
  • Tracking problems across layers using distributed tracing utilities

Automation and Resolution

  • Activating automated scripts or workflows from incidents
  • Integrating with ITSM platforms (ServiceNow, Jira)
  • Application examples: self-healing, scaling, traffic redirection

Open Source and Commercial AIOps Solutions

  • Overview of tools: Prometheus, Grafana, ELK, Moogsoft, Dynatrace
  • Assessment criteria for choosing an AIOps platform
  • Demonstration and practical work with a chosen stack

Recap and Future Directions

Requirements

  • A solid grasp of IT operations and system oversight principles
  • Practical experience with oversight instruments or control panels
  • Basic knowledge of standard log and indicator structures

Target Audience

  • Operational teams managing infrastructure and applications
  • Site Reliability Engineers (SREs)
  • IT oversight and observability specialists

Number of participants


Price per participant

Upcoming Courses

Related Categories