Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Duration 14 hours
Course Outline
Introduction to AIOps with Open-Source Tools
- Overview of AIOps concepts and their advantages
- The role of Prometheus and Grafana within the observability stack
- The place of ML in AIOps: comparing predictive and reactive analytics
Configuring Prometheus and Grafana
- Installing and setting up Prometheus for time series data collection
- Building dashboards in Grafana using live metrics
- Investigating exporters, relabeling techniques, and service discovery
Preprocessing Data for ML
- Retrieving and modifying Prometheus metrics
- Curating datasets for anomaly detection and forecasting tasks
- Leveraging Grafana transformations or Python-based pipelines
Utilizing ML for Anomaly Detection
- Implementing basic ML models for outlier identification (e.g., Isolation Forest, One-Class SVM)
- Training and assessing models on time series data
- Displaying detected anomalies within Grafana dashboards
Forecasting Metrics via ML
- Developing simple forecasting models (Introduction to ARIMA, Prophet, LSTM)
- Anticipating system load or resource consumption
- Utilizing forecasts for early warnings and scaling strategies
Connecting ML with Alerting and Automation
- Establishing alert rules driven by ML outputs or specific thresholds
- Employing Alertmanager and managing notification routing
- Initiating scripts or automation workflows upon anomaly detection
Scaling and Operationalizing AIOps
- Connecting external observability tools (e.g., ELK stack, Moogsoft, Dynatrace)
- Integrating ML models into observability workflows
- Best practices for implementing AIOps at scale
Recap and Future Directions
Requirements
- A solid grasp of system monitoring and observability principles
- Practical experience with Grafana or Prometheus
- Proficiency in Python and an understanding of fundamental machine learning concepts
Target Audience
- Observability engineers
- Infrastructure and DevOps teams
- Monitoring platform architects and site reliability engineers (SREs)