Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Performance Fundamentals and Key Metrics
- Analyzing latency, throughput, energy consumption, and resource utilization.
- Distinguishing between system-level and model-level constraints.
- Differentiating profiling approaches for inference versus training.
Profiling Techniques on Huawei Ascend
- Leveraging CANN Profiler and MindInsight for detailed analysis.
- Diagnosing issues at the kernel and operator levels.
- Managing offload patterns and memory mapping strategies.
Profiling Techniques on Biren GPU
- Utilizing Biren SDK features for performance monitoring.
- Optimizing kernel fusion, memory alignment, and execution queues.
- Incorporating power and temperature data into profiling workflows.
Profiling Techniques on Cambricon MLU
- Applying BANGPy and Neuware performance tools.
- Interpreting kernel-level visibility and system logs.
- Integrating the MLU profiler with various deployment frameworks.
Graph and Model-Level Enhancement
- Strategies for graph pruning and quantization.
- Techniques for operator fusion and computational graph restructuring.
- Standardizing input sizes and tuning batch parameters.
Memory and Kernel Refinement
- Improving memory layout and data reuse efficiency.
- Managing buffers effectively across different chipsets.
- Platform-specific kernel-level tuning methods.
Cross-Platform Best Practices
- Achieving performance portability through abstraction strategies.
- Developing unified tuning pipelines for multi-chip environments.
- Case study: optimizing an object detection model across Ascend, Biren, and MLU platforms.
Conclusion and Future Directions
Requirements
- Practical experience in AI model training or deployment pipelines.
- A solid grasp of GPU/MLU compute mechanics and model optimization concepts.
- Familiarity with fundamental performance profiling tools and key metrics.
Intended Audience
- Performance engineers.
- Machine learning infrastructure teams.
- AI system architects.
21 Hours