Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Performance Fundamentals and Key Metrics
- Analyzing latency, throughput, power consumption, and resource utilization
- Distinguishing between system-level and model-level constraints
- Differentiating profiling approaches for inference versus training
Profiling Techniques on Huawei Ascend
- Leveraging CANN Profiler and MindInsight
- Diagnosing kernels and operators
- Managing offload patterns and memory mapping
Profiling Techniques on Biren GPU
- Utilizing Biren SDK for performance monitoring
- Optimizing kernel fusion, memory alignment, and execution queues
- Conducting power and temperature-aware analysis
Profiling Techniques on Cambricon MLU
- Employing BANGPy and Neuware performance tools
- Gaining kernel-level visibility and interpreting logs
- Integrating the MLU profiler with deployment frameworks
Graph and Model-Level Enhancements
- Strategies for graph pruning and quantization
- Restructuring computational graphs and fusing operators
- Standardizing input sizes and tuning batch processing
Memory and Kernel Refinement
- Optimizing memory layouts and maximizing reuse
- Managing buffers efficiently across different chipsets
- Applying specific kernel-level tuning techniques for each platform
Cross-Platform Best Practices
- Achieving performance portability through abstraction strategies
- Establishing shared tuning pipelines for multi-chip setups
- Case study: Optimizing an object detection model across Ascend, Biren, and MLU
Conclusion and Future Directions
Requirements
- Practical experience with AI model training or deployment workflows
- Solid grasp of GPU/MLU computing principles and model optimization strategies
- Familiarity with fundamental performance profiling tools and key metrics
Target Audience
- Performance engineers
- Machine learning infrastructure teams
- AI system architects
21 Hours