Get in Touch

Course Outline

Performance Fundamentals and Key Metrics

  • Analyzing latency, throughput, power consumption, and resource utilization
  • Distinguishing between system-level and model-level constraints
  • Differentiating profiling approaches for inference versus training

Profiling Techniques on Huawei Ascend

  • Leveraging CANN Profiler and MindInsight
  • Diagnosing kernels and operators
  • Managing offload patterns and memory mapping

Profiling Techniques on Biren GPU

  • Utilizing Biren SDK for performance monitoring
  • Optimizing kernel fusion, memory alignment, and execution queues
  • Conducting power and temperature-aware analysis

Profiling Techniques on Cambricon MLU

  • Employing BANGPy and Neuware performance tools
  • Gaining kernel-level visibility and interpreting logs
  • Integrating the MLU profiler with deployment frameworks

Graph and Model-Level Enhancements

  • Strategies for graph pruning and quantization
  • Restructuring computational graphs and fusing operators
  • Standardizing input sizes and tuning batch processing

Memory and Kernel Refinement

  • Optimizing memory layouts and maximizing reuse
  • Managing buffers efficiently across different chipsets
  • Applying specific kernel-level tuning techniques for each platform

Cross-Platform Best Practices

  • Achieving performance portability through abstraction strategies
  • Establishing shared tuning pipelines for multi-chip setups
  • Case study: Optimizing an object detection model across Ascend, Biren, and MLU

Conclusion and Future Directions

Requirements

  • Practical experience with AI model training or deployment workflows
  • Solid grasp of GPU/MLU computing principles and model optimization strategies
  • Familiarity with fundamental performance profiling tools and key metrics

Target Audience

  • Performance engineers
  • Machine learning infrastructure teams
  • AI system architects
 21 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories