Get in Touch
 Duration 14 hours

Course Outline

Preparing Machine Learning Models for Deployment

  • Packaging models using Docker
  • Exporting models from TensorFlow and PyTorch
  • Strategic considerations for versioning and storage

Model Serving on Kubernetes

  • An overview of inference servers
  • Deploying TensorFlow Serving and TorchServe
  • Configuring model endpoints

Inference Optimization Techniques

  • Implementing batching strategies
  • Managing concurrent request handling
  • Tuning for optimal latency and throughput

Autoscaling ML Workloads

  • Utilizing the Horizontal Pod Autoscaler (HPA)
  • Applying the Vertical Pod Autoscaler (VPA)
  • Using Kubernetes Event-Driven Autoscaling (KEDA)

GPU Provisioning and Resource Management

  • Configuration of GPU nodes
  • Overview of the NVIDIA device plugin
  • Setting resource requests and limits for ML workloads

Model Rollout and Release Strategies

  • Executing blue/green deployments
  • Implementing canary rollout patterns
  • Conducting A/B testing for model evaluation

Monitoring and Observability for ML in Production

  • Tracking metrics for inference workloads
  • Adhering to logging and tracing best practices
  • Establishing dashboards and alerting mechanisms

Security and Reliability Considerations

  • Securing model endpoints
  • Enforcing network policies and access control
  • Safeguarding high availability

Summary and Next Steps

Requirements

  • A solid grasp of containerized application workflows
  • Practical experience with Python-based machine learning models
  • Foundational knowledge of Kubernetes core concepts

Target Audience

  • ML engineers
  • DevOps engineers
  • Platform engineering teams

Number of participants


Price per participant

Testimonials (4)

Upcoming Courses

Related Categories