Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
Foundations of Agentic AI in Operations
- Transitioning from static runbooks to reasoning agents: the progression of IT automation
- Anatomy of an agent: reasoning loops, tool utilization, memory, and planning
- Strategic decision-making: when to automate vs. when to retain human oversight
Agent Frameworks and System Architectures
- Single-agent methodologies: ReAct, Plan-and-Execute, and tool-invocation cycles
- Multi-agent structures: supervisor models, hierarchical designs, and swarm patterns
- Framework analysis: LangGraph, CrewAI, AutoGen, and custom-built agents
- Developing your first operational agent: querying monitors, diagnosing issues, and proposing solutions
Tool Integration for IT Operations
- Linking agents to Prometheus, Grafana, Datadog, and PagerDuty interfaces
- Agent-driven log analysis: integrating Elasticsearch, Loki, and Splunk
- Leveraging infrastructure tools: executing kubectl, Terraform, and Ansible via agent actions
- Crafting secure tool interfaces with parameter validation and idempotency
Automating Incident Response
- Automated triage: classifying severity and routing incidents
- Generating root cause hypotheses and collecting supporting evidence
- Automated remediation: executing restart, scaling, rollback, and failover tasks
- Creating an incident runbook agent with tiered autonomy levels
Safety, Guardrails, and Human Oversight
- Classifying actions: read-only, low-risk, high-risk, and destructive
- Defining approval gates and escalation protocols for critical tasks
- Implementing guardrail patterns: action allowlists, blast radius constraints, and rollback assurances
- Establishing audit trails and decision provenance for compliance
Multi-Agent Orchestration for Complex Scenarios
- Coordinating specialist agents: triage, diagnosis, and remediation roles
- Managing inter-agent communication and shared context
- Resolving conflicts when agents suggest opposing actions
- Simulating end-to-end major incidents with multi-agent responses
Observability and Performance Evaluation
- Tracing agent reasoning paths for debugging and auditing
- Assessing decision quality: precision, recall, and resolution time
- Establishing feedback loops: learning from operator overrides and outcomes
- Monitoring costs and analyzing token economics for operational agents
Production Deployment and Maintenance
- Deploying agents as services: leveraging APIs, webhooks, and scheduled tasks
- Phased autonomy rollout: from shadow mode to full auto-remediation
- Developing runbooks for agent failures: handling breakdowns in the agent system
- Building the business case and calculating ROI for autonomous operations
Requirements
- Practical experience in IT operations, DevOps, or SRE workflows.
- Proficiency in Python scripting and RESTful APIs.
- Fundamental knowledge of Large Language Model (LLM) capabilities and prompt engineering.
Target Audience
- SRE and DevOps engineers investigating AI-driven automation strategies.
- Platform engineers developing self-healing infrastructure solutions.
- IT operations leaders assessing agentic AI for incident management.
14 Hours