Get in Touch

Course Outline

Foundations of Agentic AI in Operations

  • Transitioning from static runbooks to reasoning agents: the progression of IT automation
  • Anatomy of an agent: reasoning loops, tool utilization, memory, and planning
  • Strategic decision-making: when to automate vs. when to retain human oversight

Agent Frameworks and System Architectures

  • Single-agent methodologies: ReAct, Plan-and-Execute, and tool-invocation cycles
  • Multi-agent structures: supervisor models, hierarchical designs, and swarm patterns
  • Framework analysis: LangGraph, CrewAI, AutoGen, and custom-built agents
  • Developing your first operational agent: querying monitors, diagnosing issues, and proposing solutions

Tool Integration for IT Operations

  • Linking agents to Prometheus, Grafana, Datadog, and PagerDuty interfaces
  • Agent-driven log analysis: integrating Elasticsearch, Loki, and Splunk
  • Leveraging infrastructure tools: executing kubectl, Terraform, and Ansible via agent actions
  • Crafting secure tool interfaces with parameter validation and idempotency

Automating Incident Response

  • Automated triage: classifying severity and routing incidents
  • Generating root cause hypotheses and collecting supporting evidence
  • Automated remediation: executing restart, scaling, rollback, and failover tasks
  • Creating an incident runbook agent with tiered autonomy levels

Safety, Guardrails, and Human Oversight

  • Classifying actions: read-only, low-risk, high-risk, and destructive
  • Defining approval gates and escalation protocols for critical tasks
  • Implementing guardrail patterns: action allowlists, blast radius constraints, and rollback assurances
  • Establishing audit trails and decision provenance for compliance

Multi-Agent Orchestration for Complex Scenarios

  • Coordinating specialist agents: triage, diagnosis, and remediation roles
  • Managing inter-agent communication and shared context
  • Resolving conflicts when agents suggest opposing actions
  • Simulating end-to-end major incidents with multi-agent responses

Observability and Performance Evaluation

  • Tracing agent reasoning paths for debugging and auditing
  • Assessing decision quality: precision, recall, and resolution time
  • Establishing feedback loops: learning from operator overrides and outcomes
  • Monitoring costs and analyzing token economics for operational agents

Production Deployment and Maintenance

  • Deploying agents as services: leveraging APIs, webhooks, and scheduled tasks
  • Phased autonomy rollout: from shadow mode to full auto-remediation
  • Developing runbooks for agent failures: handling breakdowns in the agent system
  • Building the business case and calculating ROI for autonomous operations

Requirements

  • Practical experience in IT operations, DevOps, or SRE workflows.
  • Proficiency in Python scripting and RESTful APIs.
  • Fundamental knowledge of Large Language Model (LLM) capabilities and prompt engineering.

Target Audience

  • SRE and DevOps engineers investigating AI-driven automation strategies.
  • Platform engineers developing self-healing infrastructure solutions.
  • IT operations leaders assessing agentic AI for incident management.
 14 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories