Agentic AI DevOps Bootcamp · Demo Session · Summary Notes

The story in one line: today you log in, run commands and read dashboards to find out what’s wrong. Tomorrow you build an agent that does it for you, and you simply talk to it.

1. Why this bootcamp? The shift that’s already happening

As a DevOps, platform, cloud or SRE engineer, your day probably looks like this:

  • Log in to wherever the application runs: EC2 instances, Kubernetes pods, Docker containers.
  • Run commands to pull logs, or open one dashboard after another to check status.
  • Every time there’s an issue, you repeat it all: more commands, more dashboards, then you analyze the data yourself to find the reason for the failure.

Once AI enters the picture, three areas start solving this manual work together:

Area What it brings
Generative AI Models that understand and generate language, so you can ask instead of query
AIOps Using AI on operations data: logs, metrics, alerts, root-cause detection
Agentic AI Agents that decide what to do, use your tools, and complete the task

The shift in your role: instead of doing every build, deployment and monitoring task by hand, you build AI agents and deploy them to take care of those deliverables. The DevOps engineer moves from operating the pipeline to building and supervising the agents that operate it. That is the move from a DevOps engineer to an Agentic DevOps engineer, and it is becoming an expected upgrade for these roles.

2. Who is this bootcamp for?

  • Engineers already working in, or who have learned, DevOps, cloud, SRE or platform engineering and want to add AI skills.
  • Platform engineers who want to see how AI helps deploy and run platform infrastructure.
  • Any IT professional exploring a DevOps + AI career path.

What we assume you already know: the basics of Linux, containers, Git and cloud. This is not a fundamentals-of-AI course. We cover the AI and ML basics you need, then go straight to building agents.

3. The learning path

Step What we cover
1. DevOps tool refresher Advanced Kubernetes, AWS, CI/CD, Docker, Jenkins, Terraform and observability, so you know the tools we’re about to connect to AI
2. Generative AI What an LLM is, how it works, prompt engineering, Amazon Bedrock, LangChain, Terraform for IaC
3. AIOps First real use cases: log analysis, root-cause detection, suggested fixes (Prometheus and Grafana integrated with AI agents)
4. Agentic AI Building agents that act autonomously. The backbone is MCP, plus agent frameworks, Python, and harness engineering
5. Capstone projects Building agents from scratch on a real platform
Bonus: MLOps Amazon SageMaker, MLflow and deploying models in containers

Update from Session 2: we start with AI and ML foundations, and we refresh each DevOps tool just before we integrate it into an AI system, rather than doing all the tools up front.

Why MLOps is a bonus, not part of the agentic flow: MLOps is about training, tracking and deploying models. Agentic DevOps is about building agents. They are two different use cases and two different skill sets, so we keep them separate. You’ll likely need both as a DevOps engineer, which is why it’s in the curriculum.

4. How an agentic DevOps platform works

Today your applications run on infrastructure you’ve built with tools like Terraform, EKS, Argo CD, Jenkins, Prometheus and Grafana, and you are the integrator who knows how they all fit together.

In the agentic model, you teach the agent what you know:

The way a DevOps engineer knows all these tools, now the agent knows them: it chooses the right tool, calls it, and executes on your behalf.

The three building blocks to understand

1. Agents. One per responsibility: an SRE agent, a deployment agent, an operations agent. You teach each one how to do its job, the same way you’d explain it to a new colleague.

2. MCP: how agents reach your tools. Each tool gets its own MCP server (a Kubernetes MCP, a Terraform MCP, and so on). The agent goes through the MCP layer to connect to the tool and get things done. MCP is the main backbone of agentic AI, and we’ll spend real time on it.

3. RAG: how agents use knowledge. Where MCP lets an agent act through tools, RAG (Retrieval-Augmented Generation) lets it look up information, such as docs, runbooks and course content, and answer from it. Chatbots and assistants you use every day are built this way. (You’ll see one in the demo.)

The agent loop

This is how an agent “works like a human”:

  1. Understand the intent of your request.
  2. Pick the right tool. Ask “create an EC2 instance on AWS” and it knows to use AWS resources.
  3. Execute the tool and get the result.
  4. Analyze the result.
  5. Report back to you.
  6. For certain actions, ask for human approval first.

Harness engineering

The newest buzzword this year. Harness engineering is how we equip and control the agent, giving it the tools, context, limits and approvals it needs to do its job reliably, like a human colleague. It’s the foundation that makes the agent loop work nicely, and we cover it in detail.

5. The capstone: your own agentic DevOps platform

You will build the work that, today, different people on a DevOps team do:

Agent Takes over
SRE agent Provisioning the entire infrastructure
Deployment agent CI/CD and deployments
Operations agent Monitoring applications and fixing issues
Console One dashboard to see and manage all the agents

These agents connect to the cloud, Kubernetes, Argo CD and the rest of your stack, and deliver like a DevOps engineer would.

6. The tool stack

AWS · Kubernetes · Docker · Terraform · Jenkins · GitHub · Argo CD · Prometheus · Grafana · Amazon Bedrock (LLMs) · Claude · MCP · LangChain and other agent frameworks · Python · Bonus: Amazon SageMaker and MLflow

We focus on the main components and don’t go into every sub-tool.

7. The demo: “Just talk to your cluster”

The big idea: AI doesn’t change what your infrastructure is. The application still runs on a CPU and memory somewhere, just as it did on physical servers, then VMs, then containers, then Kubernetes pods. What changes is how simply you interact with it.

Analogy: Google search gives you a list of results, and you analyze and summarize them yourself. ChatGPT gives you the refined answer directly. Agents do the same for DevOps.

In the demo, a chat client (Claude Desktop) was connected to the cluster through MCP. Here’s what happened:

What I asked What the agent did The point
“Check if Grafana is running on my cluster.” (typo and all) Scanned its available MCP tools and replied: running, which port, private IP, 0 restarts Agents tolerate typos and find the right tool themselves
“List the available namespaces.” Listed them No need to remember kubectl commands. Just chat
“Check if the payment app is running.” Asked permission to use the tool, then found the app deployed but container not ready, restarted ~81 times Instant problem detection
“Why is this failing?” Pulled the right logs, sifted the errors and explained the root cause No manual log digging
“Restart the pod.” Asked for human permission first, then restarted it Destructive actions need approval
“Any CPU spike on the Grafana app?” Queried Prometheus: no spike, with actual numbers Metrics on demand

What’s the manual alternative? Log in to Grafana, find the app, pull the data, analyze it and write the email to your manager. The agent can do all of that, including sending the report, once you build it with those capabilities. That is exactly what we learn to do in this bootcamp.

Who decides when the agent asks for permission?

You do, as the builder. Which actions the agent may run on its own and which need a human’s OK is a design decision. We’ll learn how to build that approval into the agent.

Bonus demo: the EDUeki assistant

A second agent, a chatbot trained on the bootcamp’s information, answered questions about what you’ll learn, the fee and the support details. It’s a RAG-based assistant, the same architecture behind most customer-support chatbots, and we’ll learn how to build one for DevOps use cases.

The point of the demo: you’ll understand the capabilities of agentic AI today, and why upgrading yourself to this skill set matters.

8. Format and what’s included

  • Sessions: three days a week, Tuesday, Wednesday and Thursday, 7 AM IST.
  • Extra doubt-clarification session: once a week or once in two weeks, an open session to unblock you wherever you’re stuck in practice.
  • Duration: roughly 60 days, because we cover DevOps refreshers, GenAI basics, AIOps, MLOps and the capstone projects.
  • Weekly assignments and capstone projects are part of the curriculum.
  • All sessions are recorded. Recordings are uploaded to the same place on the learning portal once you register and log in.
  • Resume help and mock interviews at the end of the course, for those interested.
  • Monthly workshop on agentic AI. The field changes fast, so every month there’s a workshop on the latest updates (a workshop on the RAG platform is coming up). All workshops are included.

9. Your resources

  • Course tracker sheet: the day-by-day agenda of what we cover.
  • Question tracker sheet: drop your questions here, and they’re answered in the question-hour section of the sessions.
  • The EDUeki website: browse courses (many are free), find the Agentic DevOps Engineer course, see the curriculum, and enroll. Create an account and log in to access recordings.
  • Questions or link problems? Email support@edueki.com or WhatsApp +91 99856 38036.

Quick recap

  1. Today’s DevOps is manual and repetitive: logins, commands, dashboards, analysis.
  2. Generative AI + AIOps + Agentic AI together automate that work.
  3. Your role shifts to building and supervising agents.
  4. Agents reach tools through MCP, use knowledge through RAG, and run an agent loop (understand → pick tool → execute → analyze → report), asking for human approval when needed.
  5. Harness engineering is what makes agents reliable.
  6. In the capstone you build SRE, Deployment and Operations agents plus a Console.
  7. MLOps is a separate, bonus track.

Think about it

  1. Which three tasks in your current role would you most like an agent to handle?
  2. Which of those would you want a human to approve first, and why?
  3. What do you think the difference is between giving an agent a tool (MCP) and giving it knowledge (RAG)?

Up next

Session 2: AI & ML foundations. What AI is, how machines learn from data, and the key terms that everything else in the course builds on.

0 Shares:
Leave a Reply

Your email address will not be published. Required fields are marked *

You May Also Like