Agentic AI DevOps Bootcamp · Session 3 · Summary Notes
In one sentence: a machine learning model is just a simple equation whose numbers were found by trial and error against real data. Once you see that, most of ML stops being mysterious.
The big picture
Last session we covered training vs. inference, structured vs. unstructured data, and prediction vs. recommendation systems. The common thread: to predict anything, a machine first learns patterns from historical data, then applies them to new data.
Today we opened the hood and looked at how that learning happens, using one story that builds step by step:
- What kinds of problems does ML solve? → Classification and Regression
- How does a model learn? → Guess a line, measure the error, adjust, repeat
- What is the model really? →
y = mx + b - What if there are many inputs? → One weight per input
- What if the answer is a category, not a number? → Add an activation function (sigmoid)
1. Two kinds of ML problems
Classification: “Which bucket does this belong to?”
The answer comes from a fixed list of categories (classes).
| Example | Classes | Type |
|---|---|---|
| Is this email spam? | Spam / Not spam | Binary (2 classes) |
| Is this credit card transaction fraud? | Fraud / Not fraud | Binary |
| Which topic is this article about? | Business / International / Finance / Banking … | Multi-class (3+ classes) |
| Is this image a car or a bus? | Car / Bus | Binary (unstructured data, still classification) |
Regression: “What is the number?”
The answer is a continuous number, not a pick from a fixed list.
| Example | Prediction |
|---|---|
| House price from area, bedrooms, etc. | 500k, 700k, 1,000k… any value |
| Salary from years of experience | Any value |
| Speech rate from an audio file | Words per minute |
Quick test: Can I list every possible answer in advance? Yes → classification. No, it’s a number on a scale → regression.
Why it matters: the problem type decides how you train the model, so always identify it first.
2. Features, target, and the model
Every training dataset has two parts:
- Features (inputs): what describes the thing. For a house: area, bedrooms.
- Target (label): the answer you want to predict. For a house: the price.
| Phase | What the model gets | What it produces |
|---|---|---|
| Training | Features and target (it learns by seeing the answers) | A trained model |
| Inference | Features only | A predicted target |
Training in practice: you write a Python program (the ML program), feed it the dataset, and run it. The program studies the data, finds the pattern, and saves what it learned as an artifact. That artifact is the ML model.
Analogy: run a Java build and you get a
.jarfile. Run ML training and you get a model file. Same idea: one file that carries the logic.
Two things follow:
- A model only works for the problem it was trained on. A house-price model cannot predict loan approvals. One problem, one model.
- Prediction quality depends on training quality and data quality. No model is 100% correct.
3. How does a machine learn? The salary story
The data: years of experience → salary (in thousands)
| Experience | 1 | 2 | 3 | 4 | 5 | 6 |
|---|---|---|---|---|---|---|
| Salary | 10 | 20 | 30 | 40 | 50 | 60 |
The question: what is the salary for 8 years?
Your gut says 80. But a machine can’t “just see it”. It must prove it mathematically. Here is how it works:
Step 1: Guess a line
Plot the data and draw a random straight line through it. Read predictions off that line.
Step 2: Measure the error
Compare each prediction with the real answer. The gap is the error (delta).
- Real salary at 4 years = 40. Our line says 55 → error = 15.
- Real salary at 6 years = 60. Our line says 80 → error = 20.
Add up the errors across all rows and you get the loss: one number that says how wrong the model is.
- Loss = 0 → predictions match reality perfectly.
- Loss > 0 → the line is wrong. Try again.
Step 3: Adjust and repeat
Draw a different line, recompute the predictions, recompute the loss. Keep going until the loss is zero, or as close to zero as possible.
Analogy: exam prep. Cover the answers, attempt the question, then uncover the real answer and compare. A mismatch tells you to study again. A match tells you that you’ve learned it. ML training is exactly this loop.
4. Turning the line into math: y = mx + b
A program can’t look at graphs, so the line becomes an equation:
y = m × x + b
| Symbol | Name | Meaning | In our example |
|---|---|---|---|
| y | Target | What we predict | Salary |
| x | Input feature | What we know | Years of experience |
| m | Slope (also called weight / coefficient) | How much y changes when x increases by 1 | 10 |
| b | Bias (also called intercept) | The value of y when x = 0 | 0 |
How we got the numbers:
– Slope: pick any two points, then (y₂ − y₁) / (x₂ − x₁). For example (60 − 50) / (6 − 5) = 10.
– Bias: where the line crosses the y-axis. Our first random line crossed at about 3–4 (wrong); the perfect line passes through the origin, so b = 0.
Making a prediction (8 years):
y = 10 × 8 + 0 = 80
10 years? 10 × 10 + 0 = 100.
The key insight
m and b are the learnable parameters. Training is simply the search for the best values of m and b. Once found, they are saved inside the model. Inference is then just plugging new inputs into the saved equation.
In this simple example the slope is fixed by the data, so the search is mostly about adjusting
b. In real models, all parameters are adjusted together.
5. Gradient descent and epochs (two terms to know)
- Gradient descent: the technique of repeatedly nudging the parameters in the direction that reduces the loss, until the loss is as low as it can get. It’s the “guess, measure, adjust” loop from above, done systematically.
- Epoch: one round of that training loop. If the model needed 10 rounds to find good values, that’s 10 epochs. We can’t know in advance how many are needed.
Will the loss always reach zero? Not necessarily. It depends on the data. The goal is as close to zero as we can get, and we improve results through better data, tuning and more training.
6. Many inputs: multiple features
Real problems rarely depend on one thing. Salary also depends on education, projects, certifications. House prices depend on area, bedrooms, and neighbourhood safety, nearby shops, and more.
The equation stays linear. You just have more xs, each with its own weight:
y = w₁·x₁ + w₂·x₂ + w₃·x₃ + w₄·x₄ + b
- Every feature gets its own weight (
w, the ML word for slopem). - There is still only one bias
bfor the whole dataset. - Text-like inputs (e.g., education level) are converted to numbers first.
Your job as the engineer: include every input you think could influence the target. If it isn’t in the training data, the model can’t learn it and can’t use it at inference time. (This also answers the “what about qualitative factors like a safe neighbourhood?” question: turn them into features.)
7. Training vs. inference: the full picture
| Training | Inference | |
|---|---|---|
| Input | Features (x) and answers (y) | Features (x) only |
| What happens | Gradient descent searches for the best weights and bias | Plug x into the saved equation |
| Output | A model containing the learned w and b |
A prediction |
| Where quality is decided | Here | Nowhere. It just does the math |
Big takeaway: inference does no learning. If predictions are wrong, the problem is in the training: bad data, missing features, or too little training. Go back and improve training, not inference.
8. When the answer is a category: classification math
Example: loan approval. Input: credit score. Output: 0 = approved, 1 = rejected.
Problem 1: the data isn’t a neat straight line, because it’s all 0s and 1s.
Problem 2: if you run y = mx + b with m = 10, b = 0, x = 300, you get 3000. What does 3000 mean? Nothing. We need 0 or 1.
Fix: add an activation function on top. For classification we still compute mx + b, then pass the result through an activation function that squashes it into the range 0 to 1.
The one we used is the sigmoid:
sigmoid(z) = 1 / (1 + e^(−z)) where z = mx + b
Reading the result: the decision boundary at 0.5
| Sigmoid output | Interpreted as | Loan decision |
|---|---|---|
| 0 → 0.5 | Class 0 | Approved |
| 0.5 → 1 | Class 1 | Rejected |
So the plot of a classifier is an S-shaped sigmoid curve, not a straight line.
Remember: regression =
mx + b. Classification =mx + b+ activation function.
Key terms cheat sheet
| Term | Plain-English meaning |
|---|---|
| Feature | An input that describes the thing (area, experience, credit score) |
| Target / Label | The answer we want to predict |
| Classification | Predict a category from a fixed set |
| Regression | Predict a continuous number |
| Model | The saved result of training: the learned parameters plus the equation |
| Loss | A single number measuring how wrong the predictions are |
| Slope (m) / Weight (w) | How strongly an input influences the output |
| Bias (b) / Intercept | The output when all inputs are zero |
| Gradient descent | Repeatedly adjusting parameters to reduce loss |
| Epoch | One round of the training loop |
| Activation function | Converts a raw number into a usable range (e.g., sigmoid → 0 to 1) |
Quick recap
- ML problems are classification (pick a category) or regression (predict a number).
- Training data has features + target; inference has features only.
- A model learns by guessing, measuring the loss, and adjusting until the loss is near zero (gradient descent).
- The model boils down to
y = mx + b: the learnedmandbare saved in the model file. - Many features → one weight per feature, one bias.
- Wrong predictions mean fix the training, not the inference.
- For classification, add an activation function like sigmoid with a 0.5 decision boundary.
Check yourself
- Is “predict tomorrow’s CPU usage %” classification or regression? What about “will this deployment fail: yes/no”?
- In
y = mx + b, which values are learned during training, and where are they stored? - Your model gives poor predictions in production. Is the issue in training or inference? Why?
- Why can’t a plain
mx + boutput decide “approved vs. rejected” on its own? What fixes it? - What does a loss of zero mean?
Answers
- CPU % is regression (continuous number). Deployment fail yes/no is binary classification.
m(weights) andb(bias). They are saved in the model artifact.- Training. Inference only applies the saved values.
- It outputs any number (like 3000), not a class. An activation function (sigmoid) squashes it to 0–1, and a 0.5 boundary picks the class.
- Predictions exactly match the real answers in the training data.
Why are we learning this for an Agentic DevOps course?
Honest answer from the session: agentic AI work does not use these ML internals directly. When you build agents, you call an LLM and write code around it. But:
- This is the foundation for understanding how machines learn and “think”. Without it, you’re running agents without knowing what’s under the hood.
- It’s directly useful for MLOps, which we cover at the end of the course. Training, models as artifacts, and retraining all build on today’s ideas.
- As DevOps engineers, you’ll be asked to operate ML systems, not just agents.
Up next
Hands-on Python. We turn today’s theory into code: training and using a real model.
