Agentic AI DevOps Bootcamp · Session 3 · Summary Notes

In one sentence: a machine learning model is just a simple equation whose numbers were found by trial and error against real data. Once you see that, most of ML stops being mysterious.

The big picture

Last session we covered training vs. inference, structured vs. unstructured data, and prediction vs. recommendation systems. The common thread: to predict anything, a machine first learns patterns from historical data, then applies them to new data.

Today we opened the hood and looked at how that learning happens, using one story that builds step by step:

  1. What kinds of problems does ML solve? → Classification and Regression
  2. How does a model learn? → Guess a line, measure the error, adjust, repeat
  3. What is the model really? → y = mx + b
  4. What if there are many inputs? → One weight per input
  5. What if the answer is a category, not a number? → Add an activation function (sigmoid)

1. Two kinds of ML problems

Classification: “Which bucket does this belong to?”

The answer comes from a fixed list of categories (classes).

Example Classes Type
Is this email spam? Spam / Not spam Binary (2 classes)
Is this credit card transaction fraud? Fraud / Not fraud Binary
Which topic is this article about? Business / International / Finance / Banking … Multi-class (3+ classes)
Is this image a car or a bus? Car / Bus Binary (unstructured data, still classification)

Regression: “What is the number?”

The answer is a continuous number, not a pick from a fixed list.

Example Prediction
House price from area, bedrooms, etc. 500k, 700k, 1,000k… any value
Salary from years of experience Any value
Speech rate from an audio file Words per minute

Quick test: Can I list every possible answer in advance? Yes → classification. No, it’s a number on a scale → regression.

Why it matters: the problem type decides how you train the model, so always identify it first.

2. Features, target, and the model

Every training dataset has two parts:

  • Features (inputs): what describes the thing. For a house: area, bedrooms.
  • Target (label): the answer you want to predict. For a house: the price.
Phase What the model gets What it produces
Training Features and target (it learns by seeing the answers) A trained model
Inference Features only A predicted target

Training in practice: you write a Python program (the ML program), feed it the dataset, and run it. The program studies the data, finds the pattern, and saves what it learned as an artifact. That artifact is the ML model.

Analogy: run a Java build and you get a .jar file. Run ML training and you get a model file. Same idea: one file that carries the logic.

Two things follow:

  • A model only works for the problem it was trained on. A house-price model cannot predict loan approvals. One problem, one model.
  • Prediction quality depends on training quality and data quality. No model is 100% correct.

3. How does a machine learn? The salary story

The data: years of experience → salary (in thousands)

Experience 1 2 3 4 5 6
Salary 10 20 30 40 50 60

The question: what is the salary for 8 years?

Your gut says 80. But a machine can’t “just see it”. It must prove it mathematically. Here is how it works:

Step 1: Guess a line

Plot the data and draw a random straight line through it. Read predictions off that line.

Step 2: Measure the error

Compare each prediction with the real answer. The gap is the error (delta).

  • Real salary at 4 years = 40. Our line says 55 → error = 15.
  • Real salary at 6 years = 60. Our line says 80 → error = 20.

Add up the errors across all rows and you get the loss: one number that says how wrong the model is.

  • Loss = 0 → predictions match reality perfectly.
  • Loss > 0 → the line is wrong. Try again.

Step 3: Adjust and repeat

Draw a different line, recompute the predictions, recompute the loss. Keep going until the loss is zero, or as close to zero as possible.

Analogy: exam prep. Cover the answers, attempt the question, then uncover the real answer and compare. A mismatch tells you to study again. A match tells you that you’ve learned it. ML training is exactly this loop.

4. Turning the line into math: y = mx + b

A program can’t look at graphs, so the line becomes an equation:

y = m × x + b
Symbol Name Meaning In our example
y Target What we predict Salary
x Input feature What we know Years of experience
m Slope (also called weight / coefficient) How much y changes when x increases by 1 10
b Bias (also called intercept) The value of y when x = 0 0

How we got the numbers:
– Slope: pick any two points, then (y₂ − y₁) / (x₂ − x₁). For example (60 − 50) / (6 − 5) = 10.
– Bias: where the line crosses the y-axis. Our first random line crossed at about 3–4 (wrong); the perfect line passes through the origin, so b = 0.

Making a prediction (8 years):

y = 10 × 8 + 0 = 80

10 years? 10 × 10 + 0 = 100.

The key insight

m and b are the learnable parameters. Training is simply the search for the best values of m and b. Once found, they are saved inside the model. Inference is then just plugging new inputs into the saved equation.

In this simple example the slope is fixed by the data, so the search is mostly about adjusting b. In real models, all parameters are adjusted together.

5. Gradient descent and epochs (two terms to know)

  • Gradient descent: the technique of repeatedly nudging the parameters in the direction that reduces the loss, until the loss is as low as it can get. It’s the “guess, measure, adjust” loop from above, done systematically.
  • Epoch: one round of that training loop. If the model needed 10 rounds to find good values, that’s 10 epochs. We can’t know in advance how many are needed.

Will the loss always reach zero? Not necessarily. It depends on the data. The goal is as close to zero as we can get, and we improve results through better data, tuning and more training.

6. Many inputs: multiple features

Real problems rarely depend on one thing. Salary also depends on education, projects, certifications. House prices depend on area, bedrooms, and neighbourhood safety, nearby shops, and more.

The equation stays linear. You just have more xs, each with its own weight:

y = w₁·x₁ + w₂·x₂ + w₃·x₃ + w₄·x₄ + b
  • Every feature gets its own weight (w, the ML word for slope m).
  • There is still only one bias b for the whole dataset.
  • Text-like inputs (e.g., education level) are converted to numbers first.

Your job as the engineer: include every input you think could influence the target. If it isn’t in the training data, the model can’t learn it and can’t use it at inference time. (This also answers the “what about qualitative factors like a safe neighbourhood?” question: turn them into features.)

7. Training vs. inference: the full picture

Training Inference
Input Features (x) and answers (y) Features (x) only
What happens Gradient descent searches for the best weights and bias Plug x into the saved equation
Output A model containing the learned w and b A prediction
Where quality is decided Here Nowhere. It just does the math

Big takeaway: inference does no learning. If predictions are wrong, the problem is in the training: bad data, missing features, or too little training. Go back and improve training, not inference.

8. When the answer is a category: classification math

Example: loan approval. Input: credit score. Output: 0 = approved, 1 = rejected.

Problem 1: the data isn’t a neat straight line, because it’s all 0s and 1s.

Problem 2: if you run y = mx + b with m = 10, b = 0, x = 300, you get 3000. What does 3000 mean? Nothing. We need 0 or 1.

Fix: add an activation function on top. For classification we still compute mx + b, then pass the result through an activation function that squashes it into the range 0 to 1.

The one we used is the sigmoid:

sigmoid(z) = 1 / (1 + e^(−z))      where z = mx + b

Reading the result: the decision boundary at 0.5

Sigmoid output Interpreted as Loan decision
0 → 0.5 Class 0 Approved
0.5 → 1 Class 1 Rejected

So the plot of a classifier is an S-shaped sigmoid curve, not a straight line.

Remember: regression = mx + b. Classification = mx + b + activation function.

Key terms cheat sheet

Term Plain-English meaning
Feature An input that describes the thing (area, experience, credit score)
Target / Label The answer we want to predict
Classification Predict a category from a fixed set
Regression Predict a continuous number
Model The saved result of training: the learned parameters plus the equation
Loss A single number measuring how wrong the predictions are
Slope (m) / Weight (w) How strongly an input influences the output
Bias (b) / Intercept The output when all inputs are zero
Gradient descent Repeatedly adjusting parameters to reduce loss
Epoch One round of the training loop
Activation function Converts a raw number into a usable range (e.g., sigmoid → 0 to 1)

Quick recap

  1. ML problems are classification (pick a category) or regression (predict a number).
  2. Training data has features + target; inference has features only.
  3. A model learns by guessing, measuring the loss, and adjusting until the loss is near zero (gradient descent).
  4. The model boils down to y = mx + b: the learned m and b are saved in the model file.
  5. Many features → one weight per feature, one bias.
  6. Wrong predictions mean fix the training, not the inference.
  7. For classification, add an activation function like sigmoid with a 0.5 decision boundary.

Check yourself

  1. Is “predict tomorrow’s CPU usage %” classification or regression? What about “will this deployment fail: yes/no”?
  2. In y = mx + b, which values are learned during training, and where are they stored?
  3. Your model gives poor predictions in production. Is the issue in training or inference? Why?
  4. Why can’t a plain mx + b output decide “approved vs. rejected” on its own? What fixes it?
  5. What does a loss of zero mean?
Answers
  1. CPU % is regression (continuous number). Deployment fail yes/no is binary classification.
  2. m (weights) and b (bias). They are saved in the model artifact.
  3. Training. Inference only applies the saved values.
  4. It outputs any number (like 3000), not a class. An activation function (sigmoid) squashes it to 0–1, and a 0.5 boundary picks the class.
  5. Predictions exactly match the real answers in the training data.

Why are we learning this for an Agentic DevOps course?

Honest answer from the session: agentic AI work does not use these ML internals directly. When you build agents, you call an LLM and write code around it. But:

  • This is the foundation for understanding how machines learn and “think”. Without it, you’re running agents without knowing what’s under the hood.
  • It’s directly useful for MLOps, which we cover at the end of the course. Training, models as artifacts, and retraining all build on today’s ideas.
  • As DevOps engineers, you’ll be asked to operate ML systems, not just agents.

Up next

Hands-on Python. We turn today’s theory into code: training and using a real model.

0 Shares:
Leave a Reply

Your email address will not be published. Required fields are marked *

You May Also Like