What exactly is 'training'? ML questions my colleagues keep asking me

Last Friday, during lunch, I was eating chicken rice at Maxwell Food Centre. A colleague from the marketing team sitting across from me suddenly asked, "You always talk about 'training' models. What does that actually mean? Does the computer study?" I paused mid-bite. It's something I do every day, but when I tried to explain it, I wasn't sure where to start.

It's actually not the first time I've gotten a question like this. Working in a Singapore office, non-engineer colleagues — PMs, designers, sales — get curious about AI or machine learning pretty often. Since I'm still learning every day myself, I figured writing these questions down would be a good review for me too.

"What do you mean by 'training'?"

This is the question I hear most often. The analogy I like to use goes like this. Imagine a friend visiting a hawker centre for the first time. You show them hundreds of photos of laksa, char kway teow, and kaya toast, telling them, "This is laksa, this is char kway teow." They get it wrong at first, but the more photos they see, the better they get at telling them apart.

Training a machine learning model is similar. You give the model data, tell it the correct answers, and work on reducing the gap between what the model predicts and the actual answers (this is called the loss). This process is repeated thousands, even tens of thousands of times. In code, it looks something like this:

for epoch in range(num_epochs):
    predictions = model(training_data)
    loss = compute_loss(predictions, actual_labels)
    loss.backward()
    optimizer.step()

This short loop is essentially the core of "training." With each iteration, the model slightly adjusts its internal numbers (parameters), trying to make more accurate predictions the next time around.

Conceptual illustration of a machine learning training loop with data flowing through a model and loss decreasing

"Is more data always better?"

This is the second most common question I get. The answer is "yes and no." Once, for an internal project, I built a model to classify customer feedback text. We had around 50,000 data points, which seemed like plenty, but the labelling was a mess. The same sentence would be tagged as "positive" in one case and "neutral" in another. Data quality mattered more than quantity. In the end, three of us spent two days cleaning up the labels before we finally saw a noticeable improvement in performance.

On the other hand, when you have limited data, you can use methods like transfer learning. You take a model that has already been trained on a large dataset and fine-tune it with your smaller dataset. Instead of teaching from scratch, it's like taking a friend who has already seen tons of food photos and giving them a bit more training: "Now try to distinguish these specific types."

"So once training is done, can you use it right away?"

This one comes up a lot too, but the reality isn't that straightforward. A model that has finished training needs to be tested. You evaluate its performance on a separate set of data (a test set) that wasn't used during training. If the test scores are bad, you need to go back and adjust — gather more data, change the model architecture, or tweak the hyperparameters.

In my experience, this cycle of "train, evaluate, adjust" takes up a huge chunk of actual work time. Writing the model code is only part of the job; the rest goes to data cleaning, experiment logging, and analysing results. Filling Jupyter notebooks with detailed experiment logs — leaving notes like "lowered the learning rate from 0.001 to 0.0005 and the validation loss decreased by this much" — is just part of the daily routine.

Jupyter notebook with training experiment logs on a laptop screen beside a cup of kopi

You learn by explaining

Honestly, answering colleagues' questions sometimes exposes areas I only vaguely understood and glossed over. If I can't give a clear answer to "Why did you pick that value for the learning rate?", I end up going home that evening and looking it up again. Teaching really is one of the best ways to learn.

Next time someone asks me "How does AI learn?", I plan to start with the hawker centre analogy. This weekend, I'm going to take my laptop to a newly opened café near Tiong Bahru and do some cleanup on the dataset for an image classification side project I've been working on.

Comments