Pushpjeet Cholkar

Blog

What to Learn in AI/ML: A Practical Roadmap, Stage by Stage

October 3, 2026 · Pushpjeet Cholkar

By Pushpjeet Cholkar: 17 years of corporate experience, starting in databases and now working in cloud, AI and ML.

“What should I learn in AI/ML, and in what order?” Engineers and students in my batches ask me this often. Most have already watched plenty of videos. What they lack is an order, a way to know when a topic is “good enough”, and a project that proves it.

This post is that order. It is a map, not a course: each stage links to the tutorial on this site that covers it in depth.

The short answer (OPINION): learn Python, SQL and only the maths you will actually use. Then learn classic machine learning on tabular data, the basics of deep learning, and then LLMs, RAG and agents. Finish with deploying and running models on the cloud. Put off training large models from scratch until a job needs it.

The AI/ML learning roadmap at a glance

StageWhat to learn“Good enough” meansProject that proves itOn this site
0. Working setupLinux shell, Git, Python environmentsYou can rebuild your setup on a new machine without helpPush a small, reproducible repoLinux from Scratch
1. FoundationsPython, pandas, SQL, practical mathsYou can clean and explore a messy dataset and explain the numbersData-quality report on a public datasetAI & ML Foundations, AI & ML Quest
2. Classic MLscikit-learn, validation, metrics, regression, treesYou beat a simple baseline honestly and can say whyTabular prediction model with a written evaluationLinear Regression: Beginner and Expert, Bias-variance video
3. Deep learning basicsPyTorch, training loops, embeddings, transfer learningYou can fine-tune a pretrained model and read its loss curvesSmall image or text classifierGANs, VAEs and Latent Space
4. GenAI, RAG, agentsLLM APIs, prompting, embeddings, retrieval, evaluation, tool useYour answers are grounded, measured and safe to show a userRAG assistant over your own documents, with a test setRAG Tutorial in Python, AI Agents
5. Cloud and MLOpsPackaging, endpoints, IAM, cost, monitoring, pipelinesSomeone else can deploy, roll back and tear down your modelModel or RAG app deployed on AWS with monitoringML Solutions Quest, AWS training

Stages 0 and 1 can overlap, as can 4 and 5 if you come from cloud. Don’t skip Stage 2: everything after it assumes you can tell whether a model is good.

Stage 0: A working setup

What to learn: the Linux command line, Git, and Python virtual environments (venv or conda, pick one). Add enough Bash to run a script on a remote machine and read its logs.

Why: real ML work runs on Linux servers and containers you reach through a terminal. Skip this and you lose hours to broken environments.

Good enough: you can set up a clean environment from a requirements.txt, run a notebook or script on a cloud VM, and move files back and forth.

Prove it: a small Git repository that anyone can clone and run with two commands.

On this site: Linux from Scratch covers the terminal, files, permissions, processes and scripting, with labs.

Stage 1: Foundations (Python, data handling, SQL, and maths that matters)

What to learn:

Why (OPINION): most ML failures I have seen in projects were data failures, not model failures.

Good enough: you can take an unfamiliar dataset and quickly say what is missing, what is skewed, and what would leak the answer into the features.

Prove it: write a one-page data-quality report on a public dataset, covering row counts, nulls, duplicates, distributions and three questions you would ask the data owner.

On this site: start with the AI & ML Foundations study guide for the mental model of AI vs ML vs deep learning vs generative AI. The AI & ML Quest is a gamified warm-up on the same ideas. The RAG tutorial works through dot products and cosine similarity by hand, which is the clearest reason to learn vectors.

Stage 2: Classic machine learning

What to learn: supervised learning on tabular data with scikit-learn. That means train/validation/test splits, cross-validation, linear and logistic regression, decision trees, random forests and gradient-boosted trees. It also means choosing metrics (accuracy, precision and recall for classification; RMSE and R² for regression), handling class imbalance, and spotting data leakage. Add a little unsupervised learning, like clustering, so you recognise the problem type.

Why (OPINION): many business ML problems are still tabular: churn, demand, fraud, risk, pricing. On those problems a well-validated tree model or regression is often the right answer, and it is cheaper to run and easier to explain than a neural network. This stage also teaches the habit the rest of the roadmap depends on, which is honest evaluation.

Good enough: you can explain why a model overfits or underfits, pick a metric that matches the business cost of an error, and always compare against a simple baseline before saying a model “works”.

Prove it: build a churn or price-prediction model on a public dataset. Include a baseline, cross-validation, a confusion matrix or residual plot, and a short paragraph on where the model fails.

On this site:

Stage 3: Deep learning basics

What to learn: how a neural network trains: forward pass, loss, backpropagation and an optimiser step. Learn one framework; I suggest PyTorch. Learn transfer learning, which means starting from a pretrained model instead of training from scratch. Learn what embeddings are, and the basic idea of attention and transformers, since LLMs are built on them.

Why: you won’t invent new architectures, but you must know what “fine-tuning”, “embedding” and “context window” mean in practice, or every GenAI decision becomes guesswork.

Good enough: you can fine-tune a small pretrained model on your own labelled data. You can read a training and validation loss curve and tell whether it is learning, overfitting or broken.

Prove it: fine-tune a small pretrained image or text classifier on a few hundred labelled examples. Report how it does on held-out data against your Stage 2 model, if the task allows.

On this site: GANs, VAEs and Latent Space explains latent space and generative models, with a video lesson. Read it once you’re comfortable with training basics.

Stage 4: Generative AI (LLMs, RAG and agents)

What to learn, in this order:

  1. Calling an LLM through an API: tokens, context limits, temperature, system prompts, structured (JSON) output, and cost per call.
  2. Prompting as engineering: versioned prompts and test cases, not one-off clever wording.
  3. Embeddings and retrieval: chunking, vector search and similarity.
  4. RAG: retrieving your own documents so the model answers from them, then measuring retrieval quality.
  5. Agents: models that call tools in a loop. This includes tool design, permissions, human approval for risky actions, and guardrails.
  6. Evaluation and safety across all of the above: test sets, failure cases, prompt injection and personal data.

Why (OPINION): most teams build on foundation models rather than training them, so the valuable skills are grounding, evaluating and securing them. Coming from databases, I see RAG as a data problem first: if retrieval is poor, no model choice fixes the answers.

Good enough: your assistant answers from your documents, says “I don’t know” when it should, and you can show a retrieval hit rate on a fixed test set. For agents, tool calls are logged and side-effecting actions need approval.

Prove it: build a RAG assistant over a set of documents you know well, such as runbooks, policies or product docs. Write twenty test questions with expected sources, measure retrieval, change one thing (chunk size, for example) and measure again.

On this site:

Stage 5: Deploying and running ML on the cloud (MLOps)

What to learn: packaging a model behind an API, containers, and the difference between batch and real-time inference. Learn identity and access (who can call the model, and what can the model reach), logging and monitoring, data and model drift, model versioning, pipelines, and cost control. On AWS, the main services to know are:

Why (OPINION): a model that only runs in your notebook gives the business nothing. In my consulting work, the hard part is rarely the model; it is security, cost and keeping it working after go-live.

Good enough: someone else can deploy your model from your repository, roll it back, see when it degrades, and delete everything so no forgotten endpoint keeps billing.

Prove it: deploy your Stage 2 model or your Stage 4 RAG app on AWS with least-privilege IAM, a monitoring dashboard or alarm, and a written teardown script. Then run the teardown.

On this site: the ML Solutions Quest walks through the ML lifecycle, SageMaker AI, evaluation, deployment options and MLOps. If cloud itself is the gap, the AWS Solutions Architect training covers the architecture foundation that ML workloads sit on.

What to skip or postpone (OPINION)

Your route by background

The stages are the same for everyone; where you start and what you can skim are not. This is how I advise people in corporate batches.

Cloud or DevOps engineer

Your advantage: Stages 0 and 5 are already mostly yours. You know IAM, containers, CI/CD, monitoring and cost.

Your gap: statistics and evaluation. Spend real time on Stages 1 and 2. Without them you will deploy models you can’t judge.

Fastest route: Stage 1 maths and pandas → Stage 2 → Stage 4 (RAG on Bedrock) → Stage 5, where you can become the person who makes ML work in production. Read the ML Solutions Quest early; its MLOps half will feel familiar.

DBA or SQL person

Your advantage: SQL, data quality, indexing, and the reliability mindset that comes from running production systems. This was my own route.

Your gap: Python and pandas, and moving from “is the data correct?” to “is this data predictive?”

Fastest route: Stage 1 Python (map pandas operations to the SQL you already know) → Stage 2 → Stage 4 with a focus on retrieval, vector search and data pipelines for RAG. That is where database skills carry over most directly.

Software developer

Your advantage: code, APIs, testing and Git. Stage 4 will feel natural: calling an LLM API and building agent tools is ordinary software work.

Your gap: the habit of evaluating models. Developers often ship an LLM feature after it looks right on five examples.

Fastest route: a quick Stage 1 (maths and statistics) → Stage 2 for evaluation discipline → Stage 4 in depth → Stage 5. Treat prompts and test sets like code: version them and run them in CI.

Fresh student

Your advantage: time, and no bad habits to unlearn.

Your gap: context. You haven’t yet seen real company data, or why a slightly less accurate but explainable model can win.

Fastest route: follow the stages in order with no skipping, and finish every project. Your portfolio is the stage projects, not a list of courses. Start with the Foundations guide and the Linux course in parallel.

Where AWS certifications fit

If you work on AWS, the current AI/ML certifications map onto the stages (per AWS’s official certification pages, checked October 2026):

The older AWS Certified Machine Learning – Specialty has been retired. Its last exam date was 31 March 2026, so don’t plan around it.

OPINION: take the exam that matches the stage you have finished, as a check on it, not as a substitute for the project.

Common mistakes I see

  1. Starting with deep learning or LLMs. Without Stage 2 you can’t tell a good model from a lucky one.
  2. Testing on training data. The classic reason a model looks great in a notebook and fails in production.
  3. Using accuracy on imbalanced data. A fraud model that always predicts “not fraud” can score high and catch nothing.
  4. Tutorial loops. Watching a fourth course instead of building the first project.
  5. Ignoring data and SQL. The model is usually the smallest part of a production ML system.
  6. Shipping GenAI without an evaluation set. If you can’t measure it, you can’t safely change the prompt, the chunk size or the model.
  7. Forgetting cost and teardown in the cloud. Idle endpoints and forgotten notebooks keep billing.

FAQ

What should I learn first in AI/ML?

Learn Python, basic SQL and data handling with pandas first, along with practical maths: vectors, gradients and basic statistics. Then move to classic machine learning with scikit-learn. These foundations support every later stage, including LLMs and RAG.

Do I need advanced maths for machine learning?

No, not to start. You need vectors and dot products, an intuitive idea of gradients, and basic probability and statistics. Deeper maths, such as the matrix form of regression, is worth learning when a real problem calls for it.

Should I learn machine learning or generative AI first?

Learn classic machine learning first, at least until you can validate a model honestly. GenAI work depends on the same evaluation and data skills; skipping ML usually leads to GenAI systems nobody can measure.

How long does it take to learn AI/ML?

It depends on your background and weekly time. Measure progress by finished projects, not hours. Move on when you can complete a stage’s project without following a tutorial step by step.

Is coding required to work in AI/ML?

For hands-on engineering roles, yes. Python is the standard, and SQL is close behind. Some roles focus on strategy, product or governance and need less code, but they still need a solid grasp of how models are built and evaluated.

Which AWS certification is best for AI/ML?

Start with AWS Certified AI Practitioner for the concepts. Take AWS Certified Machine Learning Engineer – Associate once you can build and deploy models, and AWS Certified Generative AI Developer – Professional once you build GenAI applications. Always check the official AWS page for the current exam version.

Where to go next

Pick your route above, find your current stage in the table, and do that stage’s project first. The Tutorials page lists everything referenced here: the Foundations guide, the linear regression tutorials, the RAG tutorial, the AI agents lesson and the ML Solutions Quest.

To learn the cloud side with live labs, see AWS training and the AWS Solutions Architect course. For the GenAI side, the RAG & LLM Applications training covers chunking, embeddings, vector stores and evaluation, the core of Stage 4.

Newsletter

Enjoyed this post?

Get new posts, AI/ML tutorials, AWS batch dates and Oracle tips by email. No spam, unsubscribe anytime.

I’m interested in

Double opt-in: you’ll get a confirmation email first.