Overall strategy
Begin after Part 1's engineering and interview readiness check. Build on that foundation with ML, PyTorch, AI application development and production ML skills.
| Stage | Weeks | Main goal |
|---|---|---|
| 0 | Week 0 | Environment, API budget and data boundary, market scan, two learn-tests — ML libraries, hosting and GPU are set up in the week that first needs them |
| 1 | Weeks 1–5 | Python, data, SQL, statistics, math foundations — Week 5 is the extension week that turns the Week 0 learn-tests into passes |
| 1.5 | Week 6 | Early win — ship a small AI demo, tag it v0 |
| 2 | Weeks 7–18 | Classical ML, Scikit-Learn, serving, Project #1 — Week 12 consolidates before serving |
| 3 | Weeks 19–28 | PyTorch, deep learning, transformers, Project #2 — Week 22 consolidates before training deep networks |
| 4 | Weeks 29–33 | LLMs + Applied AI — and Project #3's foundation, built as you learn |
| 5 | Weeks 34–40 | Production ML/MLOps + hardening Project #3 into the flagship |
| DSA | Weeks 1–40, every other week | 1 hr fortnightly re-solving from your Part 1 log, so the habit does not lapse |
| Design | Weeks 14–36, twelve sessions | ML system design only — serving, training pipelines, ranking, batch inference, monitoring — then Project #3 as a design answer |
| Search | When you choose to apply | Refresh Part 1 interview skills and add role-specific ML practice; budget separately from full-time ML study |
10 hrs/weekKeep current jobBuild while learning3 serious projects3 buffer weeks + 3 extension weeks~410 hrs in Part 2 — ~380 ML learning + building, ~20 DSA, ~12 ML system design
The honest arithmetic. Ten hours a week is a hard limit, not a target: every week's allocation below sums to ten, and anything added to a week (a retake, a dataset choice, a deployment) replaces something inside those ten rather than sitting on top. Part 1 may let you pass some foundation checks; use the actual results to remove extension weeks rather than assuming they are needed. Pass both Week 0 learn-tests and Week 5 disappears. Treat 41 as the target, not a promise — if a stage runs long, the duration moves, the content does not get compressed to fit.
Weeks 30–35 do double duty: they are instruction weeks, but every exercise is committed to the Project #3 repository rather than to scratch files. That is what lets the hardening weeks harden rather than build — see Stage 5.
Stage 0 — Set yourself up
Duration: Week 0
The goal is to make sure you can spend the next several months building without environment problems slowing you down — and to gather the two pieces of information you will need much later, while gathering them is still cheap.
This week is full: ten hours covers the market scan, the installs, the cost and data decisions, and the two learn-tests. If installs go badly, the scan and the learn-tests are the parts to protect. Everything else here can be finished in Week 1.
Scan your actual market first
Ninety minutes, before you install anything. Pull 20 real job postings from the market you can actually work in — your city, your country, or remote roles that hire from it — across the titles listed in the ML interview section. For each, note the title, the modelling depth asked for, whether cloud, Kubernetes, or an orchestrator appears, whether they want RAG and evaluation or classical ML, and the education line — including whether it says "or equivalent experience." That last phrase is common, and counting how often it appears in your market tells you far more than any general claim about whether degrees matter.
You are not choosing your specialism today. The Applied AI and MLE tracks in this plan are nearly identical until Week 25, so committing now would mean deciding with the least information you will ever have. You are building the evidence for the decision in Week 18, once you have shipped Project #1 and know what you actually enjoy. Save the file; you will re-read it then.
Install and configure
Foundation checks after Part 1
Recheck Python and SQL before choosing extension weeks. You have already practised them in Part 1, but not window functions or the idioms data code leans on; pass or repeat based on the artifact, not the calendar. Use the stated conditions below and record results. Missing setup is untested and gets resolved before interpreting the result as a skill gap.
| Week | The artifact you must be able to produce, in one sitting | If you fail in Week 0 |
|---|---|---|
| 1 — Python | Write a class with a context manager, a generator, and type hints, and explain a decorator you did not write | Week 1 in full, plus 4 hours in Week 5. Retake at the end of Week 5. |
| 2 — the SQL half | Write a three-table join with a group-by and a window function from memory, with no reference open | Week 2 in full, plus 4 hours in Week 5. Retake at the end of Week 5. |
If you pass one, that is how an extension week disappears. Pass Python and SQL and Week 5 is skipped. There is no FastAPI, Docker or CI test: Part 1 built all three (E1, E12), and Weeks 14 and 34 extend them to models. Nothing gets banked to Week 25, 30 or Project #3 — the extension weeks are the honest replacement for a cushion.
Your compute and cost plan — some now, the rest on a date
Week 6 makes your first paid LLM call, Week 17 puts Project #1 at a public URL, and Week 24 fine-tunes a pretrained vision model. Each fails late and expensively if you find out the week you need it. The API budget, storage and secrets are settled now; hosting and GPU are checked on the dates below, each with weeks of slack before the week that depends on it.
.env file, never in code or notebooks. Add .env to .gitignore before your first commit, not after.The book is already mapped
Pre-write your cut-list
Three buffer weeks across 41 is roughly 7% slack — thin for nine months alongside a full-time job. The buffers will absorb a bad fortnight; they will not absorb a bad quarter. So decide now, while you are rested and nothing has gone wrong, what gets sacrificed when something does. The extension weeks (5, 12, 22) are not buffers — they are instruction time you have already been told you need. Do not spend them on catch-up and then arrive at Week 18 with both gone.
The rule: if you reach a buffer week already behind, cut from this list rather than extend the schedule. Breadth goes before project quality — the earlier cuts sit in the weeks where a junior candidate is least likely to be interviewed on the detail. In order:
- The ML practice hours — spend them on catch-up for the current stage before cutting any topic below.
- Week 13 clustering — read for awareness, build nothing.
- Week 10 tree internals — Gini versus entropy is not worth a week of slippage; you will still train and tune trees.
- Week 11 beyond Random Forest and Gradient Boosting — AdaBoost, Extra Trees, stacking to awareness only.
- Week 3 confidence intervals and sampling theory — know the vocabulary, move on. It sits late in this list because a weak statistics base costs you in Stage 2.
- Week 24 computer vision — only once you have committed to the text path for Project #2.
- ML design sessions outside Weeks 31–36 become concept-only — read the concept, skip the practice design. The Project #3 presentations are never cut.
v0 tag; and anything else in Stages 4 or 5. Those weeks are the portfolio. If you are behind and everything above has been cut, the schedule extends — you do not quietly reclaim these. The fortnightly DSA hour also stays: it is cheap, and a lapsed habit is harder to restart than to keep.Suggested learning repository
ml-transition/ ├── python/ ├── numpy/ ├── pandas/ ├── sql/ ├── math/ ├── sklearn/ ├── pytorch/ ├── notebooks/ └── projects/
Four files exist: the market scan of 20 postings, the cut-list, the written data boundary, and the learn-test results. None takes long. All are worth more in Week 18 than they are today.
Materials and resources
This section is the finished version of the Week 0 mapping task. The edition is confirmed, the chapters are assigned, and every week the book does not cover has a named primary source.
The book — edition confirmed
The copy in this folder is the PyTorch edition: 19 chapters plus two appendices (Autodiff; Mixed Precision and Quantization). Chapter 10 is Building Neural Networks with PyTorch and Chapter 15 is Transformers for Natural Language Processing and Chatbots. Stage 3 therefore works exactly as written and you need no substitute PyTorch book — only the official tutorials as a reference.
Chapter map
| Week | Chapter | Note |
|---|---|---|
| 1–4 | — | Not in the book. It assumes NumPy and Pandas already — see the next table. |
| 7 | 1–2 | Landscape, then the end-to-end project |
| 8 | 3 | Classification |
| 9 | 4 | Training Models |
| 10 | 5 | Decision Trees |
| 11 | 6 | Ensemble Learning and Random Forests |
| 13 | 7 + 8 | PCA from 7; K-Means only from 8. Skip DBSCAN, Gaussian mixtures, anomaly detection. |
| 19 | 9 | Introduction to Artificial Neural Networks |
| 20–21 | 10 | Read in 20, build from it in 21 |
| 23 | 11 | Training Deep Neural Networks |
| 24 | 12 | Deep Computer Vision with CNNs — trade away first if you chose Applied AI |
| 25 | 15 | Transformers. Skip 13–14 — RNN sequence models, already cut from the plan. |
Chapters 16–19 (vision/multimodal transformers, speeding up transformers, autoencoders and diffusion, reinforcement learning) are the optional tail the plan already flags. Appendix A is worth thirty minutes in Week 20 if autograd does not click.
Weeks the book does not cover
One primary source each. All are first-party or long-stable, and none of them costs anything to read.
| Week | Topic | Primary source |
|---|---|---|
| 1, 5 | Python | docs.python.org — the official Tutorial, chapters 3–9 only, then the Week 0 learn-test artifact rebuilt without notes |
| 2 | NumPy, Pandas, SQL | numpy.org "absolute beginner's guide"; pandas.pydata.org "10 minutes to pandas"; sqlbolt.com for joins and aggregates, then the PostgreSQL tutorial on window functions and CTEs — SQLBolt does not teach either, and the Week 0 SQL learn-test asks for a window function |
| 3 | Statistics and probability | khanacademy.org — the descriptive statistics, probability, random variables, and normal distribution units only; skip inference beyond confidence intervals |
| 4 | Linear algebra, calculus | 3blue1brown.com — Essence of Linear Algebra and Essence of Calculus. Intuition, not proofs. |
| 6, 29–31 | LLM APIs, embeddings, RAG | Your chosen provider's documentation — the only source that stays current on models, pricing, and limits |
| 14 | Serving a model | Your Part 1 shop app — reuse its FastAPI and Docker setup; fastapi.tiangolo.com as reference |
| 15 | Experiment tracking | mlflow.org quickstart |
| 32 | Prompt injection, untrusted content | genai.owasp.org — Top 10 for LLM Applications. This is the security half of the week. |
| 32 | Tool calling | Provider documentation |
| 33 | Evaluation | No canonical text. Your own suite from Week 30 is the material. |
| 34 | CI/CD | docs.github.com — GitHub Actions |
| 35 | Monitoring and drift | Evidently or equivalent — read the concept docs, not the API reference |
| 14–36 | ML system design | developers.google.com — Rules of Machine Learning, the section that matches each session; the practice designs are yours |
Data structures and algorithms — a fortnightly habit
One hour every other week, Weeks 1–40, including project and extension weeks — about 20 sessions. The coding skills were built in Part 1 Stage I; this hour only keeps them from lapsing. It is not a curriculum, and it does not replace fresh mocks before a search. The other week, the Tuesday hour is an ML practice hour — see Weekly routine.
Each session: re-solve one earlier problem from your Part 1 dsa/solved.md cold, then attempt one unseen problem from the pattern your log shows is weakest. Same rules as Part 1: Python, 35–45 minutes on the new problem, state the complexity, test edge cases, and log assistance honestly in the same file. Rotate through the Part 1 patterns rather than restarting at arrays; add tries or 2-D DP only if your target roles ask for them.
Coding, debugging, system design and behavioral preparation were completed before this part. If fresh mocks show gaps when a search approaches, refresh them using Part 1's application budget and reduce or pause ML hours.
ML system design
The backend designs — URL shortener, rate limiter, notification service — were practised in Part 1 and are not repeated here. This track covers only what is new: systems built around a model. Twelve Wednesday sessions between Week 14 and Week 36, each in the week whose material it draws on. Every other Wednesday from Week 7 is an ML practice hour.
Each session: one concept, then one small practice design done by you alone — 35 minutes on paper or a whiteboard, phone photo, requirements → estimates → API → data model → high-level components → one deep-dive → bottlenecks. Only then open an AI and paste the photo with this instruction: "You are the interviewer. Ask me the follow-up questions a senior engineer would ask about this design, one at a time. Do not propose a design. At the end, list what I missed." Note the questions you could not answer in system-design/wNN.md and commit it. If the model ever starts designing for you, the session is void — the skill being trained is defending your own diagram under pressure.
| Week | Concept | Practice design (you, alone, then attacked) |
|---|---|---|
| 14 | Serving a model behind an API: request/response schema, model versioning, timeouts, what the endpoint returns when the model fails | Your own Project #1 endpoint as a design answer — you are building it this week |
| 15 | Training pipelines: data snapshots, reproducibility, experiment tracking, model registry, promotion rules | Scheduled retraining for Project #1 — what triggers it, how a new model is compared with the live one, and how you roll back |
| 20 | Recommendation and ranking: candidate generation → ranking → re-ranking; offline vs online features | A product recommender, v1 |
| 21 | Training–serving skew, feature stores, feedback loops, cold start | Recommender v2 — the same design, attacked on freshness, skew and new users |
| 24 | Serving ML online: batch vs real-time inference, model versioning and rollout, feature consistency | An online prediction service — your deployed Project #1, scaled to 1,000 QPS. Keep this session even if you skim the rest of Week 24. |
| 28 | Batch inference: scheduled scoring, object storage for model artifacts and outputs, backfills | Nightly scoring of 10 million rows — and what happens when the job dies halfway |
| 29 | Server-Sent Events, streaming responses, backpressure | A streaming LLM endpoint — what the client sees when the model is slow, the provider times out, or the connection drops mid-answer |
| 30 | How retrieval systems scale: vector indexes (flat vs HNSW), re-indexing, freshness | Concept only this week — you are building the real thing |
| 31 | — | Project #3 as a design answer, session 1 — the RAG pipeline you just assembled: ingestion, chunking, embedding, vector store, retrieval, reranking, LLM. Expect attacks on freshness, latency, cost per query, multi-tenancy, injection. |
| 33 | — | Project #3 as a design answer, session 2 — now with the evaluation suite and held-out numbers. Expect "how do you know it is good?" and "what breaks at 10× documents?" |
| 35 | Monitoring ML in production: data drift, prediction drift, delayed labels, alert thresholds | Monitoring for Project #1 — which signals you log, what fires an alert, and how true labels come back |
| 36 | — | Project #3 as a design answer, session 3, recorded — after hardening, with CI and monitoring. This recording goes into the Week 39–40 walkthrough package. |
The RAG pipeline is your best design material. Nobody else will have built, measured, and hardened the system they are describing; a URL shortener is table stakes, Project #3 is a differentiator. That is why three of the sessions are spent presenting it, and why the third one is recorded — it doubles as the architecture segment of your final walkthrough. Refresh the Part 1 backend designs with a mock when a search approaches.
The Week 29 session uses developer.mozilla.org on Server-Sent Events as its source, because LLM token streaming is SSE and it is the one networking topic that is genuinely load-bearing here.
Stage 1 — Foundations
Duration: Weeks 1–4, plus extension Week 5
Focus on the data and maths foundations needed for ML. Use Week 0 results to choose the Python and SQL review depth. The Tuesday hour alternates between DSA and ML practice.
Week 1 — Python for data work
Review lists, dicts, sets, tuples, comprehensions, functions, lambdas, classes, exceptions, context managers, iterators, generators, decorators, modules, packages and type hints; focus on gaps shown by the Week 0 check. Part 1 gave you Python for backend code; this week is about the idioms (comprehensions, unpacking, enumerate/zip, slicing) that ML code and NeetCode solutions both lean on. This is a learn-test week: the artifact is the Week 0 Python test, rebuilt at the end of the week without notes.
Your target is not clever Python. Your target is being able to read ML code without the language itself slowing you down.
Week 2 — NumPy + Pandas + SQL
Set up notebooks first: pip install ipykernel into the project environment and VS Code runs them directly — no separate Jupyter server. This is the first week that needs them.
Practice on a CSV dataset. Explore missing values, distributions, correlations, outliers, and target balance before training any model. Then load the same dataset into a database and reproduce three of those analyses in SQL.
Week 3 — Statistics
Learn mean, median, variance, standard deviation, percentiles, correlation, probability, conditional probability, independence, random variables, normal distributions, expected value, sampling, population vs sample, bias, variance, confidence intervals, and noise.
Week 4 — Linear algebra + calculus
Focus on scalar, vector, matrix, tensor, vector addition, dot product, matrix multiplication, transpose, dimensions, functions, slope, derivative, partial derivative, gradient, and chain-rule intuition.
prediction ↓ calculate loss ↓ calculate gradient ↓ adjust parameters ↓ repeat
Stage 1 mini-project
load data ↓ clean data ↓ basic statistics ↓ plots ↓ written conclusions
External proof: the mini-project is pushed to a public GitHub repo with a README a stranger could follow, and
dsa/solved.md shows every DSA session scheduled so far logged.Stage 1.5 — Early win: ship a small AI demo
Duration: Week 6
This week is out of sequence on purpose, and it is the most important structural change in the plan.
Your leverage as a working software engineer is application engineering, not model research. That skill needs no ML theory at all — you can call an LLM API and wire up retrieval with the Python you already have. Building it now, before you have "earned" it, gives you three things you would otherwise not have until Month 8:
- Something concrete to show and talk about, months before the flagship project exists.
- A real reason to care about the abstract material in Stages 2 and 3 — you will have already felt the problems that evaluation and retrieval quality solve.
- Proof to yourself that the transition is achievable, at the point in the roadmap where motivation is most fragile.
Scope it aggressively
Nine hours — the whole week, not one weekend. Your routine gives you five hours across Saturday and Sunday; the other four come from the weekday slots, spent reading provider documentation so the weekend is pure building. Load a folder of documents, chunk them naively, embed them, store them in the simplest vector store you can find, and answer questions over them from a command line. No auth, no database, no UI, no tests, no Docker.
git status is clean of it before the first push.your documents ↓ naive chunking ↓ embeddings ↓ simple vector store ↓ question → retrieve → LLM → answer
It will be mediocre. That is the point — Weeks 30 and 33 exist to teach you exactly why, and in Week 30 you will fork this repository into Project #3 and rebuild it properly.
Then tag it
v0 and stop touching it. This repository is now a permanent exhibit, not a working branch — every improvement from here happens in the Project #3 fork. Shown side by side in Week 40, the naive first attempt and the measured rebuild tell a story about your growth that neither tells alone. You cannot show that contrast if you have quietly overwritten the first half of it.Stage 2 — Classical Machine Learning
Duration: Weeks 7–18
This is the core Scikit-Learn stage and the first major section of Géron’s book.
Week 7 — ML landscape + end-to-end workflow
Install Scikit-Learn first. Study supervised vs unsupervised learning, batch vs online learning, overfitting, underfitting, generalization, hyperparameters, validation, and test sets.
Problem ↓ Data ↓ Quick structural look — shape, dtypes, target balance ↓ Choose split strategy → reserve the test set ↓ Detailed EDA — training data only ↓ Preprocessing ↓ Baseline ↓ Train models ↓ Cross-validation ↓ Tune ↓ Evaluate on the test set — once ↓ Deploy ↓ Monitor
Time-based when the rows carry timestamps and the prediction is about the future — a transaction log, monthly demand. Train on earlier periods, test on later ones, or you are predicting the past.
Group-based when several rows share an entity — multiple rows per customer, patient, or device. Keep every row of a group on one side, or the model recognises the entity instead of learning the pattern.
Stratified random when the dataset is a single snapshot with one row per entity and no time axis — which is what many public churn datasets are. Stratify on the target so the class balance survives the split.
A churn dataset can need any of the three, depending on how it was collected. Write down which one you used and why; it is the first question a reviewer asks.
Week 8 — Classification
Learn binary and multiclass classification, confusion matrices, accuracy, precision, recall, F1, ROC, AUC, PR-AUC, calibration, and decision thresholds.
Change the decision threshold on a small classifier and observe what happens to precision, recall, and F1.
Then check calibration. If you tell a business that a customer has a 30% chance of churning, that number has to mean 30%, and a plot of predicted probability against observed frequency is how you find out it does not. Both are cheap, both come up in interviews, and both separate someone who has deployed a classifier from someone who has only trained one.
Week 9 — How models learn
Study linear regression, gradient descent, SGD, polynomial regression, regularization, Ridge, Lasso, Logistic Regression, and learning curves.
Implement a tiny linear regression model once without Scikit-Learn to understand prediction, loss, and optimization. This is the week that pays for the Week 5 chain-rule hours. Then fit a polynomial model to the same data and plot its learning curves against the linear one — that pair of plots is the clearest picture of over- and underfitting you will see all year.
Week 10 — Decision Trees
Study splits, Gini impurity, entropy, depth, overfitting, and regularization. Train trees with different maximum depths and visualize the behavior.
Week 11 — Ensembles
Learn bagging, Random Forests, Gradient Boosting, Histogram Gradient Boosting, and feature importance. Run Extra Trees, AdaBoost, and a simple stacking ensemble once each on the same data so you have seen them behave — they are the first thing on the cut-list after clustering if you slip.
Compare Logistic Regression, Decision Tree, Random Forest, and Gradient Boosting on the same problem. Extra comparisons beyond those four are stretch work, here and in Project #1.
Week 13 — Dimensionality reduction + clustering
Learn PCA, the curse of dimensionality, and K-Means. These two carry almost all of the interview and practical weight in this area. Clustering is one small exercise — cluster something, look at the clusters, write three sentences about whether they mean anything — not a week of it.
Week 14 — Serving: FastAPI + Docker
Moved forward from Stage 5, because Project #1 requires it. Wrap a trained Scikit-Learn model in a FastAPI endpoint with a request/response schema, containerize it, and run the container locally. Keep it minimal — no auth, no database. FastAPI and Docker are familiar from Part 1, so spend the time on what a model adds: load the artifact once at startup, validate the features in the request schema, and return the model version with every prediction. The Wednesday design session presents this same endpoint as a design answer — you will be asked about it in exactly that form. Pick the Project #1 hosting platform this week — see the hosting card in Stage 0 — half an hour out of the week's building time.
Week 15 — Experiment tracking
Track model versions, dataset versions, hyperparameters, metrics, artifacts, and code versions. Use a tool such as MLflow or an equivalent. You want this habit in place before Project #1, not after, so the project produces a real experiment log instead of a folder of forgotten notebook runs. Local SQLite as the tracking backend and local files for artifacts — MLflow's model registry needs a database-backed store, and a file-only setup does not support it — then register one model version and load it back by name. No remote service. The habit is the point; the infrastructure is not.
Weeks 16–17 — Project #1
Choose a structured-data problem such as customer churn, loan default, employee attrition, fraud detection, house prices, insurance claims, or customer conversion.
Mandatory, in this order: data loaded and split, baseline recorded, one trained model compared against it with the result explained either way, honest evaluation on the held-out set exactly once, the FastAPI endpoint from Week 14, a Dockerfile, a README. Stretch, and only once all of that runs: comparison across four algorithms, hyperparameter tuning, feature-importance analysis, calibration plots. A finished small project beats an unfinished thorough one, and the Week 18 buffer is not a licence to start stretch goals in Week 17.
project/ ├── data/ ├── notebooks/ │ └── exploration.ipynb ├── src/ │ ├── preprocessing.py │ ├── train.py │ ├── evaluate.py │ └── predict.py ├── tests/ ├── models/ ├── api/ │ └── main.py ├── Dockerfile ├── requirements.txt └── README.md
Your README should explain the problem, dataset, metric choice, experiments, final result, failure modes, and how to run/deploy the project.
A link a reviewer can open outweighs any amount of "and it's containerized" on a CV — and it turns Docker from something you have read about into something you have operated.
Applied AI — take the text path for Project #2, treat Week 24's computer vision as optional from here on, and spend the recovered hours on evaluation, retrieval quality, and deployment in Stages 4 and 5.
MLE — keep the full modelling breadth and Week 24, and use the platform decision in the ML interview section to budget any needed data-pipeline and cloud practice.
Pick one. The schedule is already full, and hedging across both is how people arrive at Week 40 holding two half-portfolios. Week 18 is also the honest moment for this: Weeks 1–25 are nearly identical either way, so nothing you have done is wasted whichever you choose — and you now know something Week 0 could not tell you, which is whether you actually enjoy this.
External proof: Project #1 is public, its API runs from a single
docker run, it is deployed at a URL that currently responds, and one other engineer has read the README and told you what was unclear.Stage 3 — Deep Learning + PyTorch
Duration: Weeks 19–28, including extension Week 22
Week 19 — Neural network fundamentals
Learn neurons, layers, inputs, weights, biases, activation functions, forward pass, loss, backpropagation, epochs, batch size, and learning rate.
Week 20 — PyTorch fundamentals
Learn tensors, nn.Module, autograd, Dataset/DataLoader, model saving/loading, and basic optimization.
for X, y in loader:
optimizer.zero_grad()
prediction = model(X)
loss = criterion(prediction, y)
loss.backward()
optimizer.step()
Week 21 — Build small neural networks yourself
Implement a regression neural network, a binary classifier, and a multiclass classifier. Track training loss, validation loss, and useful metrics.
loss.backward() does to a colleague.Week 23 — Training deep networks
Study vanishing gradients, initialization, ReLU, normalization, gradient clipping, Adam/AdamW, learning-rate scheduling, dropout, weight decay, and transfer learning. Learn two learning-rate schedules properly (OneCycle and cosine) and compare them on one of your Week 21 networks; then do one small transfer-learning exercise — freeze a pretrained backbone, train a new head — so that Week 24 or Week 26 is the second time you have done it, not the first.
Week 24 — Computer vision
Learn CNNs, convolution, kernels/filters, pooling, feature maps, ResNet, and transfer learning. Fine-tune a pretrained vision model on a small dataset — this is the week your Week 18 GPU check pays off.
If you chose Applied AI in Week 18, this week is optional. Skim it for vocabulary, take the text path in Weeks 26–27 where you will fine-tune a transformer instead, and move the recovered hours to Week 25 or Week 30. Keep the Wednesday design session either way — an online prediction service is your material regardless of path. If you chose MLE, keep the week in full.
Week 25 — Attention + Transformers
Go deep on (about 6 hours): tokenization, embeddings, attention, and Transformer architecture. Attention is the one idea in this week that interviewers actually probe, and it is the foundation of everything in Stage 4.
Working familiarity only (about 2 hours): load and run a pretrained model through Hugging Face, and know the BERT vs GPT-style distinction. In-context learning and instruction tuning are read, not practised — you will live them from Week 29 onward. You will use all of these constantly in Stage 4 without needing to have trained one.
The remaining hour is the Wednesday ML practice hour, spent this week to choose the Project #2 dataset and target metric, check the licence, and write the baseline down. That is nine hours plus the Tuesday hour — the week is exactly full. If something has to give, it is the two familiarity hours, not the attention hours.
Weeks 26–27 — Project #2
Two weeks, not one — and never collapsed to one. Build either an image classifier or text classifier with PyTorch. Include a Dataset/DataLoader, training loop, evaluation, experiment tracking, inference API, and an error-analysis section.
The dataset and target metric were chosen in Week 25. Separate mandatory from stretch: a working training loop, honest evaluation, and the error analysis are mandatory; architecture comparisons and hyperparameter sweeps are not. The same reproducibility rules apply — one command, pinned versions, fixed seeds, data provenance, stated limitations.
If you choose the text path, fine-tune a small pretrained transformer rather than training from scratch — this is the fine-tuning exercise deferred from Week 25, and it belongs here where the hours exist.
The error analysis is the part that distinguishes you from a tutorial follower. Look at the examples your model gets wrong, categorize the failures, and write down what you would do next. It is the deliverable that survives every cut.
Stage 4 — Applied AI / LLM Engineering
Duration: Weeks 29–33
You built a naive version of this in Week 6. Now you learn why it was mediocre. The routine does not change here — the same ten hours, the same slots.
v0 demo in Week 30 and treat these four weeks as the flagship's foundation: Week 30 becomes its ingestion and retrieval layer, Week 31 its RAG pipeline, Week 32 its one tool, Week 33 its evaluation suite.This is what makes the Stage 5 arithmetic work. Weeks 36–39 give you about 35 hours — enough to harden a system that already exists and has been measured for six weeks, not enough to build a production RAG system from scratch. Build it twice and you will finish neither.
Week 29 — LLM APIs
Learn system/user messages, context windows, tokens, temperature, structured outputs, streaming, rate limits, retries, timeouts, caching, and cost modelling. The Wednesday design session is Server-Sent Events and a streaming endpoint — LLM token streaming is SSE, and it is the one networking topic that is load-bearing here.
Caching earns a mention of its own because it decides whether the evaluation suite you are about to build is something you can afford to run on every commit or something you run twice and quietly abandon. Work out the per-run cost of that suite now, and know which lever brings it down — prompt caching, response caching, or a smaller model for iteration.
Spend about 2½ of this week's building hours writing the next 30 Project #3 questions — see Week 30.
Week 30 — Embeddings + vector search
Learn embeddings, vector similarity, cosine similarity, semantic search, chunking strategies, vector storage, and reranking.
Reranking is the cheapest large win available in retrieval, so it belongs here rather than nowhere: retrieve generously with vector search, then reorder the candidates with a stronger cross-encoder or model-based scorer before anything reaches the LLM. Measure the difference — it doubles as the clearest proof that your evaluation set actually works. If reranking does not beat the baseline on your documents, that is the result you report; a measured "no improvement" is worth more than an assumed one.
Project #3 starts this week. Fork your tagged v0 demo into a new repository and leave v0 untouched. Everything below lands in the fork.
Choose the vector store in writing. The Week 6 spike used whatever was simplest; the flagship gets a decision. Default to pgvector in Postgres — one database to run and back up, with metadata filters in plain SQL. Move to a dedicated vector database only for a need you can name: corpus size, a latency you have measured, or filtering and hybrid search pgvector cannot do. Put the choice and its reason in the README; the Week 31 design session will ask.
v0 documents — about 8 hours at five minutes each, so it is spread over three weeks: about 40 in the Week 28 buffer, 30 in Week 29, and the last 30 here. If the buffer went to catch-up, Week 30 extends by a week; the set does not shrink. For each, record the expected answer, the supporting passage — quoted text plus page or section, never a chunk ID, because Week 30 changes the chunking and chunk IDs will not survive it — and whether the system should abstain. Include ten to fifteen abstain cases: questions the documents do not answer, where the right output is "not in these documents". Then split them, and commit the split:Development set (~60 questions) — what you iterate against all through Weeks 30–32. Run it constantly. Tune chunking, embeddings, and reranking against it freely.
Held-out set (~40 questions) — commit it, then do not look at it, do not run against it, and do not tune anything on it until Week 33.
For retrieval, count Recall@k (how often the correct passage lands in the top k) and MRR (how high it ranks) — over the answerable questions only. Abstain cases have no gold passage, so they cannot be scored for retrieval; they are scored under abstention in Week 33 and must be excluded here or they silently drag both numbers down. Record both on the development set before changing anything. That pair is your baseline, and every decision from here is measured rather than guessed.
The size is not arbitrary either. A 20-question held-out set moves in 5-point jumps; 40 questions give finer score increments and more evidence, though finer increments alone do not make a difference significant. Writing a hundred questions is about eight dull hours, which is why it starts two weeks early. Even so, treat the result as a small portfolio evaluation — report the counts and the limitations alongside the number, not as proof of broad reliability.
Week 31 — RAG
Assemble the full pipeline in the Project #3 repository, running the development set as you go. Answer quality now rests on retrieval quality, and you already have numbers for the second.
Then present it. This week's system-design session is Project #3 as a design answer, session 1 — draw the pipeline you just built, alone, then let the AI interviewer attack it (freshness, latency, cost per query, what happens at ten times the documents, what happens when a chunk contains instructions). Write down the questions you could not answer; several of them are Week 32 and Week 33's to-do list.
Documents ↓ Parser ↓ Chunks ↓ Embeddings ↓ Vector DB User question ↓ Embed question ↓ Retrieve chunks ↓ LLM ↓ Answer + sources
Week 32 — Tool calling, and the failure modes that matter
Learn tool/function calling, workflow state, retries, tool failures, timeouts, structured-output failures, and human approval boundaries.
Weight this week toward security and failure handling rather than agent frameworks. Spend roughly half of it on prompt injection, untrusted retrieved content, and data leakage. That is not a detour: your system reads documents and hands them to a model that can call tools, so a hostile instruction hidden in a retrieved chunk is precisely the shape of the threat. "What happens if one of your documents tells the assistant to ignore its instructions?" is a question an interviewer can ask about the system you actually built. Generic multi-agent orchestration is not.
Use filtered retrieval: the model calls it with an explicit source or date filter when a question is scoped ("what did the 2024 handbook say about X?"). It is genuinely useful, it needs no write access, and it fails in ways you can test — bad filter values, empty result sets, filters that silently match nothing and quietly return a confident answer from the wrong document.
Write those failure tests now, while the material is fresh. If you later want an approval boundary to talk about, add a second tool that writes something — but not at the cost of the evaluation work.
Week 33 — Evaluation
Expand the Week 30 development set into a real evaluation suite: cover the failure cases you have hit since, and add cases that exercise both the tool and the injection attempts from Week 32. Track Recall@k and MRR for retrieval, citation correctness (do the cited passages actually support the answer?), faithfulness and hallucination rate, abstention (does it say "not in these documents" when it should, and only then?), task success, tool success, latency, and cost per query.
Grade answers against a written rubric, committed next to the questions: what counts as correct, partially correct, and wrong for each, so that a model-graded score means the same thing next month. Run the grader at temperature zero and pin the model version; then run the whole suite three times on the same commit and record the spread. Three runs give you the observed run-to-run variation — an initial estimate, not a floor. An improvement smaller than that spread is not yet demonstrated, and it is the number that decides how loose Week 34's CI threshold has to start.
Then check the grader is right, not just consistent. Grade about 20 development-set answers yourself against the rubric without looking at its scores, compare, and record the agreement rate next to the suite. Where you disagree, fix the rubric or the grader prompt and check again before trusting its numbers. About an hour, from this week's building time.
Make it runnable as a single command that prints a table of scores. That command is what turns Project #3 from a demo into something you can defend — and it is what Week 34's CI will run.
System-design session 2 on Project #3 happens this week, after the held-out run: the same presentation as Week 31, but now every claim has a number behind it. "How do you know it is good?" is the question most people cannot answer about their own project. You can.
The held-out number is the one you quote. Keep tuning against it and it stops being held out — so the Week 38 measurement uses fresh questions rather than these.
v0 and the current system against the same held-out set, can report the comparison with numbers, and can explain the result — including which specific changes moved which metric, and any that did not. The rebuild will almost certainly score higher; the exit criterion is the measured, explained comparison, not the direction.External proof: the Project #3 repository already contains ingestion, retrieval with reranking, a RAG pipeline, one tool with failure tests, and an evaluation suite that runs from a single command. Stage 5 hardens this. It does not build it.
Stage 5 — Production ML + flagship project
Duration: Weeks 34–40
Production skills come before the hardening weeks, so Project #3 can actually apply them instead of promising them. The system itself already exists — you have been building it since Week 30.
Week 34 — CI/CD + ML testing
Test data schemas, preprocessing, model loading, prediction shapes, feature ranges, API behavior, and evaluation regressions. Reuse the pipeline you built in Part 1 E12, extend it with these checks, and retrofit it onto Projects #1 and #2 while the material is fresh.
For Project #3, CI runs your Week 33 evaluation suite and fails the build when a metric regresses past a threshold you set. That one check is the most persuasive thing in the repository: it says you treat retrieval quality as something that can break, rather than something you measured once and hoped about.
The LLM-judged metrics — faithfulness, citation correctness, abstention, task success — cost money per run and vary between runs: run them nightly or on tag, at temperature zero with the grader model pinned, and set the starting threshold from the run-to-run variation you measured in Week 33, then tighten it as more runs accumulate.
Both schedules run the development and regression cases only. The held-out sets are never wired into CI — a held-out set that runs on every push stops being held out within a week. A CI that flakes gets ignored within a fortnight, and an eval that costs a dollar a push gets switched off. The split is also the better answer to "you run LLM evals on every commit? What does that cost?"
Week 35 — Monitoring + drift
Learn service health, latency, error rates, data drift, prediction drift, model degradation, LLM evaluation regressions, and cost monitoring. Produce one drift report with Evidently or equivalent against Project #1's data. If the dataset has a time axis, its earlier and later periods are the reference and current windows. If it is a snapshot with no time axis — which Week 7 allows — build a clearly labelled synthetic shift: copy the test set, move one or two feature distributions in a way you write down, and run the report on that copy against the original. Label it synthetic everywhere it appears; the exercise is reading a drift report, not claiming to have seen drift; wire health, latency, and cost-per-query metrics into Project #3; and add an alert that fires when the nightly evaluation run drops below the Week 34 threshold. Do not survey the tooling landscape beyond that.
Weeks 36–39 — Project #3, hardening the flagship
Four weeks at the full rate — about 35 hours: eight in Week 36, where Wednesday is the recorded design session, and nine in each of Weeks 37–39. The fortnightly DSA hour continues throughout; in the other weeks the Tuesday hour goes to the flagship too.
Thirty-five hours does not build a production RAG system. It comfortably hardens one that has existed since Week 30, has been measured since Week 30, and already carries an evaluation suite, one tool, and CI. That is the entire reason Stage 4 committed its work to this repository — do not restart here. Your target:
Documents ↓ Parsing / chunking ↓ Embeddings ↓ Vector database ↓ Retrieval + reranking ←── measured against your eval set ↓ LLM ⇄ filtered-retrieval tool ↓ API (answer + sources) ↓ Logging + eval suite in CI
| Week | Focus |
|---|---|
| 36 | Harden ingestion — real parsing, awkward and malformed documents, re-index, confirm the development-set metrics still hold. Wednesday: Project #3 design session 3, recorded — it becomes the architecture segment of the Week 39–40 walkthrough |
| 37 | FastAPI endpoint, answer generation with sources, error handling, the Week 32 tool wired in with its failure paths, and response caching — measured as a before/after on cost per query and latency, because caching is the one cut item people do ask about. Two rules: the cache key includes everything an answer depends on (question, retrieval settings, prompt version, model version, corpus version), so a change to any of them invalidates it rather than serving a stale answer; and every evaluation run — nightly, Week 38's clean measurement, the noise-floor runs — bypasses the response cache, or it measures the cache instead of the system |
| 38 | Full test suite, both halves of the evaluation running in CI on their own schedules, and a fresh set of about 40 held-out questions for one clean final measurement |
| 39 | Docker, private deployment, logging and the Week 35 monitoring hooks, README and a written walkthrough |
Ship at the end of each week. A working narrow system in Week 36 that grows is far better than an ambitious system still broken in Week 39.
Deliberately kept in: the single tool from Week 32. Your portfolio claims tool calling, and a claim with no code behind it is worse than no claim at all. One tool, with failure tests — not an agent framework.
Add the rest later if the project has momentum, or when a specific job description makes one of them relevant. An honest README listing them as known gaps reads better than a half-finished auth system.
Put it behind platform-level auth or an IP allowlist — always, including a short-lived instance you bring up for a demo and take down afterwards. Pair it with the recorded demo and the written walkthrough so nothing depends on the service being awake when someone opens your CV. Keep the hard spend cap you set in Week 0.
v0 and Project #3 side by side in that walkthrough, with each one's numbers — the contrast between the naive demo and the measured system is the strongest single artifact you own. Use it alongside Part 1's engineering evidence when preparing for ML interviews.ML interviews — add specialization to Part 1
General software-engineering interview preparation belongs to Part 1 and does not wait for these projects. When you choose to apply for ML/AI roles, refresh that foundation and practise the role-specific topics below.
Roles this roadmap trains you for
Lead with the primary role you chose in Week 18. Refresh the actual postings and interview format before applying; use Part 1's engineering evidence alongside the ML projects.
Adjacent roles — reachable, but not fully covered here
Apply to these too; just know what you are missing so a gap does not surprise you in a screen.
ML interview topics
Bias vs variance, overfitting, regularization, data leakage, cross-validation, feature engineering, class imbalance, model selection, and evaluation metrics.
Algorithm knowledge
Linear Regression, Logistic Regression, Decision Trees, Random Forests, Gradient Boosting, K-Means, PCA, and Neural Networks.
Deep-learning knowledge
Gradient descent, backpropagation, activation functions, loss functions, batch size, learning rate, Adam, dropout, normalization, embeddings, attention, transformers, and fine-tuning.
AI-engineering knowledge
RAG, embeddings, vector search, chunking, reranking, tool calling, agents/workflows, structured outputs, evaluation, hallucinations, latency, cost, caching, and prompt injection.
System design
The Part 1 backend designs — URL shortener, rate limiter, notification service — refreshed with a mock before the search. The ML designs from this part: training pipelines and retraining, a recommender with its feature and skew problems, batch inference, and monitoring. Then the two you built: an online prediction service (Project #1, scaled) and Project #3's RAG pipeline as a design answer — ingestion, chunking, embeddings, vector index, retrieval and reranking, the tool boundary, evaluation in CI, cost, freshness, and prompt injection. Lead with Project #3 whenever the interviewer lets you choose.
Coding screens
Refresh the patterns practised in Part 1, including graphs and 1-D DP, using fresh timed mocks. Review runnable code, edge cases and complexity aloud. Use the employer's actual format to choose any additional practice.
Software fundamentals
Part 1 supplies the backend project, debugging, SQL, coding and design evidence. These ML projects add model serving, evaluation and monitoring. Explain what you personally built, measured and changed.
Be ready for "how do you use AI in your own work?" — now a routine question, and one where an experienced engineer stands apart from a bootcamp graduate. The strong answer is specific about the boundary: what you delegate, what you refuse to delegate, and how you review what comes back. You will have lived that boundary for nine months by then; say so concretely rather than in generalities.
Your three-project portfolio
Project #1 — Classical ML Weeks 16–17
Demonstrates Scikit-Learn, EDA, feature engineering, cross-validation, model comparison, metric choice under class imbalance, calibration, pipelines, API design, and Docker.
The only one of the three left running at a public URL — it has no paid LLM call per request, so it is the one to expose; hosting charges remain subject to the selected platform's limits and the Week 14 hosting budget.
Project #2 — Deep Learning Weeks 26–27
Demonstrates PyTorch, Dataset/DataLoader, training loops, transfer learning, evaluation, experiment tracking, error analysis, and inference deployment.
Project #3 — Applied AI Weeks 30–39
Your flagship, and the one that gets the most calendar rather than the least. Demonstrates LLMs, RAG, embeddings, vector databases, reranking, one tool with failure tests, FastAPI, Docker, a runnable evaluation suite with a held-out set, monitoring, tests, and CI/CD.
Built across ten weeks rather than four: Weeks 30–33 produce it while you learn, Weeks 34–35 supply the production skills, and Weeks 36–39 harden it in about 35 hours. Auth, a UI, and a separate application database are cut; caching is in, measured. The measured evaluation suite is what makes the project persuasive — protect those hours ahead of everything else.
v0 spike from Week 6. Keep it public and keep it frozen. Shown beside Project #3 with both sets of numbers, the naive first attempt and the measured rebuild tell a better story about your growth than either does alone — which only works if you never went back and quietly improved the first one.Weekly routine
One routine, all 41 weeks
| Day | Hours | Work |
|---|---|---|
| Monday | 1 | Reading — the chapter or the week's theory |
| Tuesday | 1 | Alternating: DSA one week (a re-solve and one unseen problem, no AI), ML practice the next |
| Wednesday | 1 | Maths (Weeks 1–6) → from Week 7, ML system design in the twelve scheduled weeks, ML practice in the rest |
| Thursday | 1 | Chapter exercises / coding along |
| Friday | 1 | Chapter exercises / coding along |
| Saturday | 3 | Project work |
| Sunday | 2 | Project / review |
The ML practice hour is where the time saved from Part 1 goes. Redo the week's exercise with the book closed, run it on a second dataset, or do error analysis on the week's model — whichever the week's material most needs. In foundation weeks it goes to that week's topic; in project weeks, to the project.
10 hours, exactly: 3 topic + 1 Tuesday (DSA or ML practice) + 1 Wednesday (maths, ML design or ML practice) + 5 building. The same table runs from Week 0 to Week 40 — there is no reduced-hours phase because there is no parallel job search. The week types differ only in where the hours point:
| Week type | Allocation |
|---|---|
| Foundations (1–5) | 9 hours foundations, including the Wednesday maths hour + 1 Tuesday hour |
| Normal learning (7–11, 13–15, 19–21, 23–25, 29–35) | 8 hours learning and building + 1 Tuesday hour + 1 Wednesday hour — ML design where scheduled, otherwise ML practice (Week 25: the Project #2 dataset choice) |
| Dedicated project (6, 16–17, 26–27, 37–39) | 9 hours project + 1 Tuesday hour |
| Week 36 | 8 hours project + 1 Tuesday hour + 1 design (the recorded Project #3 session) |
| Extension (5, 12, 22) and buffer (18, 28, 40) | As the week's own note says; DSA on its alternate weeks, and design where the track schedules a session |
Anything a week adds — a retake, a dataset choice, a deployment, a README — replaces work inside those ten hours. Nothing sits on top.
The minimum viable week
Weeks 1–29: 3 hours on the current week's chapter or project — on a DSA week, one of them is the DSA session. Everything else — the Wednesday hour, the reading, the weekend build — pauses without guilt.
Weeks 30–40: 3 hours on Project #3, one of them DSA on a DSA week — evaluation-safe work only; a collapsed week is not the week to touch the held-out set.
A real interview loop takes priority within the same weekly cap — pause or reduce ML study to make room.
One collapsed week is not slippage and does not trigger the cut-list. Two consecutive collapsed weeks do. The one exception to "no guilt": a collapse on a project week (5, 14–15, 23–24) still costs that project its Saturday, and the following buffer absorbs it — do not let it quietly eat the stretch goals.
How to use AI while learning
You attempt it ↓ get stuck ↓ ask AI ↓ understand answer ↓ implement ↓ modify it yourself
DSA problem-solving: no AI during the attempt — not for hints, not for "explain the problem", not for generating or completing a solution, and not during a re-solve from memory. Timebox → human-written editorial → then AI as a coach only: critique your edge cases, show alternative approaches, generate two or three more problems of the same pattern. It never writes the solution and never participates mid-attempt. The screen you are training for gives you a shared editor and a person; the reps only count if they happen under those conditions.
System design: you design alone first, every time. The AI's only role is interviewer — it asks the follow-ups and, at the end, lists what you missed. It never proposes a component, never draws the diagram, never "improves" yours. If it starts to, stop the session and restart it. Presenting Project #3 as a design answer is the one exception where it may also point out what a real interviewer would ask about your system — that is still attacking your design, not producing one.
Professional mode is a separate skill and an increasingly explicit hiring signal: scoping work for agentic coding tools, giving them the right context, reviewing what comes back, and knowing where they are reliably wrong. This is most of what "AI-capable software engineer" means to the people writing the job descriptions — and this plan builds AI systems without ever teaching you to work this way.
Practise it on the parts of your projects that are not the learning objective: scaffolding, Dockerfiles, CI configuration, test boilerplate, README drafts. Never on the model code, the retrieval logic, or the evaluation suite — those are the parts you must be able to defend line by line.
The evaluation suite is the sharpest case, and the reason is structural: if the same tool writes both the system and the tests that judge it, the tests confirm what you built rather than what you needed. That is a circle, and it produces a suite that passes while measuring nothing. Week 33 is the week it would cost you most — that suite is what the entire flagship rests on.
It costs no hours; it is a decision about how you work. And it is the one AI skill that pays off on both branches, whether or not the ML search converts.
Progress checkpoints
Every checkpoint has an externally verifiable artifact. "I feel like I understand it" is not a checkpoint — at 11pm on a Sunday you will always feel like you understand it. Calendar weeks and months below count from Part 2 Week 0, which is set after the Part 1 readiness check.
dsa/solved.md with every scheduled DSA session logged.v0, with an honest limitations section.system-design/wNN.md notes committed (Weeks 14–15).v0 — with the recorded Week 36 design presentation as its architecture segment. DSA: every scheduled session logged, one every other week since Week 1. Design: twelve notes, including three Project #3 presentations on record.