Interview guide

Machine learning engineer interview guide

The ML engineer loop: fundamentals, ML system design, coding, and project deep-dives. What interviewers listen for and how to structure concept and design answers.

A machine learning engineer loop mixes four kinds of conversation: ML fundamentals (“explain X, when does it break”), ML system design (“design a ranker for Y”), coding (often data-shaped: streams, joins, sampling), and a deep-dive on a project from your résumé. Companies weight these differently. Product teams lean on design and the deep-dive; platform and research-adjacent teams lean on fundamentals and coding.

The trait that separates strong candidates is specificity. Weak answers talk about “the model” and “the data”. Strong answers name the retriever and the ranker, say what counts as a positive label, and give a number for the latency budget.

What the rounds look like

Fundamentals is a rapid conversation: bias–variance, regularisation, class imbalance, attention, evaluation metrics, calibration, drift. The interviewer is checking that you understand mechanisms, not that you can recite definitions. Every answer should end with “and here is when that stops working”.

ML system design is 45–60 minutes on one problem: a feed ranker, a fraud detector, a search relevance model, a recommendation system. The interviewer walks the stages in order — scope, pipeline, data and labels, features, model, training, serving, evaluation — and expects you to drive.

The project deep-dive is the round people under-prepare for. Pick one project, know its business metric and the ML proxy for it, the baseline, what you tried that failed, and what actually moved the number.

What interviewers listen for

How to structure an answer

For a concept question: the direct answer in one sentence, the mechanism in two, when it breaks, and what you reach for instead. Sixty to ninety seconds. Stop.

For a design question: state the scoping assumption a strong candidate states before designing (“I’ll assume we’re ranking posts from creators you don’t follow, optimising engagement as a proxy for retention”). Then give the headline architecture with one number (“two-stage: candidate retrieval, then a multi-task ranker, p99 under 200 ms at 100K QPS”). Then walk the stages in order, spending the most time where the interviewer pushes.

Questions to practise

  1. Explain the bias–variance trade-off and how you would diagnose which side you are on.
  2. A fraud dataset is 1 positive in 2,000. How do you train and evaluate a classifier?
  3. How does self-attention work, and why does a transformer need positional encoding?
  4. When would you choose a two-tower model over a cross-encoder for retrieval?
  5. What is training–serving skew and how does a feature store help?
  6. How do you evaluate a ranking model offline before an A/B test?
  7. Design the ranking system for a short-video feed.
  8. L1 versus L2 regularisation: what does each do to the weights and when do you prefer one?
  9. How would you detect data drift in production and decide when to retrain?
  10. Walk me through a project: what was the metric, the baseline, and what moved it?