A delivery-time prediction model scored a mean absolute error of 3.2 minutes on its offline test set. In production it is running at 7.8 minutes: more than double.
The model has been deployed twice with the same result. The offline evaluation has been re-run and reproduces 3.2 minutes. The serving code has been reviewed by two engineers and looks correct. Traffic is normal and the input volume matches expectations.
The team is convinced the model is fine and something in production is broken, but they cannot find it.
Diagnose this. Be specific about what you would check, in what order, and what each result would tell you.
Build an interview-ready technical plan: frame the problem, explain the ordered approach, name the trade-offs, and show how you would validate the outcome. The AI reviewer grades that reasoning against this problem’s rubric.
Four guided sections · works on any screen