CPU-based autoscaling that works fine for a web tier quietly fails on GPU inference: wrong signal, and replicas that take minutes to warm. The interviewer wants the signals you actually scale on and how you hide the cold start.
Unlock the other 750 answers · ₹2,000 / $25Your progress and mastery stay saved · 6 months · one payment · no auto-renew
