In large-scale AI inference services, when a cluster scales out, newly added inference instances must go through a series of cold-start steps before they can serve requests: container image preparation, file system mounting, runtime and inference framework initialization, and model weight loading. As image sizes and model weights continue to grow, the time spent on data preparation becomes increasingly prominent. During large-scale scaling events, many instances start simultaneously, further i...