Model Caching Feature Announced for Amazon SageMaker HyperPod
Amazon Web Services has launched a model caching feature that enables the pre-loading of model weights and container images for large language models running on HyperPod.
Amazon Web Services has announced the new model caching feature for Amazon SageMaker Inference on HyperPod to eliminate long startup times experienced during the deployment processes of large language models.
Startup Latencies in Large Language Models
When large language models are deployed on Amazon SageMaker HyperPod, a specific time interval occurs between the pod request and when it becomes ready to handle traffic. This latency process is mainly caused by two major download operations.
Download Processes and Network Dependencies
The latency period is directly related to downloading container images from Amazon Elastic Container Registry and model weights pulled from storage sources such as Amazon S3, FSx for Lustre, or HuggingFace Hub.
Challenges with High-Volume Models
While small-sized models can become ready within a few minutes, massive models like DeepSeek-R1, which exceed 600 gigabytes, can take over thirty minutes to become operational.
Advantages of the Model Caching Solution
The newly introduced model caching feature eliminates the need for pods to download over the network by pre-loading model weights and container images onto cluster nodes.
Fast Startup with Local NVMe Storage
When the new feature is activated, pods can read from local NVMe storage, which offers speeds of approximately 7 gigabytes per second, allowing them to start handling traffic in seconds instead of minutes.