Tech
vLLM's weight cache can serve another checkpoint's weights when the tensor layout matches
TL;DR : vLLM 0.30.0 can keep a model's weights in a long-running daemon so that engines restart without reloading them ( load_format="ipc_cache" ). Before an engine uses those weights, both sides compare a fingerprint, and the docs describe the checkpoint part of it as "checkpoint content". In the code it is a hash of each shard's file name and safetensors header: tensor names, shapes, dtypes and offsets, but no tensor values. Two checkpoints with the same layout, such as a base model and a fine...
Read the full discussion on Dev.to
This article was aggregated from Dev.to. Click to join the conversation.
View on Dev.to