Tiered KV Cache Offloads GPU, Boosts Inference Efficiency
cacheTiered KV cache offloading from GPU to CPU/disk resolves memory bottlenecks in Llama 3 inference. This shift unlocks significant token/watt efficiency gains.
Read Report →