Efficient Memory Management for Large Language Model Serving with PagedAttention

13/09/2023 41 min

Listen "Efficient Memory Management for Large Language Model Serving with PagedAttention"

Episode Synopsis

The paper proposes PagedAttention, an attention algorithm inspired by virtual memory and paging techniques, to address the memory inefficiencies in large language model serving systems. The proposed system, vLLM, achieves near-zero waste in memory and improves throughput by 2-4 times compared to existing systems.

https://arxiv.org/abs//2309.06180

YouTube: https://www.youtube.com/@ArxivPapers

PODCASTS:
Apple Podcasts: https://podcasts.apple.com/us/podcast/arxiv-papers/id1692476016
Spotify: https://podcasters.spotify.com/pod/show/arxiv-papers

More episodes of the podcast Arxiv Papers