Alora, the research verification company
Menu

2401.00001

Paged KV-cache for 13B-70B serving without quality loss

We show a paged KV-cache that cuts inference memory 2.4x on Llama-class 13B-70B models while matching baseline accuracy on long-context retrieval. Code and data are released. Seeds and variance are reported.

Grade B
PARTIALLY REPLICATED · Moderate

Paged KV-cache cuts inference memory 2.4x on Llama-class 13B-70B without quality loss on long-context retrieval.

Status PARTIALLY_REPLICATED. Confidence 72%. Band Moderate. Replicability grade B.

Holds when: holds when serving 13B-70B Llama-class with paged blocks

Evidence ledger