2401.00001
Paged KV-cache for 13B-70B serving without quality loss
We show a paged KV-cache that cuts inference memory 2.4x on Llama-class 13B-70B models while matching baseline accuracy on long-context retrieval. Code and data are released. Seeds and variance are reported.
Paged KV-cache cuts inference memory 2.4x on Llama-class 13B-70B without quality loss on long-context retrieval.
Status PARTIALLY_REPLICATED. Confidence 72%. Band Moderate. Replicability grade B.
Holds when: holds when serving 13B-70B Llama-class with paged blocks
Evidence ledger
- S1
Present: We show a paged KV-cache that cuts inference memory 2.4x on Llama-class 13B-70B models while matching baseline accuracy on long-context retrieval. Code and data are released. Seeds and variance are reported.
Present: https://github.com/example/paged-kv
- S2
Present: arXiv subjects: machine learning.
- S3
Present: Official repository listed.
Not available yet: Container dry-build is not available yet
- S4
Absent: No distinct-team reimplementation with matching numbers
- S5
Absent: No Hugging Face model card cites this paper.
- S6
Present: OpenAlex lookup recorded. Context classifier excluded from publish until 500 labels.
- S7
Absent: No OpenReview decision attached
- S8
Absent: No named-account repro thread
- S9
Absent: No independent harness match
- S10
Not available yet: S10 is not available yet
- S11
Not available yet: S11 is not available yet
- S12
Present: S12 band factor only. Prestige is not substantive evidence.
- S13
Present: Dataset link present
- S14
Absent: No merged serving-framework PR
- S15
Absent: No Alora audit attached at extract
- S16
Not available yet: S16 is not available yet