2402.00002
INT4 post-training quantization for self-hosted 70B decode
GPTQ-style INT4 quantization on Llama-class 70B reduces inference cost with a 0.3 point MMLU drop. We release the conversion scripts.
INT4 post-training quantization on 70B Llama-class reduces inference cost with a 0.3 point MMLU drop.
structural read; full verdict pending
Holds when: holds for post-training GPTQ-style INT4 on 70B
Evidence ledger
- S1
Present: GPTQ-style INT4 quantization on Llama-class 70B reduces inference cost with a 0.3 point MMLU drop. We release the conversion scripts.
Present: https://github.com/example/int4-serve
- S2
Present: arXiv subjects: machine learning.
- S3
Present: Official repository listed.
Not available yet: Container dry-build is not available yet
- S4
Absent: No distinct-team reimplementation with matching numbers
- S5
Absent: No Hugging Face model card cites this paper.
- S6
Present: OpenAlex lookup recorded. Context classifier excluded from publish until 500 labels.
- S7
Absent: No OpenReview decision attached
- S8
Absent: No named-account repro thread
- S9
Absent: No independent harness match
- S10
Not available yet: S10 is not available yet
- S11
Not available yet: S11 is not available yet
- S12
Present: S12 band factor only. Prestige is not substantive evidence.
- S13
Absent: No public dataset link
- S14
Absent: No merged serving-framework PR
- S15
Absent: No Alora audit attached at extract
- S16
Not available yet: S16 is not available yet