Alora, the research verification company
Menu

2402.00002

INT4 post-training quantization for self-hosted 70B decode

GPTQ-style INT4 quantization on Llama-class 70B reduces inference cost with a 0.3 point MMLU drop. We release the conversion scripts.

UNTESTED STRONG

INT4 post-training quantization on 70B Llama-class reduces inference cost with a 0.3 point MMLU drop.

structural read; full verdict pending

Holds when: holds for post-training GPTQ-style INT4 on 70B

Evidence ledger