Alora, the research verification company
Menu

2406.00007

Flash-style attention rewrite for SGLang

A serving-infra kernel for SGLang improves throughput on 8xA100 with reported 1.7x tokens/s versus stock.

UNTESTED STRONG

An SGLang attention kernel reports 1.7x tokens/s on 8xA100 versus stock.

structural read; full verdict pending

Holds when: 8xA100; SGLang

Evidence ledger