2406.00007
Flash-style attention rewrite for SGLang
A serving-infra kernel for SGLang improves throughput on 8xA100 with reported 1.7x tokens/s versus stock.
An SGLang attention kernel reports 1.7x tokens/s on 8xA100 versus stock.
structural read; full verdict pending
Holds when: 8xA100; SGLang
Evidence ledger
- S1
Present: A serving-infra kernel for SGLang improves throughput on 8xA100 with reported 1.7x tokens/s versus stock.
Present: https://github.com/example/sglang-attn
- S2
Present: arXiv subjects: distributed computing.
- S3
Present: Official repository listed.
Not available yet: Container dry-build is not available yet
- S4
Absent: No distinct-team reimplementation with matching numbers
- S5
Absent: No Hugging Face model card cites this paper.
- S6
Present: OpenAlex lookup recorded. Context classifier excluded from publish until 500 labels.
- S7
Absent: No OpenReview decision attached
- S8
Absent: No named-account repro thread
- S9
Absent: No independent harness match
- S10
Not available yet: S10 is not available yet
- S11
Not available yet: S11 is not available yet
- S12
Present: S12 band factor only. Prestige is not substantive evidence.
- S13
Absent: No public dataset link
- S14
Absent: No merged serving-framework PR
- S15
Absent: No Alora audit attached at extract
- S16
Not available yet: S16 is not available yet