Alora, the research verification company
Menu

2205.00010

IO-aware attention for long sequences on a single GPU

An IO-aware attention tiling method reduces HBM traffic for long-context decode. Complements streaming softmax reductions.

UNTESTED STRONG

IO-aware attention tiling cuts HBM traffic on long sequences and composes with a streaming softmax reduction.

structural read; full verdict pending

Holds when: single GPU; long sequences

Evidence ledger