2107.00009
Online softmax as a streaming reduction for attention
We recast online softmax as a log-sum-exp streaming reduction. The method is scale-agnostic and ships with a reference kernel.
Online softmax can be computed as a streaming log-sum-exp reduction without a full materialization of the attention row.
structural read; full verdict pending
Holds when: scale-agnostic kernel
Evidence ledger
- S1
Present: We recast online softmax as a log-sum-exp streaming reduction. The method is scale-agnostic and ships with a reference kernel.
Present: https://github.com/example/online-softmax
- S2
Present: arXiv subjects: machine learning.
- S3
Present: Official repository listed.
Not available yet: Container dry-build is not available yet
- S4
Absent: No distinct-team reimplementation with matching numbers
- S5
Absent: No Hugging Face model card cites this paper.
- S6
Present: OpenAlex lookup recorded. Context classifier excluded from publish until 500 labels.
- S7
Absent: No OpenReview decision attached
- S8
Absent: No named-account repro thread
- S9
Absent: No independent harness match
- S10
Not available yet: S10 is not available yet
- S11
Not available yet: S11 is not available yet
- S12
Present: S12 band factor only. Prestige is not substantive evidence.
- S13
Absent: No public dataset link
- S14
Absent: No merged serving-framework PR
- S15
Absent: No Alora audit attached at extract
- S16
Not available yet: S16 is not available yet