forked from Dao-AILab/flash-attention
-
Notifications
You must be signed in to change notification settings - Fork 169
Pull requests: vllm-project/flash-attention
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
Apply split-K to varlen paged KV when max_seqlen_q > 1 (spec decode)
#176
opened Aug 6, 2026 by
arbi-dev
Loading…
[WIP/RFC][SM120] Add bounded BF16 D256 paged-decode specialization
#170
opened Jul 22, 2026 by
goodnight654
•
Draft
Refactor FA3 onto the PyTorch stable API
#163
opened Jul 14, 2026 by
LucasWilkinson
Collaborator
•
Draft
[FA4] Support paged/decode compile specs in compile_flash_attn_varlen_func_from_specs
#158
opened Jul 9, 2026 by
sfc-gh-goliaro
Loading…
[Bugfix] Fix two SM80/SM120 forward kernel bugs: missing is_split_kv default, mDynamicCausal NameError
#156
opened Jun 30, 2026 by
tgmerritt
Loading…
feat(cute): add SM90 FP8 KV support with in-kernel dequantization
#147
opened Jun 12, 2026 by
qixiang-99
•
Draft
fix: handle suffix-less runtime arch (sm_103/GB300) in SM 10.x gate
#146
opened Jun 11, 2026 by
aoshen02
Loading…
Fix illegal memory access in FA2 varlen SplitKV early-exit LSE write
#139
opened May 18, 2026 by
wangyxbh
Loading…
[Perf] SM103 tcgen05.ld.red for fused TMEM load + row-max in softmax
#131
opened Apr 9, 2026 by
LopezCastroRoberto
Loading…
Combine kernel: increase pipeline depth from 4 to 8 stages
#124
opened Mar 4, 2026 by
jmkuebler
Loading…
Previous Next
ProTip!
Follow long discussions with comments:>50.