-
Notifications
You must be signed in to change notification settings - Fork 700
Pull requests: fla-org/flash-linear-attention
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
[Fix] Eliminate per-call reference cycle in launch_grid_chunked
#1235
opened Sep 9, 2026 by
Ginray
Loading…
4 of 5 tasks
[Ops] Extend fixed-capacity CUDA Graph support for GDN and KDA
enhancement
New feature or request
needs-verification
Lacks real execution evidence (CI skipped / no before-after data)
waiting-author
Reviewer acted; ball is in the author's court
#1230
opened Sep 7, 2026 by
1418036869abc-cmd
Loading…
4 of 5 tasks
[Fix] Correct packed loss boundaries and cache defaults
bug
Something isn't working
#1217
opened Sep 3, 2026 by
taking-lying-flat
Contributor
Loading…
4 of 5 tasks
[Ops] Add Triton-Ascend compatibility for GDN-2
ascend-npu
Ascend NPU (triton_ascend) related
#1213
opened Sep 1, 2026 by
hazelduan
Loading…
3 of 5 tasks
Gate solve_tril dot precision on whether TF32 is faster, not whether it exists
bug
Something isn't working
waiting-author
Reviewer acted; ball is in the author's court
#1205
opened Aug 30, 2026 by
arbi-dev
Loading…
[Rotary] Separate varlen lengths from cache capacity
bug
Something isn't working
#1204
opened Aug 30, 2026 by
taking-lying-flat
Contributor
Loading…
4 of 5 tasks
[GDN] Add configurable output gate activation
enhancement
New feature or request
minor
Low-value-density change (typo/docs/small validation), batch-process
#1182
opened Aug 27, 2026 by
alifurkanstahl
Loading…
4 of 5 tasks
[Docs] Add software health index badge
minor
Low-value-density change (typo/docs/small validation), batch-process
#1153
opened Aug 19, 2026 by
Nayjest
Loading…
5 tasks done
[Ops] Use block matrix inversion to speed up solve_tril for the Ascend NPU backend
ascend-npu
Ascend NPU (triton_ascend) related
needs-verification
Lacks real execution evidence (CI skipped / no before-after data)
#1145
opened Aug 17, 2026 by
OsirisDuan
Contributor
Loading…
4 of 5 tasks
[Test] Add realistic benchmark input profiles
enhancement
New feature or request
#1144
opened Aug 17, 2026 by
0z5a
Loading…
5 tasks done
[Perf] Accelerate KDA training with TileLang
needs-verification
Lacks real execution evidence (CI skipped / no before-after data)
performance
#1128
opened Aug 14, 2026 by
lbx154
Contributor
Loading…
5 tasks done
[Cache] Memoize kernel config lookups to avoid per-launch file access
needs-verification
Lacks real execution evidence (CI skipped / no before-after data)
#1123
opened Aug 13, 2026 by
EastZeus
Contributor
Loading…
5 tasks done
[KDA] Warn on kwargs silently dropped by chunk_kda and fused_recurrent_kda
minor
Low-value-density change (typo/docs/small validation), batch-process
#1122
opened Aug 13, 2026 by
EastZeus
Contributor
Loading…
5 tasks done
[Perf] Parallelize AttnRes long-sequence reduction
needs-verification
Lacks real execution evidence (CI skipped / no before-after data)
performance
#1114
opened Aug 8, 2026 by
lbx154
Contributor
Loading…
5 tasks done
[KDA] Add FlashKDA CUDA training backend for chunk_kda
needs-verification
Lacks real execution evidence (CI skipped / no before-after data)
#1112
opened Aug 8, 2026 by
xy200303
Contributor
Loading…
5 tasks done
[GDN] feat(triton_ascend): add fused_recurrent GDN fwd kernel and dispatch
ascend-npu
Ascend NPU (triton_ascend) related
enhancement
New feature or request
#1078
opened Jul 30, 2026 by
OsirisDuan
Contributor
Loading…
[KDA] Add TLE inference backend for NVIDIA H800 with up to 1.41x speedup
needs-verification
Lacks real execution evidence (CI skipped / no before-after data)
performance
#1075
opened Jul 30, 2026 by
kuoihao
Loading…
[Perf] TileLang backends for linear-attention chunk backward (simple_gla / gla / delta_rule / gdn)
waiting-author
Reviewer acted; ball is in the author's court
#1046
opened Jul 19, 2026 by
markovchain-builder
Contributor
Loading…
[Feature] support reusable final_state buffer in fused recurrent GDN
enhancement
New feature or request
#1043
opened Jul 18, 2026 by
Bsgg1
Loading…
[GDN] Add batch-invariant mode: bitwise-reproducible prefill/decode across batch sizes
enhancement
New feature or request
#1003
opened Jul 4, 2026 by
yichuan-w
Loading…
[GDN2] Add fused BT=16 inference kernels for GDN2 prefill
enhancement
New feature or request
#990
opened Jun 29, 2026 by
Ghd-2077
Loading…
[Attn] Remove erroneous @torch.compile on ParallelAttentionFunction class
bug
Something isn't working
[Perf] Generalize fused q/k/v short convolution across layers
performance
#977
opened Jun 23, 2026 by
zhiyuan1i
Collaborator
Loading…
Previous Next
ProTip!
no:milestone will show everything without a milestone.