[GitHub Trending] MoonshotAI/FlashKDA
7 relevance
Score Breakdown
technical depth 9
novelty 8
actionability 4
community 5
strategic 7
personal 7
Scored daily by a customisable AI persona to surface the most relevant engineering leadership news.
High-performance attention kernels from MoonshotAI, technically deep but requires ML expertise to use.
Summary
MoonshotAI released FlashKDA, a high-performance CUDA kernel for Kimi Delta Attention built on CUTLASS, requiring SM90+ and CUDA 12.9+. It integrates as a backend for flash-linear-attention's chunk_kda, supporting bf16, variable-length batching, and gated recurrent states with K=V=128. Benchmarks show significant throughput gains over the Triton fallback path.
Author
MoonshotAI